Derbent
Use this server to share project-scoped notes and hand off tasks between coding agents.
Save notes with
memory_write, including title, body, tags, and optionalsupersedesto replace an older note.Search notes with
memory_searchby query, with optional limit andall_projects.Read a full note by id with
memory_read.Create a handoff for another agent label or
*withhandoff_create, including title, body, and tags.List handoffs with
handoff_list, filtering by state,mine, orall_projects.Take an open handoff by id with
handoff_take, reserving it for one agent.Mark a taken handoff done with
handoff_doneand an optional note.
Provides a guarded MCP server and pre-tool hook integration for GitHub Copilot CLI, routing its tool calls through Derbent's rule engine for allow/deny/ask decisions and receipt logging.
Derbent
One guarded pass for all your coding agents.
Every tool call that reaches Derbent, MCP or a CLI's built-in tools, is decided by one policy and written to a tamper-evident log, and the calls you care about wait for you.

Derbent guards against mistakes and prompt injection, and keeps an audit trail. It is not a sandbox: an agent that can already run shell commands as you can get around it (see Limits).
In Turkish history, a derbent was a guarded post on a mountain pass: its keepers decided who went through and kept a record of everyone who did. Derbent does the same for your coding agents' tool calls.
Claude Code, Codex, GitHub Copilot CLI and Antigravity CLI connect to Derbent as one MCP server, and your other MCP servers sit behind it. Each CLI's pre-tool hook sends its built-in tools, such as the shell and file edits, through the same gate. Every call is decided by your rules and written down.
Support for Codex and Antigravity CLI is experimental (configured from their docs, not yet checked in a real session).

It is one binary for Windows, macOS and Linux. There is no daemon and no network listener: a SQLite file is the only shared state (ADR 0001).
What you get
Rules that allow, deny or ask, by agent, tool and argument. The first match wins. The same rules apply to MCP tools and to built-in tools such as the shell, in all four CLIs, and a repository can add project rules that only make them stricter.
Approvals. A call your rules ask about waits until you press
ain the terminal UI, or is denied after 50 seconds by default.Receipts. Every call gets a hash-chained receipt, and
derbent verifynames the first receipt where the chain breaks after one was edited, moved, inserted or removed. Keep the head hash it prints to catch the newest ones being deleted too.
Also:
Shared memory and handoffs. Agents keep notes per repository and can leave tasks for each other.
Pins hold back a downstream server's tool when its definition changes.
Budgets stop an agent stuck in a loop.
Related MCP server: AgentGate
Install
Download the binary for your system and SHA256SUMS from the
latest release, check the hash, and put it
on your PATH as derbent (derbent.exe on Windows):
# Linux on amd64; the other files differ only in the end of the name
curl -LO https://github.com/tunahanaliozturk/derbent/releases/download/v1.0.0/derbent-v1.0.0-linux-amd64
curl -LO https://github.com/tunahanaliozturk/derbent/releases/download/v1.0.0/SHA256SUMS
sha256sum --ignore-missing -c SHA256SUMS
install -m 755 derbent-v1.0.0-linux-amd64 ~/.local/bin/derbentOr build it with Go 1.27 or later:
go install github.com/tunahanaliozturk/derbent/cmd/derbent@latestHomebrew (brew install tunahanaliozturk/tap/derbent) and Scoop install it too; see
Install and set up.
Every release can be rebuilt byte for byte from its tag. Install and set up has the other systems, how to check a release, and each CLI's entries by hand.
Quick start
derbent init --preset balanced # add Derbent to each CLI it finds, and write a first config
derbent doctor # check what init wrote
derbent # open the terminal UI in a terminal of its ownderbent init shows every change and asks once before it makes any. It copies each file it changes
first and never replaces an entry you already have. With --dry-run it only shows the changes. Two of
them, from a real run with the paths shortened:
claude: ~/.claude.json
copied first to ~/.claude.json.derbent-backup-20260930T205424Z
runs:
claude mcp add --scope user derbent -- ~/bin/derbent mcp --agent claude
derbent: ~/.config/derbent/config.toml
a new file
adds the balanced preset (derbent init --preset balanced --print shows it)It also adds the derbent gate hook to each CLI's settings; Install and set up shows
every entry. Start your agents as usual. Their calls now appear in the UI, and calls the rules ask about
wait there for you. Built-in tools says what each CLI's hook covers.
To put your other MCP servers behind the gate, so each agent needs only the one derbent entry, see
Downstream servers.
How a call is decided

Rules live in config.toml in your user config directory (%AppData%\derbent\ on Windows,
~/.config/derbent/ on Linux, ~/Library/Application Support/derbent/ on macOS). A downstream MCP
server's tools are named <server>__<tool>, and a CLI's built-in tools native__<tool>:
[[rule]]
tool = "github__issue_read"
action = "allow"
[[rule]]
tool = "github__*" # every other GitHub tool waits for you
action = "ask"
[[rule]]
agent = "claude"
tool = "native__Bash" # Claude Code's shell, through the hook
args = { command = "git push*" }
action = "ask"
[[rule]]
action = "allow" # the last rule has no conditionTo see how a call would be decided before it happens:
derbent explain --agent claude --tool native__Bash --args '{"command":"git push origin main"}'It prints each rule and why it matches or not, down to the first match, then the project's rules,
budgets, pin and grant, and ends with verdict: ask (rule:3) for this call under the rules above. Rules covers globs, the three presets (watch,
balanced, strict), project rules, budgets and rule suggestions.
Receipts

derbent verify # count, head hash, and whether the chain is intact
derbent receipts --agent codex --since 1h # also --tool, --project, --limit, --jsonA receipt keeps the arguments after masking secrets, and only the size and hash of what a tool returned (ADR 0004). Someone who can write the database could still rewrite the whole chain or delete the newest receipts, so keep the head hash somewhere else if you want to be able to tell later. Receipts and verify covers exports that can be checked without the database.
Commands
Command | What it does | More |
| Terminal UI: waiting calls, agents, live receipts, memory, grants | |
| Add Derbent to each CLI, then check the setup | |
| The MCP server each CLI starts | |
| The pre-tool hook for built-in tools | |
| Answer waiting calls from any shell | |
| List and take back session grants | |
| Show how a call would be decided | |
| Print the rules your answers point to | |
| Validate the config and list each server's tools | |
| See and accept changed downstream tools | |
| Read, export and verify receipts | |
| List the tasks agents left for each other | |
| Print the release |
Overhead
Measured on GitHub's hosted runners on 2026-10-01 (Linux on an Intel Xeon Platinum 8370C, Windows on an AMD EPYC 9V74, 4 vCPUs each), median of ten runs at p50, from docs/benchmark-results:
Linux | Windows | |
MCP tool call, direct to the server | 256.5 µs | 293.0 µs |
MCP tool call, through the gate | 756.0 µs | 909.0 µs |
Starting the binary and exiting | 4.292 ms | 40.40 ms |
Hook call, | 6.421 ms | 64.51 ms |
The gate adds 499.5 µs to an MCP call on Linux and 616.0 µs on Windows, for the extra stdio hop, the rule decision, the project rules check and the receipt written to SQLite. A hook call costs about 2.1 ms more than starting the binary on Linux and 24 ms more on Windows, where most of its cost is the process start. The numbers come from one run on shared runners, and runs differ by more than one run's intervals: the day before, the same code on the same Windows CPU model put the gate 459.5 µs over a direct MCP call, against 616.0 µs here. The results page has p99, calls per second and the caveats.
Limits
The whole list, with the reasons, is under Known limits and risks. The ones to know first:
Only calls that pass through the gate are seen. Tools a CLI never shows its hook, such as Codex's hosted web search, are outside it.
The CLIs marked experimental at the top have entries and hook adapters that follow each CLI's documentation and have not been checked in a real session. Claude Code and Copilot CLI (1.0.88, on 2026-10-02) have.
Argument globs match strings, not meaning:
git push*does not matchcd repo && git push.Approvals guard against mistakes and prompt injection inside MCP. They do not stop an agent that can already run shell commands as you: it can run
derbent approveitself.A receipt export shows its last run unchanged only against a head you kept, and never that nothing was left out of it.
A suggestion refuses the shells, launchers and operators it knows, but a text rule can be fooled and those lists cannot be complete, and an
allowfor a command prefix lets any options through.A handoff's address is a label, not an identity: any agent that can call
handoff_takecan claim every open handoff addressed to*.Approvals depend on you watching. Unattended,
askmeans denied after the timeout.Pins trust the first definition they see, including a new tool that an update adds to a pinned server, so name the tools you allow for a server whose updates you do not review.
A budget can be passed by the calls in flight at the same moment.
CI runs the tests on Windows and Linux and only builds on macOS.
For teams
A team audit trail, with receipts synced off each machine and an export for a SIEM, is an idea, not a plan. If you would use it, say so in this discussion.
Docs
Install and set up: every system, checking a release, each CLI by hand
Rules: rules, presets,
explain, project rules, budgets, suggestionsApprovals and the UI: keys, grants, answering from a shell
Built-in tools: the hook, tool names per CLI, what each CLI covers
Receipts and verify: what a receipt holds, exports,
verify --fileDemo: a transcript of two real Claude Code sessions sharing one gate
Design, the contract for the build, and the decisions behind it
Changes are listed in the changelog. To build, test or send a change, see CONTRIBUTING.md. To report a vulnerability, see SECURITY.md.
Licence
Apache 2.0. See LICENSE.
Available Tools
7 toolshandoff_createB
Leave a task for another agent working on this project, addressed to its agent label, such as reviewer, or to * for any agent. Returns the handoff's id.
| Name | Required | Description | Default |
|---|---|---|---|
| to | Yes | the agent label to hand the task to, such as reviewer, or * for any agent | |
| body | Yes | the task itself, up to 16 KiB | |
| tags | No | optional labels made of letters, digits, dashes and underscores | |
| title | Yes | a short summary of the task |
Output Schema
| Name | Required | Description |
|---|---|---|
| id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral disclosure burden. It only mentions the return value (handoff id), omitting side effects, persistence/visibility, permissions, error behavior, or limits for this write operation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences front-load the core action and target agent concept. No wasted words, though the return-value sentence is somewhat redundant given the output schema.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Full schema coverage and an output schema mean parameters and return values are covered elsewhere. For a write tool with no annotations, however, the description leaves gaps around behavioral side effects and sibling routing, making it only minimally complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so each parameter is already documented. The description repeats the 'to' semantics ('agent label ... * for any agent') but adds no additional meaning beyond what the schema provides, making the baseline 3 appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Leave a task') and resource ('handoff') with the target agent concept, so the action is clear. It distinguishes the tool from siblings like handoff_take or handoff_done only implicitly, without naming alternatives.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage for delegating work to another agent by label, which is helpful context. However, it gives no explicit when-to-use versus handoff_take, handoff_list, or handoff_done, and no exclusions or prerequisites.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
handoff_doneB
Mark a handoff you took as done, with an optional note on what you did.
| Name | Required | Description | Default |
|---|---|---|---|
| id | Yes | id of the handoff you took | |
| note | No | optional: what you did, up to 4 KiB |
Output Schema
| Name | Required | Description |
|---|---|---|
| id | Yes | |
| state | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden, yet it only restates what the schema already says ('id of the handoff you took', optional note). It does not disclose whether the state change is reversible, whether it requires ownership/permission, whether it is idempotent, or what effect it has on handoff_list results. For a mutation tool with zero annotation coverage this is a real gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single, front-loaded sentence that states the action, the target, and the optional input with no wasted words. Nothing is buried and nothing is redundant.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The output schema means return values need not be described, and both parameters are covered. However, for a no-annotation mutation tool the description omits prerequisites (must have been taken), failure modes, and post-completion effects, leaving it only minimally adequate.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and both parameters ('id', 'note' with its 4 KiB limit) are documented in the schema. The description adds no syntax, format, or constraint detail beyond that, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource ('mark a handoff ... as done') and scopes it to 'a handoff you took', which cleanly separates it from handoff_take, handoff_create, and handoff_list. It stops short of explicitly naming those siblings, so it lands at a clear-but-undifferentiated 4.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
'a handoff you took' implies the precondition that handoff_take must have run first, and 'with an optional note on what you did' clarifies the typical follow-up action. There is no explicit when-not guidance, no mention of the alternative of leaving a handoff open, and no statement of what happens if the handoff was never taken.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
handoff_listARead-only
List this project's handoffs addressed to you or to any agent, open ones unless state says otherwise. mine lists the ones you created instead.
| Name | Required | Description | Default |
|---|---|---|---|
| mine | No | list the handoffs you created instead of those addressed to you | |
| state | No | open (the default), taken, done or all | |
| all_projects | No | list every project's handoffs instead of only this one's |
Output Schema
| Name | Required | Description |
|---|---|---|
| notice | Yes | |
| handoffs | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, so the safety profile is covered. The description adds useful behavioral context the annotations do not: the default is open handoffs, the scope is this project unless all_projects, and it covers handoffs addressed to any agent. It does not describe result shape or ordering, which is partly excused by the output schema.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, zero filler, with the primary scope constraint front-loaded before the mine alternative. Every clause carries information the agent needs.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only list tool with a full output schema and 100% parameter coverage, the description supplies the defaults (open state, current project) that matter for correct invocation. Only result ordering/pagination behavior is unstated, a minor gap given the output schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all three parameters are already documented in the schema, including the meaning of mine and the allowed state values. The description reinforces the mine/state defaults but adds no syntax or format detail beyond what the schema provides; baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (List) and resource (handoffs) with explicit scope: this project, addressed to you or any agent. It also clarifies the default filter (open) and the alternate mode (mine), so an agent can tell what it returns. It stops short of explicitly naming the mutation siblings (handoff_create/take/done) it should not be confused with, but the read verb makes that distinction self-evident.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
"open ones unless state says otherwise" and "mine lists the ones you created instead" imply when the tool applies and how to flip behavior, but there is no explicit when-to-use/when-not framing or reference to the sibling tools (e.g. use handoff_take to claim one). Usage is inferable rather than stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
handoff_takeA
Take an open handoff addressed to you or to any agent, by its id, and read it in full. Only one agent can take a handoff.
| Name | Required | Description | Default |
|---|---|---|---|
| id | Yes | id of the handoff |
Output Schema
| Name | Required | Description |
|---|---|---|
| id | Yes | |
| to | Yes | |
| body | Yes | |
| from | Yes | |
| tags | Yes | |
| state | Yes | |
| taken | Yes | |
| title | Yes | |
| notice | Yes | |
| created | Yes | |
| project | Yes | |
| taken_by | Yes | |
| from_session | Yes | |
| taken_session | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden. It usefully discloses the exclusive-take semantics ('Only one agent can take a handoff') and the 'open' state precondition, plus that the result is the full content. It omits what happens on conflict (error vs. silent failure), whether the handoff's state changes, and any authorization requirements.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences, zero filler, with the action and the exclusivity constraint front-loaded. Nothing needs trimming.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists so return values need not be spelled out, and the single parameter is fully documented. The description covers purpose, precondition, and the key concurrency rule, leaving only edge-case behavior (already-taken, permissions) unaddressed.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
One parameter with 100% schema description coverage, so the baseline is 3. The phrase 'by its id' merely restates the schema field and adds no format, sourcing, or validity detail beyond it.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (take) and resource (handoff), scoped by id, and clarifies the ownership scope ('addressed to you or to any agent'). It does not name or contrast with siblings such as handoff_list or handoff_done, which would have made it a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied: you need an id, presumably from handoff_list, and the handoff must be open. There is no explicit when-to-use vs. when-not guidance, no mention of alternatives, and no statement of what to do if the handoff is already taken.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
memory_readARead-only
Read one saved note in full by its id.
| Name | Required | Description | Default |
|---|---|---|---|
| id | Yes | id of the note |
Output Schema
| Name | Required | Description |
|---|---|---|
| at | Yes | |
| id | Yes | |
| body | Yes | |
| tags | Yes | |
| title | Yes | |
| author | Yes | |
| notice | Yes | |
| project | Yes | |
| superseded_by | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and idempotentHint=false, so the safety profile is covered. The description adds one useful piece of behavioral context — 'in full' implies untruncated retrieval, unlike a search summary — but says nothing about what happens when the id is absent or invalid, which is the main failure mode for an id lookup.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single short sentence, front-loaded with the verb and resource, with no filler or redundancy. Every word contributes to identifying the operation.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return-value explanation is unnecessary, and the single required parameter is fully documented. The only real gap is the absence of routing guidance against memory_search, which is minor for such a simple tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% for the single 'id' parameter, so the schema already carries the parameter documentation. The description's 'by its id' merely restates it without adding format, range, or sourcing guidance, which is the baseline-3 case for a fully documented schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description gives a specific verb ('Read'), resource ('one saved note'), and scope ('in full by its id'), which distinguishes it implicitly from the sibling memory_search (query-based retrieval vs single-item retrieval). It stops short of naming the alternative, so it is clear but not fully sibling-differentiated.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is only implied: the phrase 'by its id' signals you must already have an id, which hints that this is the follow-up to a memory_search result. There is no explicit statement of when to use this versus memory_search, nor any exclusion or prerequisite.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
memory_searchARead-only
Search the notes agents have saved for this project. Returns ids, titles and snippets, best matches first.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | maximum number of results, 1 to 50, default 10 | |
| query | Yes | words to look for | |
| all_projects | No | search every project instead of only this one |
Output Schema
| Name | Required | Description |
|---|---|---|
| hits | Yes | |
| notice | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, so safety is covered. The description adds that results are ranked 'best matches first' and returns ids/titles/snippets, but the latter duplicates the existing output schema; no note on result limits or empty-result behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences, purpose and scope front-loaded, with zero filler or repetition.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With an output schema present and annotations covering the safety profile, the description needn't explain return fields; it covers purpose, scope and ranking. The only real gap is the absence of routing guidance against the sibling memory_read/write tools.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% with all three parameters documented, so the schema carries the load. The description's 'for this project' only loosely reinforces the all_projects default and adds no syntax or format detail.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Search) and resource (notes agents have saved for this project), plus what it returns. It is clear on its own but never names the adjacent siblings memory_read/memory_write, so an agent must infer the search-vs-read distinction.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The scope phrase 'for this project' implies when the tool applies, and the all_projects parameter implies the broader variant, but there is no explicit when-to-use, when-not, or named alternative. Usage is inferable rather than stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
memory_writeA
Save a note for the other agents working on this project: a decision, a finding, a convention. Returns the note's id. Pass supersedes to replace an older note.
| Name | Required | Description | Default |
|---|---|---|---|
| body | Yes | the note itself, up to 16 KiB | |
| tags | No | optional labels made of letters, digits, dashes and underscores | |
| title | Yes | a short summary of the note | |
| supersedes | No | id of an earlier note this one replaces |
Output Schema
| Name | Required | Description |
|---|---|---|
| id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations, so the description carries the full burden. It does disclose two useful behaviors: the return value (the note's id) and the supersession semantics ('pass supersedes to replace an older note'). It says nothing about permissions, durability, deduplication, or whether superseding makes the old note unreadable.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, front-loaded with the action and purpose, then the return value, then the one non-obvious parameter behavior. No filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Output schema exists and the description still confirms the id return, and all four parameters are schema-documented. Adequate for a write tool, though lifecycle details (what superseding does to the old note, whether notes persist) remain unstated.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so every parameter is already documented (body size cap, tag charset, supersedes meaning). The description only restates the supersedes intent, adding no format or constraint detail beyond the schema, so baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Specific verb+resource ('Save a note') plus the kinds of content the note holds (decision, finding, convention) and the destination (other agents on this project). It reads clearly as the write counterpart to memory_read/memory_search, but it never names a sibling to sharpen the boundary.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives practical when-to-use context by enumerating the note categories (decision, finding, convention) and implying a shared, cross-agent audience. It stops short of explicit alternatives or exclusions, e.g. when to prefer handoff_create or how it relates to memory_read.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
7 tool updates
v0.1.0- First observed
handoff_create - First observed
handoff_done - First observed
handoff_list - First observed
handoff_take - First observed
memory_read - First observed
memory_search - First observed
memory_write
TDQS
Scored across 7 tools
The seven tools split cleanly into two non-overlapping domains (handoffs and memory notes), and within each domain the operations target distinct lifecycle stages. handoff_create/take/list/done are unambiguous, and memory_read (by id) versus memory_search (by query) are clearly differentiated.
All names use a predictable snake_case namespace_verb pattern (handoff_*, memory_*), which is easy to scan. The only minor wrinkle is handoff_done, which is a state adjective rather than a verb like the others, but consistency overall is strong.
Seven tools is a well-scoped set for a coordination server: four handoff operations and three memory operations, each earning its place. Nothing is redundant or missing at the count level.
Coverage is solid: handoffs support create/list/take/done, and memory supports write/read/search plus supersedes for replacement. Minor gaps remain (e.g., no delete or list-all for memory notes, no cancel/release for a taken handoff), but agents can work around these.
Maintenance
Related MCP Connectors
Shared control plane for AI coding agents — tasks, memory, decisions, file locks. 12 tools.
See, price, and control every tool call your AI agents make: policy checks, cost, and audit tools.
Zero-trust gateway for AI agents: score tool calls, verify agent cards, enforce policy, audit.
AgentGuard — 20-tool AI safety MCP: policy preflight, risk scoring, audit logging, rate limits.
Related MCP Servers
- AlicenseAqualityAmaintenanceLocal zero-trust permission gateway for AI agents. Enforces policy-based tool authorization, human approvals, scoped permissions, and cryptographically verifiable audit logs.456 PyPI5Apache 2.0
- AlicenseAqualityAmaintenanceHuman-in-the-loop approval gateway for agent tool calls: agents request, policies decide, humans approve via Slack/Discord/web — with an OWASP-LLM-Top-10-tagged audit trail. Self-hostable.1018 npm33MIT
- AlicenseNot gradedqualityBmaintenanceSafe, reversible tool execution for AI agents. It sits between an agent and its tool servers, adding contracts, dry-run planning, policy, approvals, saga execution, and rewind.MIT
- AlicenseNot gradedqualityBmaintenanceGates agent tool execution with human approval, audit trails, and replay-resistant permits, enabling safe use of tools in agent loops.MIT