mcp-tool-bridge
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@mcp-tool-bridgeemail invoice INV-12 to its customer"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
mcp-tool-bridge
A generic Model Context Protocol server that exposes a catalog of declared tools to a model, filters them by role, validates every argument and holds sensitive calls behind a confirmation the model cannot give itself.
You declare what a tool is (its arguments, how much harm it can do, whether it can be undone, who may use it); the bridge enforces it on every call, whatever the model says.
Status: 1.1. Tool declarations, the registry, role-based access, argument validation, the bridge, the confirmation guard with lifetimes per class of tool, an audit log that keeps no content by default, the stdio server and three adapters (
envelope,adapt,importDefinitions), with a runnable example inexamples/minimal. The HTTP transport comes next.
Requirements
Node.js 20 or later. The package ships as both ES modules and CommonJS. Zod is optional:
install zod@^4 only if you use mcp-tool-bridge/zod.
Related MCP server: enterprise-agent-lab
A first look
import { defineTool, jsonSchema, ToolRegistry, text } from 'mcp-tool-bridge'
const sendInvoice = defineTool({
name: 'send_invoice',
description: 'Emails an existing invoice to the customer it belongs to.',
args: jsonSchema({
type: 'object',
properties: { invoiceId: { type: 'string', pattern: '^INV-[0-9]+$' } },
required: ['invoiceId'],
additionalProperties: false,
}),
sensitivity: 'high', // none | low | medium | high | critical
reversible: false, // an email cannot be unsent
roles: ['billing'],
summarize: ({ invoiceId }) => `Email invoice ${invoiceId} to its customer`,
handler: async ({ invoiceId }) => {
// `invoiceId` is a string here: the type comes from the schema above.
return text(`Invoice ${invoiceId} sent.`)
},
})
const registry = new ToolRegistry().register(sendInvoice)With Zod:
import * as z from 'zod'
import { zodSchema } from 'mcp-tool-bridge/zod'
const args = zodSchema(z.object({ invoiceId: z.string().regex(/^INV-[0-9]+$/) }))The bridge is the only way to run a tool. It works without any transport, which is how a host calls it directly and how the tests exercise it:
import { createBridge } from 'mcp-tool-bridge'
const bridge = createBridge({
registry,
context: (principal) => ({ db, tenantId: principal.id }), // what handlers get as call.context
})
const alice = { id: 'alice', roles: ['billing'] }
bridge.listTools(alice) // what the model may see
const outcome = await bridge.callTool(alice, {
name: 'send_invoice',
arguments: { invoiceId: 'INV-12' },
})
// send_invoice is high-sensitivity and irreversible: nothing ran.
// outcome.status === 'confirmation_required'
// outcome.confirmation = { token, expiresAt, tool, summary: 'Email invoice INV-12 to its customer' }
// Later, once a human has said yes in the host's own interface:
await bridge.executeConfirmed(alice, outcome.confirmation.token) // { status: 'ok', … }
// …or no:
await bridge.revokeConfirmation(alice, outcome.confirmation.token)Every outcome is a value, never an exception: ok, tool_error, invalid_arguments,
confirmation_required or rejected, each with the callId found in the audit log.
Serving over MCP
serveStdio serves the bridge over stdin/stdout, for one principal fixed when the
process starts:
import { parsePrincipal, serveStdio } from 'mcp-tool-bridge'
await serveStdio(bridge, {
info: { name: 'billing-tools', version: '1.0.0' },
principal: parsePrincipal({ id: 'alice', roles: ['billing'] }),
})For another transport, createMcpServer(bridge, options) returns the SDK's Server,
ready to connect.
On the wire:
tools/listreturns the principal's tools only. The server announcestools.listChangedand notifies the client whenever the registry changes.An unknown tool, or a tool the principal may not use, is a JSON-RPC error (
-32602 Unknown tool: …), the same for both. Invalid arguments and handler failures are results withisError: true, which the model can read and act on. A refused confirmation token reads the same whatever the reason: "This confirmation cannot be used. Nothing was done."A call that needs confirmation is resolved in one of two ways:
The client supports elicitation: the server asks the user through the client and runs the call on a yes. A no, or a dismissed question, withdraws it.
It does not, or elicitation is turned off (
elicitConfirmations: false): the result says that confirmation is pending, and the token travels in the result's_metaundermcp-tool-bridge/confirmation. To redeem it, the host repeats the same call with{ token }under the same key in the request's_meta, after a human said yes. The model has no way to do this: it writes arguments, not_meta, and a token placed in the arguments is ignored.
Install it in an MCP client such as Claude Code with
claude mcp add billing -- npx tsx path/to/server.ts.
Plugging in existing code
Most tools already exist somewhere, with their own input shape and their own way of reporting failure. Three adapters connect them without rewriting them:
envelope()reads a result envelope such as{ success, data, error }or{ ok, detail }. A failed envelope becomes aToolError, so the model reads its message; a successful one becomes the tool output.adapt()wraps an existing function. The bridge validates the arguments,inputmaps them to what the function expects, andoutput(often anenvelope) maps the result back.handler: adapt((action: LegacyAction) => legacy.execute(action), { input: (args, call) => ({ userId: call.principal.id, data: JSON.stringify(args) }), output: fromLegacy, }),importDefinitions()turns tool definitions written for an LLM API (Anthropicinput_schema, MCPinputSchema, OpenAIparameters) and one dispatcher into tools. What those formats do not say has to be declared, for every tool:const tools = importDefinitions(definitions, (name, args, call) => run(name, args, call), { search_orders: { sensitivity: 'none', reversible: true, roles: ['support'] }, refund_order: { sensitivity: 'high', reversible: false, roles: ['billing'] }, })A definition without governance, or governance for a name that has no definition, fails at startup with every name listed. When one tool is unclassified, none starts. Imported schemas are compiled in the same strict mode as
jsonSchema().
What the audit log keeps
An audit log that records arguments and results in full is a second copy of every email body, every value written to a spreadsheet, every document read. It is also the place nobody thinks of when data is deleted. So the default is the opposite: the audit log keeps no content at all.
Every event carries metadata only:
Field | What it is |
| Who called what, when, and through which channel. |
| SHA-256 of the canonical arguments. It tells two calls apart and links a call to its confirmation, without keeping what was sent. |
| The verdict: started, succeeded, failed, rejected, and why. |
| How long it took, and whether the result was an error. |
A tool that needs to keep some content says so, field by field, with JSON Pointers into its arguments and into its structured result:
defineTool({
name: 'read_email',
// …
audit: { args: ['/messageId'] }, // which message was read; never its content
})
defineTool({
name: 'search_email',
// …
audit: { result: ['/messages/*/id'] }, // which messages came back
})Kept fields appear in the event keyed by their pointer: "args": { "/messageId": "m-1" }. A * segment collects every match. Keeping content is a written decision,
in the declaration, visible in review. This is the same rule as governance: the default
path is the safe one, and a forgotten line costs a missing detail in the log, not a
leaked email.
Some guarantees hold whatever a tool declares:
Reduced on the way in. Arguments are reduced to their digest and their declared fields when the call's audit scope is built, before any event exists. The full value never lives in an audit event object, so it cannot resurface in an exception trace or a debug dump. Serialization only ever sees the reduced value.
Second line of defence. Inside a kept value, keys that look like secrets (
password,token,apiKey,authorization,secret,cookie…) are masked. Kept strings are cut at 500 characters, and at most 100 matches are kept per pointer.No summary, no text. The confirmation summary is written from the arguments ("Send 'Invoice 12' to ada@…"), so it goes to the host, never to the audit log. Result text and binary content are never kept, only declared fields of the structured result.
Every refused token is recorded with its reason:
unknown,consumed(a replay),expired,principal_mismatch,tool_mismatch,arguments_mismatch. These are the lines that show whether something is trying to force its way. The model gets one generic refusal and never learns which check stopped it.
Design decisions
At a glance
Tools are frozen descriptors that only the bridge can run. defineTool returns an
immutable description that holds no reference to the handler. The handler is kept
privately and only the bridge can reach it, so every execution goes through the same
checks: role, arguments, confirmation, audit. No code in the host can call a handler
by mistake, and no one can widen a tool's roles after its declaration has been checked.
roles is required and cannot be empty. A tool without roles would be open to
everyone or to no one. Either way, that is a decision someone must make, and make
visibly. Failing at startup turns a forgotten line into an error the developer sees
immediately. Otherwise it would surface in production as a tool silently open to all,
or silently dead.
The same function decides access for the list and for the call. The access rule
lives in one place, canAccess. The list hides what a principal cannot use. The call
checks again, because a client can send any tool name and roles can change in between.
Two implementations would end up disagreeing. Either the model is shown tools it cannot
call, or, worse, it can call tools it is never shown.
Why a declarative registry
A tool is two things: code that does something, and facts about that code. How sensitive
is it? Can it be undone? Who may use it? What arguments does it accept? In most agent
code these facts live in the head of whoever wrote the handler. At best they are
scattered: an if in the handler, a sentence in the system prompt, a comment.
Here they are fields of the declaration, and the governance fields (sensitivity,
reversible, roles) have no default. A tool nobody classified does not start. It
does not silently become "low sensitivity, everyone". Because the facts are data, the
bridge can apply one policy to every tool. It can also hand the facts to a reviewer, or
to an audit, as a table rather than as code to read.
The registry only accepts descriptors created by defineTool, which checks the whole
declaration at startup. A descriptor carries no handle on the handler: the only way to
run a tool is through the bridge and its checks. A copy made with { ...tool }
type-checks, but registering it is refused.
Why filter twice
The access rule is one function, canAccess(tool, principal): the principal must hold
at least one of the tool's roles. Names are compared exactly. There is no wildcard and
no hierarchy, so anyone can read who has access to what from the declarations alone.
The bridge applies this rule twice: when it lists tools for a principal, and again when it executes a call. Two reasons:
The list is a convenience, not a barrier. Hiding a tool keeps it out of the model's context. It saves tokens and avoids tempting the model with a tool it cannot use. But nothing forces a client to call only the tools it was shown. A model can hallucinate a name, a prompt injection can supply one, a client can be buggy or hostile. The check that matters is the one made when the call is executed.
Time passes between the two. Roles can be revoked and tools unregistered while a session is open. The decision is taken again, at the moment the effect would happen.
A call refused for lack of a role gets the same answer as a call to a tool that does not
exist (unknown_tool), so probing names reveals nothing. The audit log still records
the real reason.
Why validate on the server
The argument schema is declared once and serves two purposes: the JSON Schema shown to
the model, and the check enforced before the handler runs. Both adapters (jsonSchema
and zodSchema) derive the two from the same declaration, so they cannot drift apart.
The handler's argument type is inferred from the same source.
Validation is strict on purpose:
No coercion.
"3"is not a number. A model that gets a type wrong should be told so, not second-guessed.Defaults are not applied. In a JSON Schema,
defaulttells the model what happens if it leaves a field out. The handler decides what that means. The inferred type keeps defaulted properties optional, so the handler cannot assume they are present.Only plain JSON gets in. Arguments are deep-copied, and anything that is not JSON (dates, functions, class instances, cycles,
NaN) is rejected. The handler, the audit log and the confirmation guard each see a value that nobody else can mutate.Bad schemas fail at startup. JSON Schemas are compiled with Ajv in strict mode. Unknown keywords, unknown formats and
requiredproperties that are never declared are errors. A Zod type with no JSON equivalent (z.date(),z.bigint()) is refused too, because the model could never send it.
Rejected arguments come back to the model as a tool result, with one JSON Pointer and
one message per problem (/to: must match format "email"). The model can then correct
itself, and the handler never sees the bad input.
Why the confirmation lives on the server, not in the prompt
"Ask the user before sending anything" in a system prompt is a request made to the model, and the model is the very component the guard protects against. It can lose the instruction in a long context. A prompt injection can override it. It can judge that this case does not count. And nothing in the code enforces it.
The guard is code on the execution path. A tool whose declaration says
sensitivity: 'high' (or 'medium' and irreversible) returns confirmation_required
instead of running, and the handler cannot be reached without a valid token. The
decision comes from the declaration, not from the model's reading of the situation: the
same tool is always confirmed, or never.
The token is built so that approving one thing cannot authorise another:
Bound to the call. A token belongs to one principal, one tool and the exact arguments. Key order does not matter, values do. Approving "email invoice INV-12" cannot email INV-13, and a token presented by anyone else is refused.
Single-use. Redeeming takes the record out of the store in one atomic step, and that same step checks the expiry. Two concurrent redemptions cannot both run, and a second presentation is recorded as a replay (
consumed).Short-lived, by class. The lifetime can be a table per sensitivity and reversibility: an irreversible critical action deserves a shorter window than a reversible medium one, because past it the context of the decision is gone. The expiry is computed when the confirmation is issued and stored with it; changing the configuration later does not move confirmations already issued. Five minutes by default.
Never stored. The store keeps a SHA-256 hash. Whoever can read the store cannot redeem anything.
Checked again. At redemption the role is checked again and the arguments are revalidated against the tool registered at that moment. The arguments that run are the ones that were confirmed, frozen when the confirmation was issued.
Not for the model. The token is meant for the host. The MCP server carries it in the result's
_meta, which MCP clients are not expected to pass to the model, while the model only reads that confirmation is pending. The host redeems it after a human decision: by repeating the call with the token, by callingexecuteConfirmed()later, or through MCP elicitation when the client supports it. There is deliberately noconfirm_actiontool. A model under prompt injection would simply call it, and the guard would be reduced to a delay.
What the guard does not cover
The guard stops the model from running a sensitive call on its own. It does not make the system safe by itself. Each limit below comes with what you should put in place around it.
A compromised host or client. Whoever controls the MCP client can attach a token it was given and replay a decision: the guard protects against the model, not against the host. Put in place: run the client in a component you control, authenticate it (the
authenticatehook of the upcoming HTTP transport), and keep the tokens it receives in memory, out of logs and transcripts.A host that shows the token to the model. If a host copies
_meta, or the whole outcome, into the conversation, the model can confirm its own calls, and the guard is gone. Put in place: strip_metaand confirmation outcomes before anything reaches the model's context, and add a test that fails ifmtb_(the token prefix) ever appears in a transcript.Misclassified tools. Below the threshold, tools run directly. A tool declared
lowthat actually deletes data is not caught. Put in place: review declarations like permissions: printregistry.list()as a table (name, sensitivity, reversible, roles) in code review, and require a second reviewer for any new or reclassified tool.What the handler really does. The guard confirms a call, not the behaviour of the code behind it. Put in place: give each handler credentials scoped to the one effect its declaration describes, so that it cannot do more even by mistake.
Information leaving through reads.
sensitivity: 'none'tools run freely. A model can read data and repeat it elsewhere. Confirmation governs actions, not information flow. Put in place: restrict reads of personal or confidential data by role, return only the fields the task needs, and watch the audit log for unusual read volumes.The quality of the human decision. The guard makes sure someone confirmed. It cannot make sure they read what they confirmed. Put in place: write
summarizefor the person who confirms (what happens, to whom, with what consequence) and show it next to the arguments, never as a bare "Confirm?".A world that changed. Roles and arguments are checked again at redemption, but the invoice may have been paid in the meantime. Put in place: have handlers check their preconditions when they run (invoice still unpaid, slot still free) and throw a
ToolErrorotherwise, and shortenttlMswhere the context moves fast.Timeouts. A handler that ignores its abort signal keeps running after the bridge has given up. Its result is discarded, but its effects are not undone. Put in place: pass
call.signalto every I/O the handler starts (fetchand most database drivers accept one) and make side effects idempotent, so that a cancelled call either stops or can be retried safely.Several processes. The default store lives in one process. A token issued by one instance cannot be redeemed on another. Put in place: before running a second instance, implement
ConfirmationStoreover shared storage with an atomictake(RedisGETDEL, SQLDELETE … RETURNING).
Why the audit log records before running
Each call that runs leaves two events under one callId: call.started, then
call.succeeded or call.failed. A confirmed call is preceded by confirmation.issued
and possibly confirmation.declined, under the same id. Rejections are recorded with
their real reason (forbidden, invalid_arguments, an invalid or expired confirmation),
even when the model is told unknown_tool.
call.started is written before the handler runs, so the log knows about the call even
if the process dies during it. With auditFailure: 'block', a call or a confirmation
that could not be recorded does not happen. The default, continue, warns on stderr and
proceeds. Events that follow an effect (call.succeeded, call.failed) can only be
reported: the effect has already happened.
What these events contain is the subject of What the audit log keeps: by default, metadata only. Messages of unexpected exceptions go to the log only, cut to 500 characters; the model gets a generic failure. The default sink writes JSON lines to stderr, because under stdio, stdout is the protocol channel.
Declaring tools: reference
Field | Required | Meaning |
| yes | Unique in a registry. MCP allows |
| yes | What the model reads to decide whether and how to call the tool. |
| yes |
|
| yes |
|
| yes | Can the effect be undone? |
| yes | Non-empty. A principal needs one of them. |
| no |
|
| no | One sentence describing a specific call, shown to whoever confirms it. |
| no | What the audit log may keep: |
| no | Per-call time limit. |
| no | Human-readable name. |
Behaviour hints sent to MCP clients are derived from the declaration. readOnlyHint
is true for sensitivity: "none". destructiveHint is true for irreversible tools that
have an effect. Roles and sensitivity levels are never sent to the client.
Handlers throw ToolError for failures the model may read ("no invoice INV-12"). Any
other exception reaches the model as a generic failure; its details go to the audit log.
Roadmap
v1.0: declarations and registry, role-based access, validation (JSON Schema and
Zod), the confirmation guard with single-use tokens and MCP elicitation, the audit log,
the envelope adapter, the stdio transport, a three-tool example.
v1.1: adapt() and importDefinitions() (see
Plugging in existing code); confirmation lifetimes per
sensitivity and reversibility, an atomic take that reports why a token was refused;
an audit log that keeps no content unless a tool declares it (see
What the audit log keeps).
Next:
HTTP transport: Streamable HTTP, with a per-request
authenticatehook, Origin/Host checks and sessions bound to the principal who opened them. The deprecated HTTP+SSE transport will be available behindlegacySse: true.
Out of scope for now: an OAuth authorization server, MCP resources and prompts, retries and rate limiting.
Development
npm ci
npm run lint && npm run format:check && npm run typecheck
npm test
npm run build && npm run check:dist
npm run check:mutationsThe tests make no network calls.
The safety guards are checked by targeted mutation testing: npm run check:mutations removes each guard in turn and fails if the test suite still passes.
Eighteen guards are covered, among them:
the role is checked again at call time;
a confirmation token is single-use;
a token only runs the exact arguments it was issued for;
the role is checked again when a confirmation is redeemed;
the token never appears in what the model reads;
invalid arguments never reach a handler;
an imported tool without governance never starts;
the store checks the expiry in the same step as the take, and the bridge checks it again from the expiry stored at issue time;
the audit keeps no argument or result content unless the tool declares it, and no confirmation summary;
every refused token leaves its reason in the audit, and the model never learns it.
CI runs this on every push. See CONTRIBUTING for the mutants and the tests that catch them.
License
This server cannot be deployed
Maintenance
Related MCP Connectors
Runtime permission, approval, and audit layer for AI agent tool execution.
See, price, and control every tool call your AI agents make: policy checks, cost, and audit tools.
Find, vet, and run MCP tools through a secure audited gateway with prompt-injection risk scoring
Zero-trust gateway for AI agents: score tool calls, verify agent cards, enforce policy, audit.
Related MCP Servers
- AlicenseNot gradedqualityBmaintenanceGates agent tool execution with human approval, audit trails, and replay-resistant permits, enabling safe use of tools in agent loops.MIT
- FlicenseNot gradedqualityCmaintenanceEnables controlled AI-agent access to enterprise-shaped tools with a deny-by-default gated write path, human approval, dry-run execution, and append-only audit logging.1-
- AlicenseNot gradedqualityCmaintenanceProvides permission gates and tamper-evident audit logging for AI agent tool executions, with declarative policies, consent ladders, and hash-chained verification.MIT
- FlicenseNot gradedqualityCmaintenanceEnables agents to securely discover and invoke a centrally governed catalog of tools from distributed internal and external providers, with policy enforcement, quotas, inspection, and audit controls.-