vaultshell
Resolves secrets from 1Password via the authenticated op CLI using op://vault/item/field references; read-only access for injecting secrets into shell child processes at exec time.
Provides a credential-injecting reverse proxy for Stripe API calls, letting commands call http://127.0.0.1:<port>/v1/charges instead of https://api.stripe.com/v1/charges; injects the Authorization header in memory only and auto-stops after a TTL.
Resolves secrets from HashiCorp Vault via the authenticated vault CLI using vault://path#field references; read-only access for injecting secrets into shell child processes at exec time.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@vaultshellexecute the deploy script using the production secret profile"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
vaultshell
English | 简体中文
An MCP server that stores secrets in local secure storage, injects them into shell child processes at exec time based on rules, and guarantees secret plaintext never reaches the model context — not in tool responses, not in logs, not in audit records.
The closed loop: reference-based storage + conditional injection + mandatory output redaction.
Status: M3 — one-shot commands + persistent sessions + credential proxy + external CLI backends. See CHANGELOG and docs/design-spec.md.
The iron rule
No MCP tool returns secret plaintext — ever. Responses contain only names, masks (
[REDACTED:NAME]), and redacted output.Resolvers are only invoked inside the launcher; values go straight into the child process
envand never pass through the tool layer.shell_exechas no free-formenvparameter. Secrets enter by rule reference only.Redactor failures are fail-closed: output is discarded, never returned raw.
Related MCP server: Keiko
Quick start
Requires Node.js 20+.
Install from GitHub (no npm registry account needed)
# Run on demand directly from the repo (npm installs devDependencies and
# builds dist/ via the `prepare` script):
npx -y github:SoWhatI/vaultshell
# Or install globally:
npm i -g github:SoWhatI/vaultshellThis is the same code as the npm registry package, installed from git —
npm clones the repo, runs prepare (→ npm run build) and links the
vaultshell bin. Pin a tag for reproducibility:
npx -y github:SoWhatI/vaultshell#v0.1.1.
Docker (ghcr.io)
docker run -i --rm \
-e MASTER_KEY=<64 hex chars> \
-v ~/.vaultshell:/home/node/.vaultshell \
ghcr.io/sowhati/vaultshell:latestMCP client config using Docker:
{
"mcpServers": {
"vaultshell": {
"command": "docker",
"args": [
"run", "-i", "--rm",
"-e", "MASTER_KEY=<64 hex chars>",
"-v", "~/.vaultshell:/home/node/.vaultshell",
"ghcr.io/sowhati/vaultshell:latest"
]
}
}
}The image runs as the non-root node user (HOME=/home/node), so the data
volume mounts at /home/node/.vaultshell. node-pty is excluded from the
image — sessions automatically fall back to pipe mode (pty: false) in
containers; injection and redaction are unchanged. The web subcommand
passes through (docker run … ghcr.io/sowhati/vaultshell web — loopback
inside the container, of limited use).
From source
# Run via npx (after publish) or from source:
npm install && npm run build
# 1. Create a master key for the encrypted-file backend
export MASTER_KEY=$(openssl rand -hex 32)
# 2. Create ~/.vaultshell/config.yaml and ~/.vaultshell/rules.yaml
# (full annotated examples: docs/user-guide.md)
# 3. Store a secret (never echoed back) and run with injection
# via your MCP client's tools: secret_set, then shell_execClaude Desktop (claude_desktop_config.json):
{
"mcpServers": {
"vaultshell": {
"command": "npx",
"args": ["-y", "vaultshell"],
"env": { "MASTER_KEY": "<64 hex chars>" }
}
}
}Full setup, configuration reference, rule-writing guide, client integration, and FAQ: docs/user-guide.md.
MCP tools
Tool | Purpose | Returns |
| Run a one-shot command with per-rule injection. Params: |
|
| Open a persistent session; secrets injected once at creation |
|
| Send a command to a session |
|
| Kill a session; secrets die with the process |
|
| List live sessions | metadata only |
| Kill and rebuild the session without secrets |
|
| Start a credential-injecting reverse proxy on 127.0.0.1 |
|
| Stop a proxy |
|
| List running proxies | metadata only (incl. secret name, never value) |
| List names + metadata (backend ref, resolvable) | never values |
| Write to backend; persisted, never echoed |
|
| Delete and unregister |
|
| Check a ref resolves |
|
| Rules + static warnings (e.g. inject-everywhere) | — |
| Static findings: unknown secrets, unreachable rules, inject-everywhere, union risk, requireConfirm capability |
|
| Read audit entries | names + redacted commands only |
Execution limits (M3)
dryRun: trueonshell_execreports the would-beruleId,injectednames, final env key list (never values) and the deny-list verdict — without executing anything.timeoutSeconds(defaultdefaults.execTimeoutSeconds= 120, capped atkills the command and returns
timedOut: true.
defaults.maxOutputBytes(default 1 MiB) truncates each output stream and setstruncated: true. Truncation happens after the Redactor — it can never bypass redaction.defaults.maxConcurrentExecs(default 8) rejects excess concurrentshell_execcalls with a clear error (no queue).All of the above are recorded in the audit log (
dryRun/timedOut/truncatedflags).
Web config UI
vaultshell web # loopback-only HTTP UI on a random port
vaultshell web --port 5317Prints a one-time-token URL like http://127.0.0.1:5317/?token=… — open it
to manage secrets (write-only), edit rules with live static-validation
warnings and a matcher dry-run, edit config, and browse the redacted audit
log. The token dies with the process; every API call needs it, mutations
require same-origin application/json requests, and no endpoint can ever
return a secret value. Do not port-forward it. Design & threat model:
docs/web-config-ui-design.md.
Credential proxy (M3)
For high-sensitivity API tokens, prefer not landing the secret in env at
all. Declare a proxy in config.yaml:
proxies:
- id: stripe
upstreamHost: https://api.stripe.com
secretRef: STRIPE_KEY # registry name or full ref
headerTemplate: "Authorization: Bearer ${value}"Then shell_proxy_start { "id": "stripe" } → { port, expiresAt }, and
commands call http://127.0.0.1:<port>/v1/charges instead of
https://api.stripe.com/v1/charges. The proxy injects the header in
memory only and auto-stops after its TTL (default 300s). Audit records the
host and secret name, never the header value.
Boundaries: http:///https:// upstreams only (proxy→upstream TLS is
properly validated; no CONNECT tunneling, no TLS termination); the
client→proxy leg is plaintext on 127.0.0.1, which inherits the same-UID
local-process boundary from the threat model.
Persistent sessions (M2)
shell_session_open spawns a long-lived shell with the matched rule's
secrets injected at creation. The session is recycled after an idle TTL
(defaults.sessionTtlSeconds, default 900s, overridable per rule via
ttlSeconds or per call) — when the process dies, the secrets die with it.
shell_session_revoke kills a possibly-compromised session immediately and
rebuilds one without any secrets.
Sessions use a real PTY via node-pty
(an optional dependency — native build, needs Xcode CLT on macOS /
build-essential + python3 on Linux). If node-pty cannot be loaded or spawned,
vaultshell automatically falls back to a plain child_process shell and the
shell_session_open response says pty: false; injection and redaction
behave identically, only interactivity is degraded. Session output merges
stdout/stderr (PTY semantics) and passes through the same Redactor.
Backend support matrix
Backend | Ref scheme | Status | Platforms | Notes |
Encrypted file |
| ✅ Implemented | all | AES-256-GCM; master key from |
Local keychain |
| ⚠️ macOS only | macOS | spawns |
Process env |
| ✅ Implemented | all | read-only; good for CI |
File |
| ✅ Implemented | all | first line; refuses permissions > 0600 |
Inline |
| ⚠️ Restricted | all | plaintext in config; disabled by default, startup warning when enabled |
1Password |
| ✅ via CLI | all | needs authenticated |
Vault |
| ✅ via CLI | all | needs authenticated |
Infisical |
| ✅ via CLI | all | needs authenticated |
Doppler |
| ✅ via CLI | all | needs authenticated |
Details, per-backend setup, and how to register a plugin resolver: docs/backends.md.
Threat model
Threat | Mitigation |
Model-context leakage | Structural: values never cross the tool layer; only names and masks |
Command output echoing secrets | Mandatory redaction (exact + URL/Base64 variants + generic patterns), fail-closed |
Agent dumping env ( | Hard-blocked deny-list (configurable via |
Same-UID local process reading | OS-level limit, cannot be fully fixed; per-command injection shrinks the window. Explicit boundary. |
Long secret residency | Per-command injection by default; sessions have idle-TTL recycling + instant |
Misconfigured rule causing inject-everywhere |
|
Config files leaking | Config stores refs only, never values; |
Full discussion: docs/design-spec.md §7. Report vulnerabilities privately per SECURITY.md.
Data layout
~/.vaultshell/
config.yaml # main config
rules.yaml # injection rules + secret refs (refs only, never values)
secrets.enc # encrypted-file backend (0600)
audit/ # JSONL audit, redactedVAULTSHELL_HOME overrides the data directory (used by tests).
Development
npm run build # tsc
npm test # vitest — redactor property tests, backend round-trips,
# rule matching, shell_exec end-to-endSee CONTRIBUTING.md — including the security red lines every PR must respect.
License
MIT © 2026 vaultshell contributors
Available Tools
16 toolsaudit_queryB
Query recent audit entries (newest last). Entries contain names and redacted commands only, never values.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | max entries, default 50 |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description bears the full behavioral burden. It usefully discloses ordering (newest last) and the redaction policy (names and redacted commands, never values), but it says nothing about permissions, rate limits, pagination, or whether the read is safe/reversible beyond the implied 'query' verb.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences with no filler. The ordering and redaction constraints are front-loaded and every clause carries information the agent needs.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a single-optional-param read tool with no output schema, the description discloses the return shape (names and redacted commands) and ordering, which is what an agent most needs. The only minor gap is that 'recent' is undefined and no time-window behavior is given.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The single parameter has 100% schema coverage including bounds (1-500) and default (50), so the schema does the heavy lifting. The description adds no additional meaning about the limit parameter, which matches the baseline of 3 when coverage is high.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Query) and resource (audit entries) with scoping ('recent') and ordering ('newest last'). It is clearly distinguishable from secret/shell/rule siblings by name, but it does not explicitly differentiate itself from any alternative or explain what an audit entry is.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description offers no when-to-use or when-not-to-use guidance, and does not reference any alternative such as secret_list or rule_list. Usage is only implied by the tool name and resource.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
rule_listA
List injection rules with static warnings (e.g. rules matching every directory → potential unintended full injection).
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden. 'List' reasonably implies a read-only operation, and the parenthetical discloses a non-obvious behavior — that results include statically computed warnings (e.g. rules matching every directory). However, it says nothing about permissions, pagination, or result size.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single, front-loaded sentence with no filler. The parenthetical example earns its place by illustrating the warning concept concretely.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a no-parameter read tool with no output schema, the description conveys what is returned (rules plus static warnings) well enough to call it correctly. It could go slightly further on output shape, but nothing essential is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so there is no parameter semantics to document; the baseline for a 0-param tool is 4. The description adds nothing misleading about inputs.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('List injection rules') and adds scope via 'with static warnings,' which tells an agent what the result set contains. It does not explicitly differentiate itself from the sibling rule_validate, but the verbs are distinct enough that confusion is unlikely.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No when-to-use, when-not-to-use, or alternative guidance is given. The agent must infer that this is the read/inspection counterpart to rule_validate purely from the name; nothing in the description confirms that or states prerequisites.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
rule_validateA
Statically validate rules.yaml and return findings: unknown secret refs, unreachable rules (shadowed by earlier ones), inject-everywhere rules (high severity), requireConfirm capability notes, union mergeStrategy risk notes.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does well: 'statically' signals no side effects or execution, and it discloses the specific finding categories (unknown secret refs, unreachable/shadowed rules, inject-everywhere, requireConfirm notes, union mergeStrategy risk notes) plus a severity note. It stops short of stating explicitly that nothing is modified.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence that leads with the action and then lists findings. The parenthetical catalog is dense but every item is informative, so little is wasted, though the enumeration could be shortened with a summary phrase.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema, so the description must convey what comes back; it does so by naming the finding categories and flagging high-severity and risk-note outputs. That is sufficient for an agent to interpret a call, with only minor gaps about result format/ordering.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so there is no parameter semantics to document; the schema coverage is 100% and the baseline for a no-param tool is 4.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Names a specific verb (statically validate) and resource (rules.yaml) and enumerates the exact class of outputs it produces. It is clearly distinguishable from siblings like rule_list, which merely enumerates rules rather than analyzing them.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The word 'statically' implies this is a non-executing, pre-flight check, which hints at when to use it, but there is no explicit guidance about when to run it versus rule_list or what preconditions apply. Usage is implied rather than stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
secret_deleteB
Delete a secret from the backend and unregister it from rules.yaml.
| Name | Required | Description | Default |
|---|---|---|---|
| name | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It usefully discloses a non-obvious side effect—unregistering from rules.yaml—but omits other critical behavioral context such as whether deletion is permanent, what permissions are required, or what happens if the secret is in active use.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, front-loaded sentence that efficiently conveys the core action and an important side effect with no wasted words.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple one-parameter delete tool with no annotations or output schema, the description covers the action and side effect but leaves significant gaps: no parameter details, no warning about irreversibility or permissions, and no return information.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema has zero description coverage for the single 'name' parameter. The description does not compensate by explaining what the name represents (e.g., its format, source, or relation to rules.yaml), leaving the agent without guidance beyond the basic type.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb (Delete) and resource (a secret), and adds the side effect of unregistering from rules.yaml. This clearly distinguishes it from siblings like secret_set (create/update), secret_list, and secret_probe, which serve different operations.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no explicit guidance on when to use this tool versus alternatives, when not to use it, or any prerequisites. It merely states what the tool does, leaving usage entirely to inference from the name.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
secret_listA
List registered secret names with metadata (backend ref, resolvable). Never returns values.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does disclose two meaningful behaviors: it returns names plus metadata fields (backend ref, resolvable) and it deliberately withholds secret values. It does not state read-only safety, permission requirements, or how many entries are returned, so it falls short of exhaustive.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single sentence that front-loads the action and resource, then adds the two facts an agent most needs (what metadata comes back, what is not returned). Every clause earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter list tool with no output schema, the description conveys the return shape (names + backend ref + resolvable) and the critical exclusion of values. It could go one step further by stating the read-only nature or pagination/size behavior, but nothing essential to correct invocation is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters (empty object schema), so there is no parameter semantics to explain and the baseline is 4. Nothing in the description misrepresents the parameter surface.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Names a specific verb and resource ("List registered secret names") and immediately bounds the result set with "with metadata (backend ref, resolvable)" plus the negative scope "Never returns values." This cleanly separates it from siblings like secret_probe (which presumably retrieves values) and secret_set/secret_delete.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The "Never returns values" clause implies the tool is for enumeration only and that value retrieval belongs to a sibling, but no alternative is named and no explicit when-to-use or prerequisite is stated. Usage is inferable rather than guided.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
secret_probeB
Check whether a secret ref can be resolved. Returns only {ok, masked}; masked never contains value fragments unless defaults.maskTail > 0.
| Name | Required | Description | Default |
|---|---|---|---|
| name | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden. It usefully discloses the return shape {ok, masked} and a masking constraint, but omits side-effect/read-only status, required permissions, and failure behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two concise sentences with zero waste. The purpose is front-loaded, and the return/masking constraint follows efficiently.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple one-parameter probe with no annotations or output schema, the description gives purpose and return shape. It remains incomplete on usage routing, permissions, and failure behavior, so it is only moderately complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema has only one required string parameter with 0% description coverage. The description adds that the parameter represents a secret ref, which is meaningful, but it does not explain format, validation, or examples.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: checking whether a secret ref can be resolved. It is clear on its own, but it does not name or distinguish itself from siblings such as secret_list or secret_set, leaving routing to inference.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No when-to-use guidance, prerequisites, or alternatives are given. The description says what the tool does, but not when an agent should choose it over secret_list or secret_set.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
secret_setB
Store a secret value into the configured backend. The value is persisted, never echoed back, and the name is registered in rules.yaml (ref only).
| Name | Required | Description | Default |
|---|---|---|---|
| ref | No | optional explicit ref, e.g. keychain://svc/account; default: configured backend | |
| name | Yes | secret name, e.g. PROD_DB_URL | |
| value | Yes | secret value; stored only, never returned |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden but does deliver real behavioral facts: the value is persisted, never echoed back, and the name is registered in rules.yaml as a ref only. It still omits overwrite semantics for an existing name and any auth/backend-failure behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two compact sentences, front-loaded with the core action, and each clause adds non-redundant information (persistence, non-echo, rules.yaml side effect).
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Adequate for a 3-param mutation tool with full schema coverage, but with no annotations and no output schema the description should cover more of the write's side effects and failure modes. It covers the headline behaviors without being complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the schema already documents name, value, and ref with examples. The description adds the ref-only registration detail but no field-level syntax beyond what the schema provides; baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ("Store") and resource ("a secret value") plus the destination ("configured backend"). The store/delete/list/probe distinction from siblings is obvious from the operation named, though no sibling is called out explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance on when to use this versus secret_probe, secret_list, or secret_delete, and no prerequisites such as auth or backend-configured state. Usage is only implied by the verb.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
shell_execA
Execute a one-shot shell command with secrets injected per matching rule. Output is redacted: injected secret values never appear in the response. Dangerous commands (env, printenv, export -p, /proc/*/environ, ...) are hard-blocked. There is no free-form env parameter; secrets are only injected by reference via rules.
| Name | Required | Description | Default |
|---|---|---|---|
| cwd | No | working directory (~ is expanded); default: server cwd | |
| command | Yes | shell command to execute | |
| profile | No | optional profile name for rule matching | |
| extraEnv | No | non-secret env vars, keys must be in defaults.extraEnvAllowlist |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does substantial work: it discloses output redaction of injected secret values, hard-blocked command patterns with examples, and the secret-injection model. It still omits exit-status behavior, timeouts, and the side-effect profile of arbitrary command execution, which matter for a mutation-capable tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four tight sentences, front-loaded with the core action and immediately followed by the safety-relevant behavior. No filler or repetition of the tool name.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 4-param command-execution tool with no annotations and no output schema, the description covers the safety profile (redaction, hard blocks, env-injection model) that an agent most needs. Remaining gaps are return-shape and failure-mode details, which are minor given the absence of an output schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the description adds meaning beyond the schema by explaining that secrets cannot be passed as env params and are injected only by rule reference — which is exactly the semantic governing 'profile' and 'extraEnv'. It does not, however, clarify rule/profile matching syntax or allowlist contents.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Execute') and resource ('shell command') plus a scoping qualifier ('one-shot') that cleanly separates it from the shell_session_* siblings. An agent can distinguish it from shell_proxy_* and session tools without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The 'one-shot' framing and the explicit constraint that there is no free-form env parameter (secrets only by rule reference) give clear context for when this tool is appropriate. It stops short of naming shell_session_open/send as the alternative for persistent or interactive work, so the routing is implied rather than stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
shell_proxy_listA
List running credential proxies (metadata only: id, port, upstreamHost, secret NAME, expiry, request count). Never includes values.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations to lean on, the description carries the full burden and does well: it declares the tool is a read-only enumeration and, critically, that output is metadata-only and 'Never includes values' — a security-relevant guarantee about secret redaction that structured fields do not convey. It stops short of stating pagination or whether stopped-but-registered proxies appear.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single sentence that front-loads the action and resource, then qualifies the return shape and the security constraint. Every clause earns its place with no redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Although there is no output schema, the description enumerates the returned fields (id, port, upstreamHost, secret NAME, expiry, request count) and the exclusion of values, which is exactly what an agent needs for a zero-parameter read tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so there is nothing to document; the baseline for a parameterless tool is 4. No parameter semantics are needed or missing.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('List') and resource ('running credential proxies') and even enumerates the returned metadata fields. It is immediately distinguishable from the sibling mutation tools shell_proxy_start and shell_proxy_stop.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The listing intent is self-evident from the name and verb, so usage is implied, but the description offers no explicit guidance on when to reach for this versus shell_proxy_start/stop or when to prefer it over alternatives. No exclusions or prerequisites are stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
shell_proxy_startA
Start a credential-injecting reverse proxy on 127.0.0.1 (random port) for a proxies entry in config.yaml. Point commands at http://127.0.0.1: instead of the real API host; the proxy injects the secret into the configured header in memory only. Auto-stops after its TTL (default defaults.proxyTtlSeconds=300s). The secret value is never returned anywhere.
| Name | Required | Description | Default |
|---|---|---|---|
| id | Yes | proxies entry id from config.yaml | |
| ttlSeconds | No | TTL override (seconds) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does so well: random localhost port, in-memory-only secret injection into a configured header, automatic stop after TTL (default 300s), and an explicit guarantee that the secret is never returned. These are exactly the behavioral traits an agent needs.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight sentences, front-loaded with the action and endpoint, then the injection model and lifetime. No sentence is filler; every clause adds operational information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema, and the description covers the key return caveat (secret never returned), but it doesn't explicitly state that the chosen port is surfaced to the caller—an agent must infer how to learn <port>. Otherwise complete for a two-parameter lifecycle tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% (baseline 3), and the description adds real meaning: `id` maps to a `proxies` entry in config.yaml, and `ttlSeconds` overrides a documented default of defaults.proxyTtlSeconds=300s. This goes beyond the schema's terse 'TTL override (seconds)'.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Start a credential-injecting reverse proxy') scoped to a `proxies` entry, which clearly distinguishes it from siblings like shell_proxy_stop and shell_proxy_list. An agent can identify exactly what this tool produces without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly tells the agent how to use the result: 'Point commands at http://127.0.0.1:<port> instead of the real API host.' This gives clear operational context, but it does not state when to prefer this over alternatives (e.g. shell_exec with a secret directly) or any preconditions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
shell_proxy_stopC
Stop a running credential proxy.
| Name | Required | Description | Default |
|---|---|---|---|
| proxyId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden, and it discloses almost nothing: whether the operation is idempotent, what happens to in-flight sessions or credentials, what errors occur if the proxy is not running, or permission requirements are all absent. 'Stop' signals a state change but no consequences.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single short sentence with no filler, and the verb and resource are front-loaded. Nothing is wasted.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a mutation tool with no annotations and no output schema, the description is too sparse: it omits where proxyId comes from, what stopping does to the proxy's credentials/sessions, and what success or failure looks like. An agent can attempt the call but cannot predict its effects.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The single parameter proxyId has 0% schema description coverage, and the description only implies it identifies the running proxy without giving format, source (e.g., shell_proxy_list), or behavior for an invalid id. This leaves the parameter under-documented.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Stop) and resource (running credential proxy), which lets an agent distinguish it from shell_proxy_start and shell_proxy_list. It does not explicitly name or contrast those siblings, so it falls short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'a running credential proxy' implies the proxy must be active, but there is no guidance on when to stop versus start/list, no prerequisites, and no alternatives named. Usage must be inferred from the tool name alone.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
shell_session_closeB
Close a session: kill the process; injected secrets die with it.
| Name | Required | Description | Default |
|---|---|---|---|
| sessionId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does disclose the key destructive side effect: the process is killed and injected secrets are destroyed. It omits other behavioral facts an agent needs, such as permissions required, behavior on an already-closed session, and whether the operation is idempotent.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with zero filler: the action comes first and the consequence follows. Nothing could be removed without losing signal.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a one-parameter tool with no annotations and no output schema, the description covers the primary destruction semantics but leaves error cases, auth requirements, and confirmation of success unexplained. It is minimally adequate rather than complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
One required parameter with 0% schema description coverage, and the description never names or explains sessionId. The phrase "Close a session" makes the parameter's meaning inferable, but no format, source, or valid-value context is added beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ("Close a session") plus a concrete consequence (process killed, injected secrets die). However, it does not distinguish this from the overlapping sibling shell_session_revoke, leaving ambiguity about which teardown operation to pick.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no explicit when-to-use guidance, no prerequisites, and no named alternative. An agent must infer from the description that this is the terminal teardown step, and gets no help choosing between it and shell_session_revoke.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
shell_session_listA
List live sessions (metadata only: id, cwd, injected names, pty, ttl, timestamps).
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It does add one genuinely useful behavioral fact — 'metadata only' — signalling that session output/buffers are NOT returned here, which steers the agent to shell_session_send for content. However, it says nothing about permissions, ordering, or whether sessions from other users appear.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with the scope stated before the parenthetical detail. Every element earns its place; there is no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-param, read-only listing tool with no output schema, the description compensates by enumerating the returned metadata fields (id, cwd, pty, ttl, timestamps), which is the key missing structured information. Only minor gaps remain (auth/scope, ordering).
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so per the rubric the baseline is 4. Nothing in the description is needed to explain inputs, and it adds the useful note that outputs are metadata fields rather than live data.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('List') and resource ('live sessions'), and the parenthetical enumerates exactly what is returned (id, cwd, injected names, pty, ttl, timestamps). This cleanly separates it from the mutating siblings shell_session_open/send/close/revoke without needing their schemas.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is only implied: an agent can infer this is the discovery step before send/close/revoke, but the description never states when to call it versus alternatives or any preconditions. No explicit when/when-not guidance is present.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
shell_session_openA
Open a persistent shell session (PTY when node-pty is available, plain child_process otherwise — the response's pty field says which). Secrets are injected once at creation per matching rule; idle sessions are killed after ttlSeconds (rule ttlSeconds overrides defaults.sessionTtlSeconds).
| Name | Required | Description | Default |
|---|---|---|---|
| cwd | No | working directory (~ is expanded); default: server cwd | |
| profile | No | optional profile name for rule matching | |
| ttlSeconds | No | idle TTL override (seconds) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does substantial work: it discloses the PTY-vs-child_process fallback plus the `pty` response field, that secrets are injected once at creation per matching rule, and idle-TTL termination semantics. Missing only auth/permission requirements and failure behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two dense sentences with the core action front-loaded; the PTY parenthetical and the TTL precedence note carry real information. Slightly heavy on embedded parentheticals but no wasted prose.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a session-opening tool with no annotations and no output schema, the description covers lifecycle, secret injection, and TTY behavior well, but omits how the session is subsequently driven (shell_session_send) and any authorization constraints, leaving the agent to infer the interaction model.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3. The description goes beyond it by explaining ttlSeconds precedence (rule ttlSeconds overrides defaults.sessionTtlSeconds), which adds meaning not present in the schema field text.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: 'Open a persistent shell session.' The word 'persistent' implicitly distinguishes it from the one-shot sibling shell_exec, but it never names or contrasts with any sibling explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No when-to-use guidance is given. With shell_exec, shell_session_send, and shell_session_close in the sibling set, the agent must infer from the tool name alone that this is the session-establishment step rather than one-shot execution. No exclusions, prerequisites, or alternatives are stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
shell_session_revokeA
Immediately kill a session and rebuild a fresh one WITHOUT any injected secrets. Use when a session may have been compromised. Returns the new sessionId.
| Name | Required | Description | Default |
|---|---|---|---|
| sessionId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden, and it does disclose strong behavioral facts: the kill is immediate, the replacement session has no injected secrets, and a new sessionId is returned. It omits whether the old sessionId is invalidated, whether any state/workspace carries over, and permission requirements.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, front-loaded with the action and the secrecy constraint, then usage, then return value. No filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a single-parameter tool with no output schema, the description covers action, condition, side effect (no secrets), and return value. It leaves open the fate of the old session and whether dependent state is inherited, but the essentials are present.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% for the single sessionId parameter, and the description only mentions 'sessionId' as a return value, not clarifying that the input is the compromised session to be destroyed. The intent is inferable but not explicit.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('kill a session and rebuild a fresh one') plus a distinguishing scope qualifier ('WITHOUT any injected secrets'), which separates it from the plain close/list siblings. An agent can tell what this does without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives a clear triggering condition ('Use when a session may have been compromised'), which is exactly the context an agent needs. It does not name alternatives such as shell_session_close or say when NOT to use it, so it falls short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
shell_session_sendA
Send a command to a persistent session. Output (merged stdout/stderr) is redacted before returning. Dangerous commands are blocked per security.dangerousCommands. Returns {output, exitCode}.
| Name | Required | Description | Default |
|---|---|---|---|
| command | Yes | ||
| sessionId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations the description carries the full burden, and it does valuable work: it discloses that stdout/stderr are merged and redacted before return, that dangerous commands are blocked per security.dangerousCommands, and that the result shape is {output, exitCode}. It omits session lifecycle/timeout behavior and permission requirements, but the security-relevant traits are unusually well surfaced.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three compact sentences, front-loaded with the action and followed by security behavior and return shape. Every sentence carries information; nothing is filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description correctly summarizes the return value, and it covers the key behavioral caveats (redaction, blocking). The remaining gap is the relationship to sibling session tools (open/close/revoke) and what happens if the session is stale.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0% with two required parameters, yet the description adds little: 'persistent session' loosely hints that sessionId identifies an existing session, and 'command' is self-evident from context. It provides no detail on command format, shell semantics, or how sessionId is obtained.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Send a command to a persistent session'), which clearly implies it operates on an already-open session rather than a one-shot exec. It stops short of naming the sibling it differs from (e.g. shell_exec or shell_session_open), so an agent must infer the distinction.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is only implied: the required sessionId and the phrase 'persistent session' suggest a session must already exist, but the description never states the shell_session_open prerequisite or contrasts with shell_exec for stateless execution. No explicit when/when-not guidance is given.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
16 tool updates
v0.1.1- First observed
audit_query - First observed
rule_list - First observed
rule_validate - First observed
secret_delete - First observed
secret_list - First observed
secret_probe - First observed
secret_set - First observed
shell_exec - First observed
shell_proxy_list - First observed
shell_proxy_start - First observed
shell_proxy_stop - First observed
shell_session_close - First observed
shell_session_list - First observed
shell_session_open - First observed
shell_session_revoke - First observed
shell_session_send
TDQS
Scored across 16 tools
Most tools target distinct resources and operations, covering secret lifecycle, shell execution, persistent sessions, proxy management, rules, and audit. The only slight overlap is between secret_probe and secret_list, since both can report secret resolvability, but their descriptions clarify the intended use.
All tool names use snake_case with clear domain prefixes such as secret_, shell_, shell_session_, shell_proxy_, rule_, and audit_. The verb_noun structure is consistent throughout, with no mixed conventions or vague naming.
The server has 16 tools across five coherent subdomains, slightly above the typical 3-15 range but still well-scoped. Each tool appears to earn its place, with no obvious redundant operations.
Core secret, shell, session, proxy, and audit workflows are well covered, including session revocation and proxy lifecycle. Rule management is read-only via list and validate, with no create/update/delete tools, which is a minor gap likely handled through the underlying rules.yaml file.
Maintenance
Related MCP Connectors
Operate Linux, macOS and Windows from your LLM. Every action runs through an auditable allowlist.
Secrets for developers and agents—secure context and workflows without exposing secret values.
I run shell commands on your private cloud environment (bash, sh, zsh)
Scoped agent execution. Server-side credentials, policy, budgets and verifiable receipts.
Related MCP Servers
- AlicenseNot gradedqualityDmaintenanceA secure secrets management server that enables LLMs to execute CLI commands using injected credentials while protecting sensitive data through output redaction and user-approved session permissions. It features an encrypted vault, secret capture from command outputs, and a macOS menu bar app for native notifications and dialogs.MIT
- FlicenseNot gradedqualityCmaintenanceEnables AI agents to securely use secrets by running commands with environment-injected credentials and sanitizing output to prevent leakage.1-
- AlicenseNot gradedqualityBmaintenanceEnables agents to use secrets without exposing them, by injecting values into commands and redacting them from all outputs and transcripts.MIT
- AlicenseNot gradedqualityCmaintenanceEnables AI agents to execute shell commands on a host macOS computer with configurable file access permissions, granting access to credentials and tools while enforcing sandbox restrictions.MIT