Sentinel Memory MCP
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@Sentinel Memory MCPresume project api-review and show open tasks and next actions"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
Sentinel Memory MCP
An open-source cybersecurity MCP server for persistent graph memory, evidence-backed vulnerability triage, and authorized API security research. Connect it to Google Antigravity, OpenAI Codex, or another stdio MCP client to carry structured investigation state across sessions.
Remember what was tested. Preserve the evidence. Reject weak conclusions. Resume with the next useful action.
Quick start · Antigravity · Codex · Self-hosting · हिंदी · AI-readable overview
Release reality: the core runs locally and is tested with official MCP clients and a synthetic security lab. Actual IDE UI verification is separate. This is an early development release, not a certified production security product. Local models are optional and disabled by default after unsuccessful specialist-output benchmarks. Detailed status · Known limitations
Why this exists
Security investigations outlive a chat window. Repeated logs bury useful facts, a fixed issue can resurface as active, and a changed object ID can be mistaken for a vulnerability. Sentinel stores typed observations, hypotheses, tasks, evidence, relationships, and versioned decisions rather than treating an entire conversation as trusted memory.
When you need to… | Sentinel provides… |
Continue an API assessment after context loss | A compact resume packet with scope, tasks, findings, corrections and next actions |
Avoid repeating completed work | Test signatures including account context, ownership and deployment state |
Check a suspected IDOR/BOLA | Evidence requirements for ownership, expected access, protected data and false positives |
Keep sensitive evidence useful | Redaction, SHA-256 content addressing, provenance and integrity checks |
Explain a finding to engineers and managers | Technical and executive reports with explicit validation gates |
Use a small local model cautiously | Offline, process-isolated advice with schema checks and deterministic fallback |
Related MCP server: GoodMemory
Quick start
Prerequisites: Git and Node.js 24.14+. No cloud API key, Neo4j, Docker or model download is required for core features.
git clone https://github.com/vikrant-project/sentinel-memory-mcp.git
cd sentinel-memory-mcp
npm ci
npm test
npm run configureThe configuration generator uses your checkout's absolute paths and preserves other entries in existing MCP JSON files. Generated machine-specific files are ignored by Git.
Host | Next step | Guide |
Google Antigravity | Open this checkout, refresh MCP servers, look for | |
OpenAI Codex | Merge | |
Other stdio MCP client | Launch | |
Docker or private VPS | Keep stdio; persist the database |
Try it without touching a real target:
npm run demoThe demo starts a loopback-only lab, connects an official MCP client, captures controlled-account evidence, distinguishes a private authorization failure from a public object, and saves synthetic results to outputs/demo-summary.json. It shuts down its processes afterward.
Use Sentinel to open project api-review. Read security_skills and project_resume.
Do not send any requests yet. Help me record the assessment owner's authorized
scope, controlled accounts, expected access rules, and next validation steps.At the end of a session:
Save the current task, evidence references, unresolved hypotheses, and next
actions in Sentinel. Next session, resume project api-review from that state.The project does not independently establish permission to test any target. Scope must come from the assessment owner.
Architecture

flowchart LR
Host[Antigravity / Codex / MCP client] --> MCP[Official MCP SDK · stdio]
MCP --> Context[Context compiler]
Context <--> Graph[(SQLite graph + FTS5)]
MCP --> Scope[Scope and request controls]
Scope --> Evidence[Redacted evidence + SHA-256]
Evidence --> Triage[Evidence gates + false-positive checks]
Triage --> Reports[Technical / executive reports]
Graph --> Model[Optional local HF process]
Model -. advisory output only .-> ContextThe deterministic path handles persistence, retrieval and validation. The local model does not authorize requests, run commands, rewrite weights or confirm findings.
Evidence-first investigation workflow

Scope: authorized origins, accounts, environments, restrictions, limits and expiry.
Observe: expected behavior and minimal controlled test data.
Preserve: redacted evidence, provenance and immutable content hashes.
Validate: reproducibility, security boundary, impact, controls and false positives.
Report: FINAL only when implemented gates pass; otherwise retain DRAFT.
Resume: save task state and check prior signatures before repeating work.
Workflow | What it checks | Examples of rejected lookalikes |
Authorization / IDOR / BOLA | Ownership, unrelated account, expected access, protected response | Public objects, sharing, expected admin access, cached responses |
Rate limits | Sensitive operation, abuse feasibility, impact and controls | “Ten requests worked,” effective throttling, no security impact |
Business logic | Backend entitlement and documented workflow rules | Free features, trials, promotions, frontend cosmetics |
Data exposure | Sensitive fields, permitted caller and protected contents | Public data, sharing, expected privileges, non-sensitive fixtures |
These are structured analyst workflows, not universal vulnerability detectors. CONFIRMED means implemented evidence gates passed; imported evidence and semantic claims still need review.
MCP interface
20 tools · 4 resource templates · 6 prompts
Area | Tools |
Project and scope |
|
Graph and context |
|
Evidence and tasks |
|
Research and reporting |
|
Review and advice |
|
Resources expose context, scope, evidence and report snapshots. Prompts support investigation, validation, two-user comparison, reporting, false-positive review and regression checks. Payload examples →
Measured results, with context

Recorded on Windows, Node 24.14.1, Ryzen 7 7435HS and about 16 GB RAM. The benchmark uses 500 graph nodes and 100 samples per operation. The chart is generated from the committed benchmark JSON.
Check | Recorded result | Interpretation |
Automated suite | 28 passing tests in the recorded run | SDK subprocesses, restart, isolation, redaction, scope, lab, backup and model failure |
Synthetic triage | 32/32 expected classifications | Engineering fixtures, not independent real-world accuracy |
Confirmed precision / recall | 100% / 20% on those fixtures | Incomplete positive cases deliberately remain unconfirmed |
Repeated-history reduction | 99.32% estimated | Synthetic repeated logs; characters/4, not a tokenizer or lossless guarantee |
Local model schema acceptance | 0/4 for each candidate | Both disabled by default; failures and timeouts retained |
Do not extrapolate these numbers to unseen targets. Methodology · Raw results · Model results · 11 hardening iterations
Example deliverables
All committed example evidence is synthetic. Real databases, credentials, model weights, local caches and temporary files are excluded from version control.
Self-hosting and data ownership
Run beside your MCP client, through a Docker stdio process, or over authenticated SSH to your VPS. There is no public HTTP endpoint in this release. A website host or reverse proxy alone does not turn it into a remote MCP service.
The hosting guide covers Docker, volumes, SSH, restarts and backups. Docker/VPS recipes are marked unverified because Docker was unavailable on the original development machine.
Documentation map
Start here | Understand and extend |
FAQ
Is this a cybersecurity MCP server or a scanner?
It is an MCP server for security investigation memory, evidence handling and triage. Its networking tool performs one bounded authorized GET. It does not mass-scan, brute-force credentials or autonomously exploit targets.
Can I use it without a local LLM?
Yes. Default features use deterministic logic, SQLite, FTS5 and graph retrieval. Model downloads are opt-in.
Does it replace a security engineer?
No. Scope, policy, interpretation and business impact need an authorized analyst. Confidence probability is intentionally uncalibrated.
Does it work with Codex and Antigravity?
Both support local MCP configuration. Guides and generated examples are included. Automated tests verify official MCP clients; a specific IDE version's UI and restart behavior need a separate check.
Can it rank first in AI search?
No repository can guarantee that. Clear descriptions, useful examples, descriptive topics and accessible documentation help discovery. This project does not use fake ratings, hidden instructions or keyword spam. Discoverability notes
License and credits
Project code, original diagrams and documentation: MIT, © 2026 vikrant-project. Dependencies and model checkpoints retain their own licenses. Built with the official MCP TypeScript SDK, Node.js, SQLite, Zod and optional Hugging Face Transformers.js.
If useful, star the project or open a reproducible issue. Contributions that improve evidence quality and reduce false positives are welcome.
Available Tools
20 toolsaudit_eventsBRead-onlyIdempotent
Read paginated append-only audit events.
| Name | Required | Description | Default |
|---|---|---|---|
| after | No | ||
| project | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, openWorldHint=false and destructiveHint=false, so the safety profile is fully covered structurally. The description adds real context beyond that with 'append-only' (events cannot be mutated) and 'paginated' (results are chunked), but it says nothing about page size, whether there is a continuation token, or maximum result counts.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One front-loaded sentence with no filler words; the scoping traits arrive immediately. It is arguably too terse for a tool with an undocumented 0%-coverage schema, but as pure conciseness it is nearly optimal.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema, so the description would need to explain the return shape, yet it does not say what an audit event record contains or how pagination termination is signalled. Annotations cover the safety dimension adequately, making the definition minimally viable but leaving the agent under-informed about results.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description carries the full burden of explaining the two parameters, and it does not. The word 'paginated' hints that 'after' is a cursor, but the description never states that 'after' is an offset/cursor value or what 'project' must contain, so an agent must guess from the raw JSON Schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Read ... audit events') plus two qualifying traits, 'paginated' and 'append-only', so the agent knows it is a read-only, immutable log stream. It does not differentiate itself from siblings such as evidence_read or project_open by naming what makes audit events distinct, so it stops short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description offers no when-to-use guidance, no prerequisites, and no alternatives. An agent cannot infer from this sentence whether audit events should be consulted before or after security_triage or finding_correct. Usage is at best implied by the resource name alone.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
evidence_readBRead-onlyIdempotent
Read and integrity-check evidence in a project.
| Name | Required | Description | Default |
|---|---|---|---|
| project | Yes | ||
| evidence | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already establish read-only, idempotent, non-destructive, and closed-world behavior. The description adds that the tool performs an integrity check, which is useful context beyond annotations, but it does not explain what the check validates or what failures look like.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single front-loaded sentence with no wasted words. It is appropriately sized for a simple read tool, even if it omits details covered elsewhere.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a low-complexity read tool with rich annotations and no output schema, the description is nearly adequate but not complete. It leaves the integrity-check behavior undefined and provides no semantic detail for the two required parameters despite zero schema description coverage.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must carry the burden of explaining both parameters. It mentions 'project' and 'evidence' only as nouns and does not clarify that evidence is a 64-character hex hash or what the project identifier represents, leaving meanings to be inferred from the schema names and patterns.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (read) and resource (evidence), plus an additional integrity-check action, so an agent can understand the core operation. It does not explicitly name sibling tools like evidence_store or explain how it differs from them, which keeps it below a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No when-to-use guidance, prerequisites, or alternatives are provided. The agent must infer appropriate usage from the tool name and the single-sentence description alone.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
evidence_storeB
Redact structured secrets and store immutable content-addressed evidence; manually review free text before import.
| Name | Required | Description | Default |
|---|---|---|---|
| kind | Yes | ||
| source | Yes | ||
| project | Yes | ||
| artifact | Yes | ||
| captured_at | Yes | ||
| description | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only declare the basic profile (write, non-idempotent, non-destructive, closed-world). The description adds genuine behavioral content beyond that: automatic secret redaction, immutability of stored evidence, content-addressing, and a human-review requirement for free text. It still omits failure modes and permission requirements, and the non-idempotent annotation sits somewhat awkwardly against 'content-addressed'.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single dense sentence that front-loads the primary action (redact and store) before the caveat. No filler, though the semicolon clause compresses a precondition into a phrase that could be clearer.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With six required parameters, two enums, a nested artifact object, and no output schema, this is a complex write tool that needs more than one sentence. Missing are parameter semantics, what counts as 'structured secrets', how redaction failures are handled, and what the caller receives on success.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% across six required parameters, including a nested free-form artifact object, so the description carries the full burden. It gestures at 'structured secrets' (artifact) and 'free text' (description) but never maps meaning to project, kind, source, or captured_at, nor explains the enum vocabularies or the artifact shape.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names two specific actions on a concrete resource: redacting structured secrets and storing immutable, content-addressed evidence. It is clearly distinguishable from siblings like evidence_read (retrieval) and memory_store (memory), though it never names those alternatives explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It supplies one real precondition — manually review free text before import — which tells the agent how to prepare input. However, it gives no guidance on when to choose this tool over evidence_read, memory_store, or finding_correct, so the usage boundary is only implied.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
finding_correctB
Record a human correction without overwriting decisions; review before training.
| Name | Required | Description | Default |
|---|---|---|---|
| reason | Yes | ||
| finding | Yes | ||
| project | Yes | ||
| reviewer | Yes | ||
| correction | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only say readOnlyHint=false, idempotentHint=false, destructiveHint=false. The description adds two real behavioral facts beyond that: corrections do not overwrite existing decisions (append/history-preserving), and corrections are reviewed before being used for training. It still omits whether the write is verified, who may act as reviewer, and whether duplicates accumulate given idempotentHint=false.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single compact sentence with the primary action front-loaded and the two key caveats attached. Nothing is padded, though the brevity comes partly at the cost of the missing detail noted elsewhere.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, so the description should explain what a successful record produces, and with 5 undocumented required parameters it should define the inputs. It covers one behavioral nuance but leaves the agent unable to construct a call confidently without reverse-engineering the schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% across 5 required parameters, so the description carries the full burden. It only gestures at the 'correction' parameter and says nothing about project, finding, reason, or reviewer, nor about the five correction enum values (NOT-A-BUG, CONFIRMED, FIXED, REGRESSION, NEEDS-MANUAL-REVIEW) or the identifier patterns.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource: 'Record a human correction'. The object is clearly tied to a finding (name 'finding_correct'), and the phrase 'without overwriting decisions' scopes it as an additive annotation rather than a mutation of prior findings. It stops short of naming a sibling or explicitly describing the finding-correction workflow, so it is clear but not fully differentiated.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No when-to-use guidance and no alternatives are named, despite a large sibling set (security_triage, security_report, evidence_store) that could plausibly overlap. 'review before training' hints at a downstream workflow condition but does not tell the agent when to invoke this tool versus those siblings.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
memory_contextCRead-onlyIdempotent
Compile compact task context; critical state may exceed soft budget and is flagged.
| Name | Required | Description | Default |
|---|---|---|---|
| query | No | ||
| budget | No | ||
| project | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, and destructiveHint=false, so the safety profile is covered. The description adds one useful trait beyond that — that critical state may exceed the soft budget and be flagged — which hints at output behavior, but says nothing about how the flag manifests or how budget pressure is handled.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with no filler, and the budget caveat is appended rather than buried. It is efficient, though its brevity borders on under-specification rather than genuine conciseness.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With 0% schema coverage, no output schema, and three parameters, one short sentence is not enough to call this tool correctly. Nothing tells the agent what a compiled context contains, how 'project' and 'query' interact, or what the flagged over-budget result looks like.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must carry parameter meaning, and it largely does not. It only gestures at 'budget' (soft budget) while leaving the required 'project' and optional 'query' parameters — and their formats, maxLength, and pattern constraints — unexplained.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The verb 'Compile' and resource 'task context' give a rough sense of purpose, but 'task context' is vague and the description never distinguishes this from siblings like memory_search, memory_graph, or project_resume. An agent cannot confidently tell when this tool applies versus adjacent memory tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no when-to-use guidance, no preconditions, and no mention of alternatives among the many memory/project siblings. The agent must infer usage entirely from the name.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
memory_graphB
Read a bounded graph neighborhood or add a relationship.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare that the tool is not read-only, not idempotent, and not destructive. The description adds useful context that one mode performs a bounded graph read and the other adds a relationship, but it does not describe permissions, return shape, relation constraints, or mutation side effects beyond what annotations imply.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single front-loaded sentence covering both operating modes with no wasted words. It is appropriately concise, though its brevity contributes to gaps in usage and behavioral detail.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The rich oneOf schema and annotations cover most parameter and safety details, and the description identifies both operations. However, with no output schema, the description does not explain return values or how to choose between the two actions, leaving some complexity unaddressed.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is high: the schema defines action, node, depth, project, source, relation, and target with patterns, enums, defaults, and bounds. The description only adds the concepts of boundedness and relationship addition, without naming or explaining parameter semantics. This meets the baseline for high schema coverage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states two specific operations: reading a bounded graph neighborhood and adding a relationship. It distinguishes the tool from generic memory tools by naming graph-neighborhood and relationship concepts, but it does not explicitly contrast with siblings like memory_search, memory_context, or memory_store.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description says the tool can do one of two things, but gives no guidance on when to choose query versus relate, nor when to use this tool instead of sibling memory tools. The action selection is left entirely to the schema.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
memory_searchBRead-onlyIdempotent
Project-isolated FTS search over latest memory versions.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | ||
| query | Yes | ||
| project | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint and a closed-world, non-destructive profile, so safety is covered. The description adds genuinely non-obvious behavior: matching is keyword/full-text rather than semantic, results are scoped to a single project, and only the latest version of each memory is returned (superseded versions are excluded). It omits ranking/ordering and truncation behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with no filler; every phrase (project-isolated, FTS, latest versions) carries meaning. It is efficient, though the terseness contributes to the coverage gaps elsewhere.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with 0% schema description coverage and no output schema, the agent is left without parameter guidance, result ordering, or pagination expectations. Annotations and the lack of an output schema relieve some burden, but the description still leaves too much unspecified for correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description carries the full burden for all three parameters, yet it explains none of them. It never clarifies that 'query' is an FTS expression (and what syntax is valid), nor what 'project' format or 'limit' semantics (default 20, max 50) mean in practice.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('search') and resource ('memory'), plus scope qualifiers ('project-isolated', 'FTS', 'latest memory versions') that tell the agent what kind of matching to expect. It does not distinguish itself from siblings like memory_context or memory_graph, which also read memory, so the agent must infer which retrieval mode applies.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no when-to-use guidance and no named alternative, despite several memory-reading siblings (memory_context, memory_graph) in the toolset. The agent gets no signal about when keyword search is preferable to graph traversal or context assembly.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
memory_storeC
Store versioned typed memory with epistemic state and provenance.
| Name | Required | Description | Default |
|---|---|---|---|
| type | Yes | ||
| layer | No | L2 | |
| state | Yes | ||
| entity | Yes | ||
| content | No | ||
| project | Yes | ||
| summary | Yes | ||
| provenance | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare a non-destructive, non-idempotent write operation, and the description aligns by saying 'Store' and adding that memory is versioned and carries epistemic state/provenance. However, it does not explain versioning behavior (e.g., overwrite vs append), required permissions, or what happens on duplicate entity/project combinations. With annotations covering safety, this is moderate added context.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single sentence with no filler, front-loading the action and the distinguishing attributes. It is very concise and structured as a headline, though its brevity is arguably too extreme for an 8-parameter tool.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For an 8-parameter, 6-required write tool with a nested provenance object, 0% schema description coverage, and no output schema, the description is far too thin. It omits parameter meanings, usage context, versioning semantics, and error/duplicate behavior.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description carries full burden for 8 parameters. It only alludes to type, state, and provenance in passing and leaves project, entity, summary, content, and layer completely undocumented beyond their enum values.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb 'Store' and resource 'versioned typed memory', and names the epistemic-state and provenance dimensions. It does not explicitly differentiate from siblings such as evidence_store or memory_search, but the core action is unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance on when to use this tool versus memory_search, memory_context, evidence_store, or task_update. An agent must infer from the name alone whether to store a new memory or update/query existing ones.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
model_adviseBRead-onlyIdempotent
Optional local specialist; deterministic fallback when disabled. Cannot alter decisions.
| Name | Required | Description | Default |
|---|---|---|---|
| query | Yes | ||
| project | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already establish the safe read-only, idempotent, non-destructive profile. The description adds genuinely useful non-structured context: the output 'Cannot alter decisions' (advisory only, non-binding) and that a deterministic fallback exists when the tool is disabled, which tells the agent how to interpret results and handle unavailability.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two compact clauses with no filler, and the advisory/non-binding constraint is front-loaded. It is efficient, though the extreme brevity borders on under-specification rather than pure conciseness.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With two undocumented parameters, no output schema, and no explanation of how advice is produced or returned, the description is too thin for the tool's complexity. It omits what the query/project inputs mean and what the agent receives back, leaving real gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% for two required parameters (project, query), and the description says nothing about what either represents, their format, or the query length limit. The description does not compensate for the complete absence of parameter documentation in the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The name plus phrases like 'local specialist' only vaguely imply that this consults a local model for advisory input; the description never states a verb+resource such as 'get non-binding advice from a local model'. 'Optional local specialist' leaves the actual action and domain unclear, though an agent can roughly infer it is an advisor tool.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
'Optional' and 'deterministic fallback when disabled' imply the tool is supplementary and safe to skip, giving some usage context. However, it never states clearly when to use it versus not, nor names any alternative, so the guidance remains implied rather than explicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
project_openC
Create/open an isolated project.
| Name | Required | Description | Default |
|---|---|---|---|
| name | Yes | ||
| project | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly=false, idempotent=false, destructive=false, so the agent knows this mutates state. The description adds nothing beyond that: it does not explain what "isolated" means, what happens on a repeat call given idempotent=false, whether existing state is affected, or what error conditions exist.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The single sentence is front-loaded and free of waste, but it is terse to the point of under-specification rather than efficiently concise. Brevity here comes at the cost of the information the agent needs.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a non-idempotent mutation with two undocumented required parameters, no output schema, and a near-identical sibling (project_resume), far more context is needed. The agent cannot predict the effect of a second call or the meaning of either argument.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% for both parameters, and the description does not compensate. It never explains the distinction between `project` (pattern-constrained identifier) and `name` (free-form string up to 8000 chars), leaving the agent to infer which is the key and which is the label.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
"Create/open an isolated project" names the resource (project) but pairs two verbs whose difference matters, and never clarifies whether this creates a new project or attaches to an existing one. It also gives no signal to distinguish it from the sibling project_resume, which sounds like the same operation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no statement of when to use this tool versus project_resume, scope_register, or any other sibling, and no prerequisites or conditions for calling it. The agent must guess the intended context entirely.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
project_resumeBRead-onlyIdempotent
Resume latest task, scope, findings, open work and next actions after restart.
| Name | Required | Description | Default |
|---|---|---|---|
| project | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, so safety and repeatability are covered. The description adds the useful trigger context (post-restart) and enumerates what gets surfaced, but says nothing about return format or pagination and does not need to with no output schema.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One sentence, front-loaded with the verb and trigger, with no filler. The enumerated list is slightly loose but each item adds distinct content.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only, single-parameter tool with no output schema this is close to adequate, and annotations cover the safety profile. It still leaves the parameter undocumented and never clarifies what a 'restart' scope is, so it is minimum-viable rather than complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The single required 'project' parameter has 0% schema description coverage and the description never mentions it. The agent gets no hint about the identifier format, how it maps to the pattern, or whether it must match prior sessions, leaving the schema to speak entirely for itself.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (resume) plus the concrete state it restores: latest task, scope, findings, open work, next actions. That is far more specific than a tautology, but it never distinguishes itself from the adjacent project_open or scope_read siblings, leaving the agent to infer the boundary.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'after restart' implies the trigger condition, which is real usage guidance. However, it gives no when-not guidance and never names an alternative (e.g., project_open for starting fresh), so the agent must guess at the routing.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
scope_readCRead-onlyIdempotent
Read scope and expiry.
| Name | Required | Description | Default |
|---|---|---|---|
| project | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, destructiveHint=false, and openWorldHint=false, so the safety profile is fully covered by structured data. The description only adds that expiry information is surfaced, which is a modest hint but not meaningful behavioral disclosure.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
It is a single short fragment with no filler, which is efficient, but that brevity comes at the cost of under-specification rather than genuine conciseness.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema and no schema descriptions, the description carries the full burden yet omits what 'scope' contains, the return shape, and how this relates to the many sibling tools. An agent still cannot call it confidently.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and the single 'project' parameter is documented only by a regex pattern. The description does not explain what the project identifier is or how it should be supplied, so it fails to compensate for the schema gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a verb and a resource ('read scope and expiry'), so an agent knows it is a read operation on scope data plus an expiry value. However, 'scope' is never defined and the fragment gives no basis for distinguishing this from siblings like scope_register or evidence_read.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no statement of when to use this tool, when not to, or which sibling to prefer. The agent must infer usage context from the name alone.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
scope_registerB
Record operator-provided authorization; never derive permission from target responses. Exact origins, accounts, environments and expiry required.
| Name | Required | Description | Default |
|---|---|---|---|
| target | Yes | ||
| project | Yes | ||
| rate_limits | Yes | ||
| scope_expiry | Yes | ||
| allow_loopback | No | ||
| allowed_origins | Yes | ||
| allowed_accounts | Yes | ||
| assessment_owner | Yes | ||
| authorized_domains | Yes | ||
| prohibited_actions | Yes | ||
| authorization_reference | Yes | ||
| authorized_environments | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare this is a non-read-only, non-idempotent, non-destructive, non-open-world write. The description adds meaningful behavioral context beyond that — authorization must originate from the operator and never be inferred from target responses — but it does not disclose what the stored record controls downstream, whether re-registration replaces or accumulates, or what the caller gets back.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two front-loaded clauses: purpose first, then the sourcing constraint and required fields. No filler, though the second clause bundles an important principle with a terse field list and neither is expanded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For an 11-required-parameter, nested-object, stateful tool that likely governs the whole engagement, a two-clause description is thin. With no output schema and no per-parameter schema documentation, the definition should say what the registered scope constrains, what happens to prior registrations, and how it pairs with scope_read.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% across 12 parameters, so the description carries the full documentation burden. It names only four of them (origins, accounts, environments, expiry), leaving target, project, prohibited_actions, rate_limits, assessment_owner, authorization_reference, and allow_loopback completely undocumented in both schema and prose.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
"Record operator-provided authorization" states a specific verb and resource, and the definition makes clear this is a write of an authorization record rather than a lookup. It does not name its natural counterpart (scope_read) or otherwise differentiate itself from siblings explicitly, so it stops short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The clause "never derive permission from target responses" is a genuine usage directive about where the authorization data must come from, and the tool must be called before/independent of any assessment activity. However, it never names an alternative (e.g., scope_read for retrieving the recorded scope) or states when this must be invoked versus those siblings, so usage is only implied.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
security_controlled_getC
One bounded GET, exact authorized origin and controlled account, pinned DNS, no redirects. No state-changing requests.
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | ||
| account | Yes | ||
| project | Yes | ||
| environment | Yes | ||
| authorization | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description asserts 'No state-changing requests', i.e. a read-only operation, while the annotations declare readOnlyHint=false. That is a direct inconsistency. Beyond the contradiction, the description does add genuinely useful behavior (pinned DNS, no redirects, bounded single request), but the read-only claim conflicts with the structured hint.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two terse sentences, front-loaded with the operation and then the safety constraints; no filler. It is arguably over-terse, but every clause earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 5-parameter, security-sensitive network tool with no output schema and 0% schema description coverage, the description should explain the remaining parameters and the account/auth model. It covers behavioral guardrails reasonably but leaves project, environment, and authorization unexplained and contradicts the read-only annotation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% across 5 parameters, so the description carries the full burden. It only loosely gestures at two of them ('authorized origin' -> url, 'controlled account' -> account) and says nothing about project, environment, or authorization, nor about required URI/pattern formats.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a concrete verb+resource: a single bounded HTTP GET to an authorized origin under a controlled account. An agent can tell this is a network-fetch tool rather than one of the read_/store_/task_ siblings, though it never names a sibling or explicitly scopes itself against them.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Constraints imply usage ('exact authorized origin', 'controlled account', 'no state-changing requests') but there is no explicit when-to-use/when-not or alternative tool named. The agent must infer that this is for safe, tightly scoped fetches rather than general HTTP access.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
security_reportC
Generate gated technical/executive reports and a minimal request template.
| Name | Required | Description | Default |
|---|---|---|---|
| finding | Yes | ||
| project | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare a non-read-only, non-idempotent, non-destructive, non-open-world operation, yet the description adds nothing about what "gated" means, what access or auth is required, whether reports are persisted, or whether repeated calls create duplicates. For a mutation-flagged tool with zero annotation-free context, this is a real gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single short sentence with no padding, but its brevity comes from omission rather than precision — the only content is a vague noun phrase, so nothing is effectively front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, two required undocumented params, and mutation-style annotations, the description should carry substantial load but delivers one clause. Missing: what the output looks like, what "gated" requires, and what identifiers project/finding accept.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and both required parameters (project, finding) are constrained only by an opaque regex. The description never mentions either parameter, so it fails to compensate for the schema gap, leaving the agent to guess the identifier format and source.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a verb+resource pair ("Generate ... reports") and mentions a request template, but "gated technical/executive reports" is opaque jargon and the sentence never says what a report covers or how it differs from sibling security tools (security_triage, security_severity, finding_correct). An agent can guess it produces a report artifact but cannot tell which one it needs.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No when-to-use guidance, no prerequisites, and no mention of alternatives among the many security_* siblings. "Gated" hints at some precondition but never states it, leaving the agent to infer when this tool is appropriate.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
security_severityBRead-onlyIdempotent
CVSS 3.1 base score with metric explanation. Does not infer metric values.
| Name | Required | Description | Default |
|---|---|---|---|
| vector | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover read-only, idempotent, non-destructive, and closed-world behavior. The description adds a meaningful constraint beyond annotations: it will not infer missing metric values, so the caller must supply a complete vector. It still omits output format and error behavior, so it is not a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two short sentences with no filler. It is front-loaded with the result and follows with the key limitation, making it appropriately sized for a simple single-parameter tool.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple CVSS scoring tool with no output schema, the description covers what is returned and one important constraint. However, it leaves the required vector format undocumented and provides no usage context, which are clear gaps against the 0% parameter coverage.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
There is one required parameter, 'vector', with 0% schema description coverage. The description implies CVSS 3.1 context but does not explain the expected vector syntax, required metric fields, or examples, so it fails to compensate for the schema gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states the specific domain output: a CVSS 3.1 base score with metric explanation. It does not use a clear action verb, but an agent can infer the tool computes or returns severity scoring. Sibling tools are not distinguished, keeping it from a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives no when-to-use guidance and does not mention alternatives such as security_triage or security_report. The sentence 'Does not infer metric values' is a behavioral boundary, not usage guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
security_skillsBRead-onlyIdempotent
List generic evidence requirements and false-positive conditions.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, destructiveHint=false and openWorldHint=false, so the safety profile is fully covered. The description adds no further behavioral context (e.g., that it returns a static reference set with no side effects), but it also does not contradict the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with no filler. It is efficient, though slightly terse given the ambiguity of what a 'skill' returns.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter, read-only list tool with no output schema and full annotation coverage, the minimum is met. Still, it never clarifies the shape or scope of what is returned, leaving a gap between the name 'security_skills' and the described content.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so there is nothing for the description to disambiguate; the baseline of 4 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The verb 'List' and the resource 'generic evidence requirements and false-positive conditions' are stated, so the agent knows the content type. However, the name 'security_skills' is never reconciled with the description, and nothing distinguishes it from siblings like security_triage or security_report.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No indication of when this tool should be invoked versus security_triage, security_severity, or security_report. The word 'generic' hints it is non-case-specific, but no explicit trigger or alternative is named.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
security_triageD
Evaluate supplied evidence-backed claims; preserve uncertainty and version decisions.
| Name | Required | Description | Default |
|---|---|---|---|
| cvss | No | ||
| skill | Yes | ||
| title | Yes | ||
| actual | Yes | ||
| claims | Yes | ||
| impact | Yes | ||
| finding | Yes | ||
| project | Yes | ||
| accounts | Yes | ||
| business | No | ||
| endpoint | Yes | ||
| expected | Yes | ||
| regression | Yes | ||
| environment | Yes | ||
| remediation | Yes | ||
| reproduction | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare readOnlyHint=false and idempotentHint=false, so this is a non-idempotent mutation of some kind, yet the description never says what is written or versioned. 'Preserve uncertainty and version decisions' weakly hints at versioning behavior, but with 14 required inputs describing a recorded finding, the omission of the actual side effect (a persisted triage record?) is a real gap. No contradiction with annotations, but minimal added context.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The single sentence is compact and front-loaded, so there is no bloat, but neither clause earns its place: both are too abstract to help an agent act. Brevity here reflects under-specification rather than disciplined conciseness for a tool with 14 required fields.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 16-parameter, 14-required mutation tool with nested objects, an enum, and no output schema, the description is essentially empty. It documents no inputs, no effects, and no return expectations; an agent has nothing to work from beyond the raw schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% across 16 parameters (14 required), including two nested objects (claims, business) and an undocumented skill enum. The description compensates for none of this: it names no parameter, no expected input shape, and no format. With the schema unable to carry meaning, the description's silence leaves the entire parameter contract unexplained.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The phrase 'Evaluate supplied evidence-backed claims' is an abstract verb+resource that never states what the tool concretely produces or persists, and 'preserve uncertainty and version decisions' is opaque meta-language. Nothing distinguishes it from siblings like security_severity, security_report, or finding_correct. The description reads as near-tautological restatement for a security-triage tool rather than a clear statement of function.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no when-to-use guidance, no preconditions, and no named alternatives despite a dense sibling set (security_severity, security_report, finding_correct). 'Preserve uncertainty and version decisions' gestures at a behavioral principle but gives no selection criteria. An agent cannot infer when this tool should be chosen over its siblings.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
task_check_duplicateCRead-onlyIdempotent
Check prior completed exact test; changed deployment must use new state_version.
| Name | Required | Description | Default |
|---|---|---|---|
| project | Yes | ||
| signature | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, so the safety profile is covered by structured data. The description adds one meaningful behavioral nuance — that state_version must change when the deployment changes — but omits what the check returns, how a duplicate is interpreted, or any side effects. With annotations carrying the safety burden, a 3 is appropriate.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single compact sentence with the check purpose front-loaded and zero wasted words, which is structurally efficient. However, the brevity tips into under-specification for a tool with a nested six-field signature, so conciseness comes at the cost of completeness.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with a nested required signature object, no output schema, and no field-level documentation, the description is far too thin. It does not explain what a duplicate result looks like, how the signature fields combine to define 'exactness,' or how the caller should act on the outcome, leaving significant gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and the nested signature object has six required fields with no per-field documentation. The description only adds meaning for one field (state_version, via the changed-deployment rule), leaving project, method, endpoint, auth_context, ownership, and test_type entirely undocumented. It does not compensate for the near-total coverage gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a check action against 'prior completed exact test,' implying a duplicate-detection function, but 'exact test' and 'prior completed' are opaque jargon that leave the exact purpose ambiguous. It does not clearly tell an agent that this returns a duplicate verdict, and no sibling is named for contrast, though the sibling list contains no near-equivalent to differentiate against.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The only guidance is the embedded constraint that 'changed deployment must use new state_version,' which is about a parameter value rather than when to invoke this tool. There is no statement of when to call task_check_duplicate versus alternatives, no prerequisites, and no exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
task_updateC
Version task state and record test signature including deployment state.
| Name | Required | Description | Default |
|---|---|---|---|
| task | Yes | ||
| status | Yes | ||
| project | Yes | ||
| summary | Yes | ||
| evidence | No | ||
| signature | No | ||
| next_actions | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=false, destructiveHint=false, idempotentHint=false, and openWorldHint=false, so the safety profile is covered. The description adds a vague notion of 'versioning' and 'deployment state' but never explains what versioning means, whether prior state is retained or overwritten, or what auth is required — leaving the mutation's actual behavior opaque.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
It is a single short sentence with no padding, but its brevity comes at the cost of clarity rather than through efficient front-loading. Reasonably sized, poorly chosen words.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
This is a mutation tool with 7 parameters, 5 required, a nested signature object, and 0% schema coverage, and it has no output schema to fall back on. The description does not explain the nested object, the evidence hashes, or what a successful update returns, so it is significantly under-specified for the tool's complexity.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must carry parameter meaning, yet it only loosely gestures at 'task state' (status) and 'test signature' (signature). The 7 parameters — including project, summary, next_actions, the defaulted evidence hash array, and the nested signature object — are essentially undocumented.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names an action on a task ('version task state') and a secondary behavior ('record test signature'), so an agent can broadly infer it mutates a task record. However, 'Version' is ambiguous jargon — it doesn't state the resource or verb cleanly (update? transition? snapshot?) — and it draws no line against siblings like task_check_duplicate.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to call this tool versus alternatives such as task_check_duplicate or project_resume. No prerequisites, no exclusions, no context on when a task should be updated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
20 tool updates
v0.1.0- First observed
audit_events - First observed
evidence_read - First observed
evidence_store - First observed
finding_correct - First observed
memory_context - First observed
memory_graph - First observed
memory_search - First observed
memory_store - First observed
model_advise - First observed
project_open - First observed
project_resume - First observed
scope_read - First observed
scope_register - First observed
security_controlled_get - First observed
security_report - First observed
security_severity - First observed
security_skills - First observed
security_triage - First observed
task_check_duplicate - First observed
task_update
TDQS
Scored across 20 tools
Each tool targets a distinct resource and action, with clear boundaries between evidence, memory, tasks, scope, and security operations. The security_* cluster is nuanced but differentiated by purpose: skills list requirements, triage evaluates claims, severity computes CVSS, report generates deliverables, and controlled_get performs a bounded request.
All names use lowercase snake_case with a consistent domain prefix, making the set easy to scan. Most follow a noun_verb pattern, but a few are noun_noun (e.g., security_skills, security_severity, memory_context, audit_events), which is a minor deviation.
20 tools is slightly above the typical 3-15 sweet spot, but the domain spans projects, evidence, tasks, security triage, scope control, memory, and audit. Each tool appears to earn its place, so the count is reasonable rather than bloated.
The surface covers core project, evidence, scope, memory, security, finding, and audit workflows. Minor gaps exist, such as no explicit task_create/list or project_list/close, but immutable and versioned designs mitigate them and agents can work around via existing tools.
Maintenance
Related MCP Connectors
Private-by-default, local-first memory/context/task orchestrator for MCP apps and agents.
SigB (Signal Bureau): maintained, source-traced record of entities and events in your named domain.
Persistent knowledge graph for AI-augmented teams. Store decisions, findings, and standing rules across agent sessions with semantic search and typed connections. Includes cross-session memory, audit trail, workspace isolation, and secret detection. Built for teams running agents that need to remember. Free until launch with team tier as default, anon trial available.
- memoricOAuthio.memoric
Provenance-first database for teams and agents: every value carries sources, rules and coverage.
Related MCP Servers
- AlicenseAqualityCmaintenanceMCP server that indexes Markdown, Word, HTML, and PDF documents into a SQLite knowledge graph with CJK+Latin full-text search and cross-document reference tracking. Runs drift audits to surface stale policies, conflicting research claims, superseded ADRs, and undocumented code exports.108MIT
- AlicenseBqualityAmaintenanceLocal-first, auditable memory for Codex, Claude Code, and MCP clients. It stores scoped user/project memory in SQLite or Postgres, serves read-only recall and inspection tools by default, and supports opt-in governed writeback with review and forget controls.8145 npm18MIT
- AlicenseNot gradedqualityBmaintenanceLocal static-analysis assistant for Android malware research that manages investigation cases, exposes MCP tools via a local server, and persists evidence-backed findings without cloud dependency.MIT
- AlicenseAqualityBmaintenanceA local-first MCP server for retrieving a small evidence set and recording reviewed conclusions, policy-gated and redacted without giving an agent general filesystem access.5MIT