Skip to main content
Glama

Sentinel Memory MCP

An open-source cybersecurity MCP server for persistent graph memory, evidence-backed vulnerability triage, and authorized API security research. Connect it to Google Antigravity, OpenAI Codex, or another stdio MCP client to carry structured investigation state across sessions.

Remember what was tested. Preserve the evidence. Reject weak conclusions. Resume with the next useful action.

Quick start · Antigravity · Codex · Self-hosting · हिंदी · AI-readable overview

Release reality: the core runs locally and is tested with official MCP clients and a synthetic security lab. Actual IDE UI verification is separate. This is an early development release, not a certified production security product. Local models are optional and disabled by default after unsuccessful specialist-output benchmarks. Detailed status · Known limitations

Why this exists

Security investigations outlive a chat window. Repeated logs bury useful facts, a fixed issue can resurface as active, and a changed object ID can be mistaken for a vulnerability. Sentinel stores typed observations, hypotheses, tasks, evidence, relationships, and versioned decisions rather than treating an entire conversation as trusted memory.

When you need to…

Sentinel provides…

Continue an API assessment after context loss

A compact resume packet with scope, tasks, findings, corrections and next actions

Avoid repeating completed work

Test signatures including account context, ownership and deployment state

Check a suspected IDOR/BOLA

Evidence requirements for ownership, expected access, protected data and false positives

Keep sensitive evidence useful

Redaction, SHA-256 content addressing, provenance and integrity checks

Explain a finding to engineers and managers

Technical and executive reports with explicit validation gates

Use a small local model cautiously

Offline, process-isolated advice with schema checks and deterministic fallback

Related MCP server: GoodMemory

Quick start

Prerequisites: Git and Node.js 24.14+. No cloud API key, Neo4j, Docker or model download is required for core features.

git clone https://github.com/vikrant-project/sentinel-memory-mcp.git
cd sentinel-memory-mcp
npm ci
npm test
npm run configure

The configuration generator uses your checkout's absolute paths and preserves other entries in existing MCP JSON files. Generated machine-specific files are ignored by Git.

Host

Next step

Guide

Google Antigravity

Open this checkout, refresh MCP servers, look for sentinel-memory

Setup

OpenAI Codex

Merge examples/codex.local.toml, or use codex mcp add

CLI and TOML

Other stdio MCP client

Launch node /absolute/path/dist/src/mcp/server.js

Generic config

Docker or private VPS

Keep stdio; persist the database

Self-hosting

Try it without touching a real target:

npm run demo

The demo starts a loopback-only lab, connects an official MCP client, captures controlled-account evidence, distinguishes a private authorization failure from a public object, and saves synthetic results to outputs/demo-summary.json. It shuts down its processes afterward.

Use Sentinel to open project api-review. Read security_skills and project_resume.
Do not send any requests yet. Help me record the assessment owner's authorized
scope, controlled accounts, expected access rules, and next validation steps.

At the end of a session:

Save the current task, evidence references, unresolved hypotheses, and next
 actions in Sentinel. Next session, resume project api-review from that state.

The project does not independently establish permission to test any target. Scope must come from the assessment owner.

Architecture

MCP gateway connects graph memory, validation, evidence, reports and optional isolated model

flowchart LR
    Host[Antigravity / Codex / MCP client] --> MCP[Official MCP SDK · stdio]
    MCP --> Context[Context compiler]
    Context <--> Graph[(SQLite graph + FTS5)]
    MCP --> Scope[Scope and request controls]
    Scope --> Evidence[Redacted evidence + SHA-256]
    Evidence --> Triage[Evidence gates + false-positive checks]
    Triage --> Reports[Technical / executive reports]
    Graph --> Model[Optional local HF process]
    Model -. advisory output only .-> Context

The deterministic path handles persistence, retrieval and validation. The local model does not authorize requests, run commands, rewrite weights or confirm findings.

Evidence-first investigation workflow

Six steps: scope, observe, preserve, validate, report, resume

  1. Scope: authorized origins, accounts, environments, restrictions, limits and expiry.

  2. Observe: expected behavior and minimal controlled test data.

  3. Preserve: redacted evidence, provenance and immutable content hashes.

  4. Validate: reproducibility, security boundary, impact, controls and false positives.

  5. Report: FINAL only when implemented gates pass; otherwise retain DRAFT.

  6. Resume: save task state and check prior signatures before repeating work.

Workflow

What it checks

Examples of rejected lookalikes

Authorization / IDOR / BOLA

Ownership, unrelated account, expected access, protected response

Public objects, sharing, expected admin access, cached responses

Rate limits

Sensitive operation, abuse feasibility, impact and controls

“Ten requests worked,” effective throttling, no security impact

Business logic

Backend entitlement and documented workflow rules

Free features, trials, promotions, frontend cosmetics

Data exposure

Sensitive fields, permitted caller and protected contents

Public data, sharing, expected privileges, non-sensitive fixtures

These are structured analyst workflows, not universal vulnerability detectors. CONFIRMED means implemented evidence gates passed; imported evidence and semantic claims still need review.

MCP interface

20 tools · 4 resource templates · 6 prompts

Area

Tools

Project and scope

project_open, project_resume, scope_register, scope_read

Graph and context

memory_store, memory_search, memory_context, memory_graph

Evidence and tasks

evidence_store, evidence_read, task_update, task_check_duplicate

Research and reporting

security_skills, security_triage, security_severity, security_report, security_controlled_get

Review and advice

finding_correct, model_advise, audit_events

Resources expose context, scope, evidence and report snapshots. Prompts support investigation, validation, two-user comparison, reporting, false-positive review and regression checks. Payload examples →

Measured results, with context

Measured p95 latency for memory search, graph retrieval and context compilation

Recorded on Windows, Node 24.14.1, Ryzen 7 7435HS and about 16 GB RAM. The benchmark uses 500 graph nodes and 100 samples per operation. The chart is generated from the committed benchmark JSON.

Check

Recorded result

Interpretation

Automated suite

28 passing tests in the recorded run

SDK subprocesses, restart, isolation, redaction, scope, lab, backup and model failure

Synthetic triage

32/32 expected classifications

Engineering fixtures, not independent real-world accuracy

Confirmed precision / recall

100% / 20% on those fixtures

Incomplete positive cases deliberately remain unconfirmed

Repeated-history reduction

99.32% estimated

Synthetic repeated logs; characters/4, not a tokenizer or lossless guarantee

Local model schema acceptance

0/4 for each candidate

Both disabled by default; failures and timeouts retained

Do not extrapolate these numbers to unseen targets. Methodology · Raw results · Model results · 11 hardening iterations

Example deliverables

All committed example evidence is synthetic. Real databases, credentials, model weights, local caches and temporary files are excluded from version control.

Self-hosting and data ownership

Run beside your MCP client, through a Docker stdio process, or over authenticated SSH to your VPS. There is no public HTTP endpoint in this release. A website host or reverse proxy alone does not turn it into a remote MCP service.

The hosting guide covers Docker, volumes, SSH, restarts and backups. Docker/VPS recipes are marked unverified because Docker was unavailable on the original development machine.

Documentation map

FAQ

Is this a cybersecurity MCP server or a scanner?
It is an MCP server for security investigation memory, evidence handling and triage. Its networking tool performs one bounded authorized GET. It does not mass-scan, brute-force credentials or autonomously exploit targets.

Can I use it without a local LLM?
Yes. Default features use deterministic logic, SQLite, FTS5 and graph retrieval. Model downloads are opt-in.

Does it replace a security engineer?
No. Scope, policy, interpretation and business impact need an authorized analyst. Confidence probability is intentionally uncalibrated.

Does it work with Codex and Antigravity?
Both support local MCP configuration. Guides and generated examples are included. Automated tests verify official MCP clients; a specific IDE version's UI and restart behavior need a separate check.

Can it rank first in AI search?
No repository can guarantee that. Clear descriptions, useful examples, descriptive topics and accessible documentation help discovery. This project does not use fake ratings, hidden instructions or keyword spam. Discoverability notes

License and credits

Project code, original diagrams and documentation: MIT, © 2026 vikrant-project. Dependencies and model checkpoints retain their own licenses. Built with the official MCP TypeScript SDK, Node.js, SQLite, Zod and optional Hugging Face Transformers.js.

If useful, star the project or open a reproducible issue. Contributions that improve evidence quality and reduce false positives are welcome.

Available Tools

20 tools
audit_eventsB
Read-onlyIdempotent

Read paginated append-only audit events.

ParametersJSON Schema
NameRequiredDescriptionDefault
afterNo
projectYes

TDQS

B3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, openWorldHint=false and destructiveHint=false, so the safety profile is fully covered structurally. The description adds real context beyond that with 'append-only' (events cannot be mutated) and 'paginated' (results are chunked), but it says nothing about page size, whether there is a continuation token, or maximum result counts.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

One front-loaded sentence with no filler words; the scoping traits arrive immediately. It is arguably too terse for a tool with an undocumented 0%-coverage schema, but as pure conciseness it is nearly optimal.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

There is no output schema, so the description would need to explain the return shape, yet it does not say what an audit event record contains or how pagination termination is signalled. Annotations cover the safety dimension adequately, making the definition minimally viable but leaving the agent under-informed about results.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description carries the full burden of explaining the two parameters, and it does not. The word 'paginated' hints that 'after' is a cursor, but the description never states that 'after' is an offset/cursor value or what 'project' must contain, so an agent must guess from the raw JSON Schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Read ... audit events') plus two qualifying traits, 'paginated' and 'append-only', so the agent knows it is a read-only, immutable log stream. It does not differentiate itself from siblings such as evidence_read or project_open by naming what makes audit events distinct, so it stops short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description offers no when-to-use guidance, no prerequisites, and no alternatives. An agent cannot infer from this sentence whether audit events should be consulted before or after security_triage or finding_correct. Usage is at best implied by the resource name alone.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

evidence_readB
Read-onlyIdempotent

Read and integrity-check evidence in a project.

ParametersJSON Schema
NameRequiredDescriptionDefault
projectYes
evidenceYes

TDQS

B3.1/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already establish read-only, idempotent, non-destructive, and closed-world behavior. The description adds that the tool performs an integrity check, which is useful context beyond annotations, but it does not explain what the check validates or what failures look like.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single front-loaded sentence with no wasted words. It is appropriately sized for a simple read tool, even if it omits details covered elsewhere.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a low-complexity read tool with rich annotations and no output schema, the description is nearly adequate but not complete. It leaves the integrity-check behavior undefined and provides no semantic detail for the two required parameters despite zero schema description coverage.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must carry the burden of explaining both parameters. It mentions 'project' and 'evidence' only as nouns and does not clarify that evidence is a 64-character hex hash or what the project identifier represents, leaving meanings to be inferred from the schema names and patterns.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (read) and resource (evidence), plus an additional integrity-check action, so an agent can understand the core operation. It does not explicitly name sibling tools like evidence_store or explain how it differs from them, which keeps it below a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No when-to-use guidance, prerequisites, or alternatives are provided. The agent must infer appropriate usage from the tool name and the single-sentence description alone.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

evidence_storeB

Redact structured secrets and store immutable content-addressed evidence; manually review free text before import.

ParametersJSON Schema
NameRequiredDescriptionDefault
kindYes
sourceYes
projectYes
artifactYes
captured_atYes
descriptionYes

TDQS

B3.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only declare the basic profile (write, non-idempotent, non-destructive, closed-world). The description adds genuine behavioral content beyond that: automatic secret redaction, immutability of stored evidence, content-addressing, and a human-review requirement for free text. It still omits failure modes and permission requirements, and the non-idempotent annotation sits somewhat awkwardly against 'content-addressed'.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single dense sentence that front-loads the primary action (redact and store) before the caveat. No filler, though the semicolon clause compresses a precondition into a phrase that could be clearer.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With six required parameters, two enums, a nested artifact object, and no output schema, this is a complex write tool that needs more than one sentence. Missing are parameter semantics, what counts as 'structured secrets', how redaction failures are handled, and what the caller receives on success.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% across six required parameters, including a nested free-form artifact object, so the description carries the full burden. It gestures at 'structured secrets' (artifact) and 'free text' (description) but never maps meaning to project, kind, source, or captured_at, nor explains the enum vocabularies or the artifact shape.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names two specific actions on a concrete resource: redacting structured secrets and storing immutable, content-addressed evidence. It is clearly distinguishable from siblings like evidence_read (retrieval) and memory_store (memory), though it never names those alternatives explicitly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It supplies one real precondition — manually review free text before import — which tells the agent how to prepare input. However, it gives no guidance on when to choose this tool over evidence_read, memory_store, or finding_correct, so the usage boundary is only implied.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

finding_correctB

Record a human correction without overwriting decisions; review before training.

ParametersJSON Schema
NameRequiredDescriptionDefault
reasonYes
findingYes
projectYes
reviewerYes
correctionYes

TDQS

B3.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only say readOnlyHint=false, idempotentHint=false, destructiveHint=false. The description adds two real behavioral facts beyond that: corrections do not overwrite existing decisions (append/history-preserving), and corrections are reviewed before being used for training. It still omits whether the write is verified, who may act as reviewer, and whether duplicates accumulate given idempotentHint=false.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single compact sentence with the primary action front-loaded and the two key caveats attached. Nothing is padded, though the brevity comes partly at the cost of the missing detail noted elsewhere.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema exists, so the description should explain what a successful record produces, and with 5 undocumented required parameters it should define the inputs. It covers one behavioral nuance but leaves the agent unable to construct a call confidently without reverse-engineering the schema.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% across 5 required parameters, so the description carries the full burden. It only gestures at the 'correction' parameter and says nothing about project, finding, reason, or reviewer, nor about the five correction enum values (NOT-A-BUG, CONFIRMED, FIXED, REGRESSION, NEEDS-MANUAL-REVIEW) or the identifier patterns.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource: 'Record a human correction'. The object is clearly tied to a finding (name 'finding_correct'), and the phrase 'without overwriting decisions' scopes it as an additive annotation rather than a mutation of prior findings. It stops short of naming a sibling or explicitly describing the finding-correction workflow, so it is clear but not fully differentiated.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No when-to-use guidance and no alternatives are named, despite a large sibling set (security_triage, security_report, evidence_store) that could plausibly overlap. 'review before training' hints at a downstream workflow condition but does not tell the agent when to invoke this tool versus those siblings.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

memory_contextC
Read-onlyIdempotent

Compile compact task context; critical state may exceed soft budget and is flagged.

ParametersJSON Schema
NameRequiredDescriptionDefault
queryNo
budgetNo
projectYes

TDQS

C2.7/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, and destructiveHint=false, so the safety profile is covered. The description adds one useful trait beyond that — that critical state may exceed the soft budget and be flagged — which hints at output behavior, but says nothing about how the flag manifests or how budget pressure is handled.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with no filler, and the budget caveat is appended rather than buried. It is efficient, though its brevity borders on under-specification rather than genuine conciseness.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With 0% schema coverage, no output schema, and three parameters, one short sentence is not enough to call this tool correctly. Nothing tells the agent what a compiled context contains, how 'project' and 'query' interact, or what the flagged over-budget result looks like.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must carry parameter meaning, and it largely does not. It only gestures at 'budget' (soft budget) while leaving the required 'project' and optional 'query' parameters — and their formats, maxLength, and pattern constraints — unexplained.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The verb 'Compile' and resource 'task context' give a rough sense of purpose, but 'task context' is vague and the description never distinguishes this from siblings like memory_search, memory_graph, or project_resume. An agent cannot confidently tell when this tool applies versus adjacent memory tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance, no preconditions, and no mention of alternatives among the many memory/project siblings. The agent must infer usage entirely from the name.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

memory_graphB

Read a bounded graph neighborhood or add a relationship.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

B3.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare that the tool is not read-only, not idempotent, and not destructive. The description adds useful context that one mode performs a bounded graph read and the other adds a relationship, but it does not describe permissions, return shape, relation constraints, or mutation side effects beyond what annotations imply.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single front-loaded sentence covering both operating modes with no wasted words. It is appropriately concise, though its brevity contributes to gaps in usage and behavioral detail.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The rich oneOf schema and annotations cover most parameter and safety details, and the description identifies both operations. However, with no output schema, the description does not explain return values or how to choose between the two actions, leaving some complexity unaddressed.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is high: the schema defines action, node, depth, project, source, relation, and target with patterns, enums, defaults, and bounds. The description only adds the concepts of boundedness and relationship addition, without naming or explaining parameter semantics. This meets the baseline for high schema coverage.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states two specific operations: reading a bounded graph neighborhood and adding a relationship. It distinguishes the tool from generic memory tools by naming graph-neighborhood and relationship concepts, but it does not explicitly contrast with siblings like memory_search, memory_context, or memory_store.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description says the tool can do one of two things, but gives no guidance on when to choose query versus relate, nor when to use this tool instead of sibling memory tools. The action selection is left entirely to the schema.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

memory_storeC

Store versioned typed memory with epistemic state and provenance.

ParametersJSON Schema
NameRequiredDescriptionDefault
typeYes
layerNoL2
stateYes
entityYes
contentNo
projectYes
summaryYes
provenanceYes

TDQS

C2.9/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare a non-destructive, non-idempotent write operation, and the description aligns by saying 'Store' and adding that memory is versioned and carries epistemic state/provenance. However, it does not explain versioning behavior (e.g., overwrite vs append), required permissions, or what happens on duplicate entity/project combinations. With annotations covering safety, this is moderate added context.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single sentence with no filler, front-loading the action and the distinguishing attributes. It is very concise and structured as a headline, though its brevity is arguably too extreme for an 8-parameter tool.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For an 8-parameter, 6-required write tool with a nested provenance object, 0% schema description coverage, and no output schema, the description is far too thin. It omits parameter meanings, usage context, versioning semantics, and error/duplicate behavior.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, so the description carries full burden for 8 parameters. It only alludes to type, state, and provenance in passing and leaves project, entity, summary, content, and layer completely undocumented beyond their enum values.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb 'Store' and resource 'versioned typed memory', and names the epistemic-state and provenance dimensions. It does not explicitly differentiate from siblings such as evidence_store or memory_search, but the core action is unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to use this tool versus memory_search, memory_context, evidence_store, or task_update. An agent must infer from the name alone whether to store a new memory or update/query existing ones.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

model_adviseB
Read-onlyIdempotent

Optional local specialist; deterministic fallback when disabled. Cannot alter decisions.

ParametersJSON Schema
NameRequiredDescriptionDefault
queryYes
projectYes

TDQS

B3.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already establish the safe read-only, idempotent, non-destructive profile. The description adds genuinely useful non-structured context: the output 'Cannot alter decisions' (advisory only, non-binding) and that a deterministic fallback exists when the tool is disabled, which tells the agent how to interpret results and handle unavailability.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two compact clauses with no filler, and the advisory/non-binding constraint is front-loaded. It is efficient, though the extreme brevity borders on under-specification rather than pure conciseness.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With two undocumented parameters, no output schema, and no explanation of how advice is produced or returned, the description is too thin for the tool's complexity. It omits what the query/project inputs mean and what the agent receives back, leaving real gaps.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% for two required parameters (project, query), and the description says nothing about what either represents, their format, or the query length limit. The description does not compensate for the complete absence of parameter documentation in the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The name plus phrases like 'local specialist' only vaguely imply that this consults a local model for advisory input; the description never states a verb+resource such as 'get non-binding advice from a local model'. 'Optional local specialist' leaves the actual action and domain unclear, though an agent can roughly infer it is an advisor tool.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

'Optional' and 'deterministic fallback when disabled' imply the tool is supplementary and safe to skip, giving some usage context. However, it never states clearly when to use it versus not, nor names any alternative, so the guidance remains implied rather than explicit.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

project_openC

Create/open an isolated project.

ParametersJSON Schema
NameRequiredDescriptionDefault
nameYes
projectYes

TDQS

C2.2/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnly=false, idempotent=false, destructive=false, so the agent knows this mutates state. The description adds nothing beyond that: it does not explain what "isolated" means, what happens on a repeat call given idempotent=false, whether existing state is affected, or what error conditions exist.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The single sentence is front-loaded and free of waste, but it is terse to the point of under-specification rather than efficiently concise. Brevity here comes at the cost of the information the agent needs.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a non-idempotent mutation with two undocumented required parameters, no output schema, and a near-identical sibling (project_resume), far more context is needed. The agent cannot predict the effect of a second call or the meaning of either argument.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% for both parameters, and the description does not compensate. It never explains the distinction between `project` (pattern-constrained identifier) and `name` (free-form string up to 8000 chars), leaving the agent to infer which is the key and which is the label.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

"Create/open an isolated project" names the resource (project) but pairs two verbs whose difference matters, and never clarifies whether this creates a new project or attaches to an existing one. It also gives no signal to distinguish it from the sibling project_resume, which sounds like the same operation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no statement of when to use this tool versus project_resume, scope_register, or any other sibling, and no prerequisites or conditions for calling it. The agent must guess the intended context entirely.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

project_resumeB
Read-onlyIdempotent

Resume latest task, scope, findings, open work and next actions after restart.

ParametersJSON Schema
NameRequiredDescriptionDefault
projectYes

TDQS

B3.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, so safety and repeatability are covered. The description adds the useful trigger context (post-restart) and enumerates what gets surfaced, but says nothing about return format or pagination and does not need to with no output schema.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

One sentence, front-loaded with the verb and trigger, with no filler. The enumerated list is slightly loose but each item adds distinct content.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a read-only, single-parameter tool with no output schema this is close to adequate, and annotations cover the safety profile. It still leaves the parameter undocumented and never clarifies what a 'restart' scope is, so it is minimum-viable rather than complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The single required 'project' parameter has 0% schema description coverage and the description never mentions it. The agent gets no hint about the identifier format, how it maps to the pattern, or whether it must match prior sessions, leaving the schema to speak entirely for itself.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (resume) plus the concrete state it restores: latest task, scope, findings, open work, next actions. That is far more specific than a tautology, but it never distinguishes itself from the adjacent project_open or scope_read siblings, leaving the agent to infer the boundary.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The phrase 'after restart' implies the trigger condition, which is real usage guidance. However, it gives no when-not guidance and never names an alternative (e.g., project_open for starting fresh), so the agent must guess at the routing.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

scope_readC
Read-onlyIdempotent

Read scope and expiry.

ParametersJSON Schema
NameRequiredDescriptionDefault
projectYes

TDQS

C2.4/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, destructiveHint=false, and openWorldHint=false, so the safety profile is fully covered by structured data. The description only adds that expiry information is surfaced, which is a modest hint but not meaningful behavioral disclosure.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

It is a single short fragment with no filler, which is efficient, but that brevity comes at the cost of under-specification rather than genuine conciseness.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema and no schema descriptions, the description carries the full burden yet omits what 'scope' contains, the return shape, and how this relates to the many sibling tools. An agent still cannot call it confidently.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% and the single 'project' parameter is documented only by a regex pattern. The description does not explain what the project identifier is or how it should be supplied, so it fails to compensate for the schema gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a verb and a resource ('read scope and expiry'), so an agent knows it is a read operation on scope data plus an expiry value. However, 'scope' is never defined and the fragment gives no basis for distinguishing this from siblings like scope_register or evidence_read.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no statement of when to use this tool, when not to, or which sibling to prefer. The agent must infer usage context from the name alone.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

scope_registerB

Record operator-provided authorization; never derive permission from target responses. Exact origins, accounts, environments and expiry required.

ParametersJSON Schema
NameRequiredDescriptionDefault
targetYes
projectYes
rate_limitsYes
scope_expiryYes
allow_loopbackNo
allowed_originsYes
allowed_accountsYes
assessment_ownerYes
authorized_domainsYes
prohibited_actionsYes
authorization_referenceYes
authorized_environmentsYes

TDQS

B3.1/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare this is a non-read-only, non-idempotent, non-destructive, non-open-world write. The description adds meaningful behavioral context beyond that — authorization must originate from the operator and never be inferred from target responses — but it does not disclose what the stored record controls downstream, whether re-registration replaces or accumulates, or what the caller gets back.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two front-loaded clauses: purpose first, then the sourcing constraint and required fields. No filler, though the second clause bundles an important principle with a terse field list and neither is expanded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For an 11-required-parameter, nested-object, stateful tool that likely governs the whole engagement, a two-clause description is thin. With no output schema and no per-parameter schema documentation, the definition should say what the registered scope constrains, what happens to prior registrations, and how it pairs with scope_read.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% across 12 parameters, so the description carries the full documentation burden. It names only four of them (origins, accounts, environments, expiry), leaving target, project, prohibited_actions, rate_limits, assessment_owner, authorization_reference, and allow_loopback completely undocumented in both schema and prose.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

"Record operator-provided authorization" states a specific verb and resource, and the definition makes clear this is a write of an authorization record rather than a lookup. It does not name its natural counterpart (scope_read) or otherwise differentiate itself from siblings explicitly, so it stops short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The clause "never derive permission from target responses" is a genuine usage directive about where the authorization data must come from, and the tool must be called before/independent of any assessment activity. However, it never names an alternative (e.g., scope_read for retrieving the recorded scope) or states when this must be invoked versus those siblings, so usage is only implied.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

security_controlled_getC

One bounded GET, exact authorized origin and controlled account, pinned DNS, no redirects. No state-changing requests.

ParametersJSON Schema
NameRequiredDescriptionDefault
urlYes
accountYes
projectYes
environmentYes
authorizationNo

TDQS

C2.7/5.0
Behavior1/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description asserts 'No state-changing requests', i.e. a read-only operation, while the annotations declare readOnlyHint=false. That is a direct inconsistency. Beyond the contradiction, the description does add genuinely useful behavior (pinned DNS, no redirects, bounded single request), but the read-only claim conflicts with the structured hint.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two terse sentences, front-loaded with the operation and then the safety constraints; no filler. It is arguably over-terse, but every clause earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 5-parameter, security-sensitive network tool with no output schema and 0% schema description coverage, the description should explain the remaining parameters and the account/auth model. It covers behavioral guardrails reasonably but leaves project, environment, and authorization unexplained and contradicts the read-only annotation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% across 5 parameters, so the description carries the full burden. It only loosely gestures at two of them ('authorized origin' -> url, 'controlled account' -> account) and says nothing about project, environment, or authorization, nor about required URI/pattern formats.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a concrete verb+resource: a single bounded HTTP GET to an authorized origin under a controlled account. An agent can tell this is a network-fetch tool rather than one of the read_/store_/task_ siblings, though it never names a sibling or explicitly scopes itself against them.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Constraints imply usage ('exact authorized origin', 'controlled account', 'no state-changing requests') but there is no explicit when-to-use/when-not or alternative tool named. The agent must infer that this is for safe, tightly scoped fetches rather than general HTTP access.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

security_reportC

Generate gated technical/executive reports and a minimal request template.

ParametersJSON Schema
NameRequiredDescriptionDefault
findingYes
projectYes

TDQS

C2.4/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare a non-read-only, non-idempotent, non-destructive, non-open-world operation, yet the description adds nothing about what "gated" means, what access or auth is required, whether reports are persisted, or whether repeated calls create duplicates. For a mutation-flagged tool with zero annotation-free context, this is a real gap.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single short sentence with no padding, but its brevity comes from omission rather than precision — the only content is a vague noun phrase, so nothing is effectively front-loaded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, two required undocumented params, and mutation-style annotations, the description should carry substantial load but delivers one clause. Missing: what the output looks like, what "gated" requires, and what identifiers project/finding accept.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% and both required parameters (project, finding) are constrained only by an opaque regex. The description never mentions either parameter, so it fails to compensate for the schema gap, leaving the agent to guess the identifier format and source.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a verb+resource pair ("Generate ... reports") and mentions a request template, but "gated technical/executive reports" is opaque jargon and the sentence never says what a report covers or how it differs from sibling security tools (security_triage, security_severity, finding_correct). An agent can guess it produces a report artifact but cannot tell which one it needs.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No when-to-use guidance, no prerequisites, and no mention of alternatives among the many security_* siblings. "Gated" hints at some precondition but never states it, leaving the agent to infer when this tool is appropriate.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

security_severityB
Read-onlyIdempotent

CVSS 3.1 base score with metric explanation. Does not infer metric values.

ParametersJSON Schema
NameRequiredDescriptionDefault
vectorYes

TDQS

B3.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover read-only, idempotent, non-destructive, and closed-world behavior. The description adds a meaningful constraint beyond annotations: it will not infer missing metric values, so the caller must supply a complete vector. It still omits output format and error behavior, so it is not a 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two short sentences with no filler. It is front-loaded with the result and follows with the key limitation, making it appropriately sized for a simple single-parameter tool.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple CVSS scoring tool with no output schema, the description covers what is returned and one important constraint. However, it leaves the required vector format undocumented and provides no usage context, which are clear gaps against the 0% parameter coverage.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

There is one required parameter, 'vector', with 0% schema description coverage. The description implies CVSS 3.1 context but does not explain the expected vector syntax, required metric fields, or examples, so it fails to compensate for the schema gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states the specific domain output: a CVSS 3.1 base score with metric explanation. It does not use a clear action verb, but an agent can infer the tool computes or returns severity scoring. Sibling tools are not distinguished, keeping it from a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives no when-to-use guidance and does not mention alternatives such as security_triage or security_report. The sentence 'Does not infer metric values' is a behavioral boundary, not usage guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

security_skillsB
Read-onlyIdempotent

List generic evidence requirements and false-positive conditions.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

B3.1/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, destructiveHint=false and openWorldHint=false, so the safety profile is fully covered. The description adds no further behavioral context (e.g., that it returns a static reference set with no side effects), but it also does not contradict the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with no filler. It is efficient, though slightly terse given the ambiguity of what a 'skill' returns.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-parameter, read-only list tool with no output schema and full annotation coverage, the minimum is met. Still, it never clarifies the shape or scope of what is returned, leaving a gap between the name 'security_skills' and the described content.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so there is nothing for the description to disambiguate; the baseline of 4 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The verb 'List' and the resource 'generic evidence requirements and false-positive conditions' are stated, so the agent knows the content type. However, the name 'security_skills' is never reconciled with the description, and nothing distinguishes it from siblings like security_triage or security_report.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No indication of when this tool should be invoked versus security_triage, security_severity, or security_report. The word 'generic' hints it is non-case-specific, but no explicit trigger or alternative is named.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

security_triageD

Evaluate supplied evidence-backed claims; preserve uncertainty and version decisions.

ParametersJSON Schema
NameRequiredDescriptionDefault
cvssNo
skillYes
titleYes
actualYes
claimsYes
impactYes
findingYes
projectYes
accountsYes
businessNo
endpointYes
expectedYes
regressionYes
environmentYes
remediationYes
reproductionYes

TDQS

D1.8/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare readOnlyHint=false and idempotentHint=false, so this is a non-idempotent mutation of some kind, yet the description never says what is written or versioned. 'Preserve uncertainty and version decisions' weakly hints at versioning behavior, but with 14 required inputs describing a recorded finding, the omission of the actual side effect (a persisted triage record?) is a real gap. No contradiction with annotations, but minimal added context.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The single sentence is compact and front-loaded, so there is no bloat, but neither clause earns its place: both are too abstract to help an agent act. Brevity here reflects under-specification rather than disciplined conciseness for a tool with 14 required fields.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 16-parameter, 14-required mutation tool with nested objects, an enum, and no output schema, the description is essentially empty. It documents no inputs, no effects, and no return expectations; an agent has nothing to work from beyond the raw schema.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% across 16 parameters (14 required), including two nested objects (claims, business) and an undocumented skill enum. The description compensates for none of this: it names no parameter, no expected input shape, and no format. With the schema unable to carry meaning, the description's silence leaves the entire parameter contract unexplained.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The phrase 'Evaluate supplied evidence-backed claims' is an abstract verb+resource that never states what the tool concretely produces or persists, and 'preserve uncertainty and version decisions' is opaque meta-language. Nothing distinguishes it from siblings like security_severity, security_report, or finding_correct. The description reads as near-tautological restatement for a security-triage tool rather than a clear statement of function.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance, no preconditions, and no named alternatives despite a dense sibling set (security_severity, security_report, finding_correct). 'Preserve uncertainty and version decisions' gestures at a behavioral principle but gives no selection criteria. An agent cannot infer when this tool should be chosen over its siblings.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

task_check_duplicateC
Read-onlyIdempotent

Check prior completed exact test; changed deployment must use new state_version.

ParametersJSON Schema
NameRequiredDescriptionDefault
projectYes
signatureYes

TDQS

C2.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, so the safety profile is covered by structured data. The description adds one meaningful behavioral nuance — that state_version must change when the deployment changes — but omits what the check returns, how a duplicate is interpreted, or any side effects. With annotations carrying the safety burden, a 3 is appropriate.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single compact sentence with the check purpose front-loaded and zero wasted words, which is structurally efficient. However, the brevity tips into under-specification for a tool with a nested six-field signature, so conciseness comes at the cost of completeness.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with a nested required signature object, no output schema, and no field-level documentation, the description is far too thin. It does not explain what a duplicate result looks like, how the signature fields combine to define 'exactness,' or how the caller should act on the outcome, leaving significant gaps.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% and the nested signature object has six required fields with no per-field documentation. The description only adds meaning for one field (state_version, via the changed-deployment rule), leaving project, method, endpoint, auth_context, ownership, and test_type entirely undocumented. It does not compensate for the near-total coverage gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a check action against 'prior completed exact test,' implying a duplicate-detection function, but 'exact test' and 'prior completed' are opaque jargon that leave the exact purpose ambiguous. It does not clearly tell an agent that this returns a duplicate verdict, and no sibling is named for contrast, though the sibling list contains no near-equivalent to differentiate against.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The only guidance is the embedded constraint that 'changed deployment must use new state_version,' which is about a parameter value rather than when to invoke this tool. There is no statement of when to call task_check_duplicate versus alternatives, no prerequisites, and no exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

task_updateC

Version task state and record test signature including deployment state.

ParametersJSON Schema
NameRequiredDescriptionDefault
taskYes
statusYes
projectYes
summaryYes
evidenceNo
signatureNo
next_actionsYes

TDQS

C2.4/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=false, destructiveHint=false, idempotentHint=false, and openWorldHint=false, so the safety profile is covered. The description adds a vague notion of 'versioning' and 'deployment state' but never explains what versioning means, whether prior state is retained or overwritten, or what auth is required — leaving the mutation's actual behavior opaque.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

It is a single short sentence with no padding, but its brevity comes at the cost of clarity rather than through efficient front-loading. Reasonably sized, poorly chosen words.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

This is a mutation tool with 7 parameters, 5 required, a nested signature object, and 0% schema coverage, and it has no output schema to fall back on. The description does not explain the nested object, the evidence hashes, or what a successful update returns, so it is significantly under-specified for the tool's complexity.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must carry parameter meaning, yet it only loosely gestures at 'task state' (status) and 'test signature' (signature). The 7 parameters — including project, summary, next_actions, the defaulted evidence hash array, and the nested signature object — are essentially undocumented.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names an action on a task ('version task state') and a secondary behavior ('record test signature'), so an agent can broadly infer it mutates a task record. However, 'Version' is ambiguous jargon — it doesn't state the resource or verb cleanly (update? transition? snapshot?) — and it draws no line against siblings like task_check_duplicate.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to call this tool versus alternatives such as task_check_duplicate or project_resume. No prerequisites, no exclusions, no context on when a task should be updated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 20 tool updatesv0.1.0
    • First observedaudit_events
    • First observedevidence_read
    • First observedevidence_store
    • First observedfinding_correct
    • First observedmemory_context
    • First observedmemory_graph
    • First observedmemory_search
    • First observedmemory_store
    • First observedmodel_advise
    • First observedproject_open
    • First observedproject_resume
    • First observedscope_read
    • First observedscope_register
    • First observedsecurity_controlled_get
    • First observedsecurity_report
    • First observedsecurity_severity
    • First observedsecurity_skills
    • First observedsecurity_triage
    • First observedtask_check_duplicate
    • First observedtask_update

TDQS

B3/5.0

Scored across 20 tools

Disambiguation5/5

Each tool targets a distinct resource and action, with clear boundaries between evidence, memory, tasks, scope, and security operations. The security_* cluster is nuanced but differentiated by purpose: skills list requirements, triage evaluates claims, severity computes CVSS, report generates deliverables, and controlled_get performs a bounded request.

Naming Consistency4/5

All names use lowercase snake_case with a consistent domain prefix, making the set easy to scan. Most follow a noun_verb pattern, but a few are noun_noun (e.g., security_skills, security_severity, memory_context, audit_events), which is a minor deviation.

Tool Count4/5

20 tools is slightly above the typical 3-15 sweet spot, but the domain spans projects, evidence, tasks, security triage, scope control, memory, and audit. Each tool appears to earn its place, so the count is reasonable rather than bloated.

Completeness4/5

The surface covers core project, evidence, scope, memory, security, finding, and audit workflows. Minor gaps exist, such as no explicit task_create/list or project_list/close, but immutable and versioned designs mitigate them and agents can work around via existing tools.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • A
    license
    A
    quality
    C
    maintenance
    MCP server that indexes Markdown, Word, HTML, and PDF documents into a SQLite knowledge graph with CJK+Latin full-text search and cross-document reference tracking. Runs drift audits to surface stale policies, conflicting research claims, superseded ADRs, and undocumented code exports.
    10
    8
    MIT
  • A
    license
    B
    quality
    A
    maintenance
    Local-first, auditable memory for Codex, Claude Code, and MCP clients. It stores scoped user/project memory in SQLite or Postgres, serves read-only recall and inspection tools by default, and supports opt-in governed writeback with review and forget controls.
    8
    145 npm
    18
    MIT
  • A
    license
    Not graded
    quality
    B
    maintenance
    Local static-analysis assistant for Android malware research that manages investigation cases, exposes MCP tools via a local server, and persists evidence-backed findings without cloud dependency.
    MIT
  • A
    license
    A
    quality
    B
    maintenance
    A local-first MCP server for retrieving a small evidence set and recording reviewed conclusions, policy-gated and redacted without giving an agent general filesystem access.
    5
    MIT