Aegis MCP
Allows auditing existing Docker containers for confidentiality readiness, gathering read-only system evidence and producing professional reports. Docker targets are audit-only; Aegis does not provision, restart, or rebuild containers.
Provides read-only audits and supported hardening for existing Linux systems (local, WSL, or SSH-accessible), including evidence collection, remediation plans, verification of changes, and report generation.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@Aegis MCPAudit my registered wsl target and generate a confidentiality report."
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
Aegis MCP
An English, ISO-aligned confidentiality audit assistant for your existing Linux systems.
Connect an MCP-compatible AI client, register your VPS, WSL instance or Docker container, and work through ten hardening steps. Aegis gathers bounded system evidence, builds reviewable remediation plans, verifies supported changes, and supplies a ready-made professional report template. There is no bundled lab, agent daemon or required AI API key.
Project subject: Summary of a Ten-Step Best-Practice Guide to the Hardening Process of an Information System in Preparation for a Confidentiality Audit.
The embedded guide and report both follow that subject: ten best-practice steps for an information system preparing for a confidentiality audit. ISO/IEC 27001:2022 and ISO/IEC 27002:2022 provide suggested control mappings. They do not turn individual technical checks into certification claims.
Report design · AI connections · Ten-step guide · Architecture · Hardening boundaries

The preview uses synthetic evidence to demonstrate the template. It is not an audit of a real system.
Start in three minutes
Requires Python 3.11+, uv and a Linux execution environment. Run the server in WSL when using a Windows machine. SSH targets and existing containers need Python 3.9+ inside the target; missing tools are recorded as missing evidence.
uv sync --frozen
cp targets.example.yaml targets.yaml
# Edit targets.yaml: keep only targets you own and intend to assess.
uv run aegis doctor
uv run aegis-mcpaegis-mcp waits for an AI client over stdio. It does not print a dashboard. Register the launch command in your MCP client using the connection examples.
You can verify setup before connecting an AI:
uv run aegis guide
uv run aegis targets
uv run aegis audit wsl
# Save the returned audit_id and use it in the following commands.
uv run aegis plan wsl --audit <audit_id>
uv run aegis report --audit <audit_id>The report command fills the existing template with deterministic summary prose. A connected AI can supply a richer narrative through get_report_context and render_report.
Related MCP server: sec-ops-mcp
Connect your AI
For clients using the common mcpServers configuration shape:
{
"mcpServers": {
"aegis": {
"command": "uv",
"args": ["--directory", "/absolute/path/to/audit_mcp", "run", "--frozen", "aegis-mcp"],
"env": {
"AEGIS_TARGETS": "/absolute/path/to/audit_mcp/targets.yaml",
"AEGIS_WORKSPACE": "/absolute/path/to/audit_mcp/.runtime",
"AEGIS_ALLOW_APPLY": "0"
}
}
}
}Use absolute paths and make uv available to the client process. Client configuration shapes differ; see connection examples. Aegis works with MCP-compatible AI hosts. A model without an MCP host needs an adapter. Your chosen host handles model credentials and inference; Aegis does not call a paid model itself.
Suggested first request:
Use Aegis to summarize the ten-step hardening best-practice guide, audit my registered
wsltarget for confidentiality readiness, and produce a professional English report using the embedded template. Keep missing evidence and human review items visible. Do not apply target changes.
The ten steps
Step | Best-practice domain | Confidentiality purpose | Suggested ISO controls |
01 | Scope, inventory and classify | Know which assets and data need protection | A.5.9, A.5.12, A.5.13 |
02 | Maintain a patched baseline | Reduce exposure from vulnerable software | A.8.8, A.8.9 |
03 | Constrain identities and access | Permit access only to approved identities | A.5.15, A.5.16, A.5.18, A.8.2, A.8.5 |
04 | Separate network trust zones | Restrict unintended paths to sensitive services | A.8.20–A.8.22 |
05 | Protect cryptography and keys | Protect confidential data in transit and storage | A.8.24, A.5.17 |
06 | Reduce system attack surface | Reduce information leaks and excess privilege | A.8.9, A.8.19, A.8.27 |
07 | Minimize data exposure | Reduce permissive files and data disclosure | A.8.3, A.8.11, A.8.12, A.5.34 |
08 | Observe and investigate access | Produce usable access and incident evidence | A.8.15–A.8.17 |
09 | Secure backups and prove recovery | Protect copies and demonstrate recovery | A.8.13, A.5.30 |
10 | Validate and build the audit dossier | Preserve evidence, review gaps and verify changes | A.5.35, A.5.36, A.8.32 |
All ten domains are always present. Technical checks remain pass, fail, unknown or not_applicable; organizational and application controls remain independent review items.
What the MCP exposes
Tool | Result |
| Exact English subject, ten steps and ISO links |
| Locally registered IDs and capabilities |
| Read-only target observations and a SHA-256 audit ID |
| Sealed plan, exact diffs, prerequisites and manual work |
| Supported changes, target backup and after evidence |
| Restore managed settings and collect new evidence |
| Status transitions between snapshots |
| Compact ten-step evidence, writing brief and narrative schema |
| Offline HTML, Markdown and JSON files |
Four resources supply the guide, report director, reusable report structure and standards mappings. Two prompts guide the assessment and professional report workflow.
Professional report, with little model effort
The design is built in: editorial cover, mint accents, clear typography, ten-step evidence sections, ISO control chips, status filters, search, a before/after comparison and a print-to-PDF layout. It requires no external fonts, analytics or network requests.
The AI receives compact facts and a strict narrative schema. It drafts the title, executive summary, priorities and optional step notes once. The server generates all tables and styling. Narrative cannot override collected counts or statuses. HTML is escaped and protected by a restrictive content-security policy.
Reports and system evidence stay in .runtime/, which is ignored by Git. Evidence may include operational hostnames, usernames and configured paths: review a report before sharing it externally.
Supported hardening
Audit collection is read-only on the target. Target changes require both AEGIS_ALLOW_APPLY=1 in the server environment and allow_apply: true for the target. The operator still controls conversational authorization through the AI host.
The current automatic catalogue is deliberately explicit:
Linux kernel information-disclosure and link-protection settings:
dmesg_restrict,kptr_restrict,protected_hardlinks,protected_symlinks.A managed login-shell
umask 027drop-in. Existing processes and service policies are unaffected.Key-only SSH and disabled direct root SSH login, after independent key access and configuration prerequisites are verified locally.
Patching, firewall changes, application IAM, volume encryption, MFA, logging custody, backup encryption and restore exercises are guided work requiring environment-specific review. Docker targets are audit-only; WSL kernel controls are inherited from its host. Aegis does not provision, restart or rebuild containers.
Read the exact prerequisites and rollback behavior before enabling target writes.
Development and verification
uv sync --frozen
uv run ruff check src tests scripts
uv run pytest -q
uv run python scripts/check_repository.py
uv run pip-audit --skip-editable
uv buildTests cover the real MCP protocol using synthetic target evidence, target/option injection, evidence tampering, local permission gates, stale plans, rollback, inherited-control handling, report escaping and content-security hashes. Adapter and operational test boundaries are documented in validation.
Repository layout
src/aegis_mcp/ MCP server, bounded collector, adapters, evidence and reporting
assets/prompts/ Ten-step workflow and one-pass report director
assets/templates/ Professional offline HTML/CSS/JS template
docs/ Setup, guide, architecture, hardening and template preview
examples/ MCP host configuration examples
tests/ Protocol, security and report regression tests
scripts/ Publication checks and synthetic template preview
targets.example.yaml Operator target configuration template
uv.lock Locked dependenciesISO mapping references: ISO/IEC 27001:2022, ISO/IEC 27002:2022. This repository contains original guidance and control identifiers, not copied ISO standard text. See SECURITY.md and MIT license.
Available Tools
9 toolsapply_hardening_planADestructive
Apply a reviewed, sealed plan to its registered target; requires explicit local write opt-ins and root access. Backs up, validates and collects after evidence.
| Name | Required | Description | Default |
|---|---|---|---|
| plan_id | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already flag destructiveHint=true, idempotentHint=false and readOnlyHint=false, so the safety profile is covered. The description adds genuinely new behavioral context — that it backs up, validates, and collects evidence afterwards — which materially affects an agent's risk assessment and expectation of recoverability.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single dense sentence that front-loads the action and follows with preconditions and side effects. No filler, though the 'collects after evidence' phrasing is slightly compressed and could be clearer.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return-value explanation is unnecessary. For a destructive mutation tool the description supplies prerequisites, side effects (backup, validation, evidence collection), and hints of recoverability, leaving only the plan_id contract underspecified.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
There is a single parameter with 0% schema description coverage, so the schema documents nothing beyond the name. The description implies the plan reference via 'reviewed, sealed plan' but gives no format, provenance, or validity requirements for plan_id. It partially compensates but not fully.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific verb and resource ('Apply ... plan to its registered target') and constrains it to a 'reviewed, sealed' plan, which implicitly separates it from create_hardening_plan and rollback_hardening_plan. It is clear what the tool does, though it never names the sibling alternatives explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
'Apply a reviewed, sealed plan' implies a plan must first be created and reviewed, and the description adds two concrete preconditions (local write opt-ins, root access). However, there is no explicit when-to-use/when-not statement and no routing to alternatives such as rollback_hardening_plan if the application fails.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
audit_targetB
Collect read-only Linux evidence and save a local audit snapshot. No target configuration changes.
| Name | Required | Description | Default |
|---|---|---|---|
| target_id | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=false, destructiveHint=false and openWorldHint=false, so the safety profile is largely covered. The description usefully reconciles the non-read-only hint by stating it writes a local snapshot while making no target configuration changes, but it omits other behavioral details such as where the snapshot lands, how long collection takes, or permission requirements.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short, front-loaded sentences; the core action comes first and the scoping caveat second. Nothing is padded or redundant.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be described. Still, for a tool whose whole job is producing a saved artifact, the description omits where the snapshot is stored and what target_id must match, leaving moderate gaps given the 0%-coverage schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The single parameter target_id has 0% schema description coverage, so the description carries the full burden for it and says nothing about what a target id is or how to obtain one (e.g., via list_targets). The phrase 'target configuration changes' only obliquely hints that target_id refers to a managed host.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific verb pair (collect, save) and concrete resources (read-only Linux evidence, local audit snapshot), so the agent knows exactly what is produced. It does not, however, differentiate itself from siblings like get_report_context or compare_audits, which also deal with audit data.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no explicit when-to-use guidance and no named alternative. The agent must infer that this is the evidence-collection step versus compare_audits or create_hardening_plan, and there is no mention of prerequisites such as a valid target existing.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
compare_auditsBRead-only
Compare evidence status changes for the same target registration.
| Name | Required | Description | Default |
|---|---|---|---|
| after_audit_id | Yes | ||
| before_audit_id | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, destructiveHint=false and closed-world scope, so the safety profile needs no restating. The description adds the meaningful constraint that both audits must belong to the same target, but says nothing about what is compared (which evidence fields), how differences are surfaced, or ordering assumptions.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with no filler or redundancy. It is appropriately sized for a two-parameter tool, though its brevity leaves the ordering semantics unresolved.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists so return values need not be explained. Still, with 0% schema description coverage the definition should clarify parameter ordering and what 'evidence status' denotes; as written it is minimally sufficient rather than complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%: both before_audit_id and after_audit_id are undocumented in the schema. The description alludes to temporal ordering via 'changes' but never states that before_audit_id is the earlier audit or how the two IDs relate to the same target, so it only partially compensates for the coverage gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Specific verb ('Compare') plus resource ('evidence status changes') and an explicit scope constraint ('for the same target registration'). This clearly separates it from siblings like audit_target or get_hardening_guide, though it never names an alternative to route against.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'for the same target registration' implies a usage precondition and rules out comparing audits of different targets. However, there is no explicit when-to-use trigger, no when-not guidance, and no named alternative, so usage must be inferred.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
create_hardening_planA
Build a sealed plan with exact diffs, supported actions, prerequisites and manual work. Does not modify the target.
| Name | Required | Description | Default |
|---|---|---|---|
| audit_id | Yes | ||
| target_id | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare readOnlyHint=false (a plan resource is created) while the description clarifies that the target itself is left untouched — a genuinely useful disambiguation, not a contradiction. It also discloses what the plan contains and that it is 'sealed'. It stops short of stating auth needs or whether plans can be regenerated.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences, front-loaded with the verb and resource, with the key non-mutation caveat at the end. No filler or repetition.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return shape need not be described, and the description covers the plan's contents adequately. However, with two opaque required parameters and no usage routing, it leaves gaps for a tool sitting between audit_target and apply_hardening_plan in a multi-step workflow.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Two required parameters (audit_id, target_id) have 0% schema description coverage, so the schema contributes nothing. The description mentions neither parameter nor their relationship, leaving their meaning and the requirement for a prior audit entirely to inference from the names.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Build a sealed plan') and enumerates the plan's contents (diffs, supported actions, prerequisites, manual work). The closing clause 'Does not modify the target' implicitly separates it from apply_hardening_plan, though the sibling is never named.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is only implied: it creates a plan, so an agent infers this precedes apply_hardening_plan. It never states when to prefer this over get_hardening_guide or that an audit must exist first, and names no alternative or prerequisite workflow.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_hardening_guideBRead-only
Return the exact English subject and the ten-step confidentiality hardening guide.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already establish readOnlyHint=true, destructiveHint=false, and openWorldHint=false, so the safety profile is covered. The description adds useful content-level detail (it returns a fixed ten-step guide plus a subject string), but omits any note on idempotency, static vs dynamic content, or rate considerations. With annotations carrying the safety burden, this is adequate but not rich.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with no filler. The only cost is that 'the exact English subject' is unexplained, which is a minor clarity issue rather than verbosity.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Output schema exists, so return values need no explanation, and there are no parameters to document. The description states what the tool yields, which is sufficient for a no-arg retrieval tool, though it never clarifies what the 'subject' is for.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so the baseline is 4 and there is no parameter meaning to document. The description correctly spends no words re-explaining an empty schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb (Return) plus the resource (the ten-step confidentiality hardening guide), which is clearly distinct from action-oriented siblings like create_hardening_plan or apply_hardening_plan. The phrase 'the exact English subject' is cryptic on its own and slightly muddies otherwise clear intent.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No when-to-use context or routing guidance is offered. With zero parameters the tool is likely always callable, but the description never says whether it should precede create_hardening_plan, be called standalone, or how it relates to get_report_context/render_report.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_report_contextBRead-only
Get the compact ten-step evidence bundle, report director and narrative schema. The server handles design and tables; draft the narrative once.
| Name | Required | Description | Default |
|---|---|---|---|
| audit_id | Yes | ||
| before_audit_id | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover the safety profile (readOnlyHint=true, destructiveHint=false). The description adds context about the division of labor—the server handles design and tables, so the agent should focus on drafting the narrative—but it does not disclose permissions, rate limits, or output size beyond 'compact.'
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences with no filler, and the main action is front-loaded. The second sentence provides a useful directive, though the term 'report director' is slightly opaque and could be clearer without adding length.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so the description needn't explain return values, and annotations cover safety. However, the complete absence of parameter semantics leaves the agent without guidance on what audit_id and before_audit_id do, which is a notable gap for a tool whose primary job is fetching context.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, and the description does not mention either parameter. With two parameters and no schema descriptions, the description needed to compensate by explaining the purpose of audit_id and before_audit_id, but it adds no semantic meaning at all.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a clear verb ('Get') and specifies the resources returned: a compact ten-step evidence bundle, report director, and narrative schema. It distinguishes the tool's role from rendering by noting that the server handles design and tables, but it does not name any sibling tools explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage by instructing to 'draft the narrative once' and noting the server handles design and tables, which gives context for when to call this tool. However, it does not state when to use this versus alternatives like render_report or compare_audits, nor does it provide any exclusions or prerequisites.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_targetsARead-only
List operator-registered target IDs and capabilities. Does not connect to targets.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and destructiveHint=false, so the safety profile is known. The description adds the meaningful guarantee that it does not connect to targets (no live side effects), which is useful context beyond the annotations. It does not contradict them.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences, front-loaded with the action and followed immediately by the important behavioral caveat. Every word earns its place with no redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return-value structure need not be explained, and for a zero-parameter read tool the description covers purpose and side-effect behavior adequately. Only minor details (ordering, empty-result behavior) are left unstated.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so the baseline is 4. The description briefly characterizes what is listed (target IDs and capabilities), which is the only semantic content possible for a no-argument tool.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('List operator-registered target IDs and capabilities') with a clear qualifier ('operator-registered') that scopes the result set. No sibling tool performs enumeration, so it is implicitly distinguished from audit_target, the hardening-plan tools, and the report tools without needing to name them.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is only implied: 'Does not connect to targets' hints that this is a cheap, offline enumeration step versus the live work done by audit_target or apply_hardening_plan. There is no explicit statement of when to call it or what alternative to prefer, so an agent must infer the workflow placement.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
render_reportB
Fill the supplied professional template and export offline HTML, Markdown and JSON. Narrative is escaped commentary; observed tables cannot be overridden.
| Name | Required | Description | Default |
|---|---|---|---|
| audit_id | Yes | ||
| narrative | No | ||
| before_audit_id | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only declare it is a non-read-only, non-destructive, closed-world operation. The description adds real behavioral context beyond that: narrative input is escaped and treated as commentary, and observed tables cannot be overridden. That tells the agent the tool's write scope is constrained, which is genuinely useful.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tightly written sentences with the action front-loaded and the constraint second. No filler, no repetition of the name or schema.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists so return values need not be explained, but for a mutating/exporting tool with three completely undocumented parameters and no usage context, the description is too thin. It should at minimum explain what audit_id and before_audit_id select.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% across three parameters, so the description must compensate and largely does not. It hints that 'narrative' is escaped commentary but says nothing about audit_id, before_audit_id, or the structure of the narrative object. The agent is left to infer required inputs from the schema alone.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a concrete verb (fill/export) and resource (report), plus the output formats (offline HTML, Markdown, JSON). It distinguishes itself reasonably from siblings like get_report_context, though it never names them explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no explicit when-to-use guidance, no prerequisites stated for the template, and no mention of alternatives among the eight sibling tools. An agent must infer that this is the final rendering step after auditing.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
rollback_hardening_planADestructive
Restore the backed-up managed configuration for a successfully applied plan; rejects drift and collects after evidence.
| Name | Required | Description | Default |
|---|---|---|---|
| plan_id | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true, idempotentHint=false and readOnlyHint=false, so the mutation profile is covered. The description adds genuinely new behavior: it rejects drift (a guard condition) and collects after-evidence, which the annotations cannot express.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with the action first and the guard conditions after the semicolon. 'Collects after evidence' is slightly jargon-heavy but not wasteful.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be described. For a single-parameter destructive mutation, the description covers the precondition, the drift-rejection behavior, and the evidence collection, leaving only plan_id provenance unstated.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
One required parameter with 0% schema description coverage, so the schema does no work here. The description adds only indirect meaning for plan_id by requiring it reference a successfully applied plan; it does not document the id format or source.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Restore) and resource (backed-up managed configuration) scoped to 'a successfully applied plan', which distinguishes it from apply_hardening_plan. It stops short of explicitly naming that sibling, so differentiation requires inference.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'for a successfully applied plan' is an implicit precondition, and 'rejects drift' hints at a failure condition, but the description never states when to roll back versus re-applying or auditing, nor any exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
9 tool updates
v0.1.0- First observed
apply_hardening_plan - First observed
audit_target - First observed
compare_audits - First observed
create_hardening_plan - First observed
get_hardening_guide - First observed
get_report_context - First observed
list_targets - First observed
render_report - First observed
rollback_hardening_plan
TDQS
Scored across 9 tools
Each tool maps to a distinct stage of the hardening workflow (guide, targets, audit, plan, apply, rollback, compare, report). The only mild overlap is between get_hardening_guide and get_report_context, both informational retrievals, but descriptions make their purposes clearly separable.
All nine tools follow a consistent snake_case verb_noun pattern (get_, list_, audit_, create_, apply_, rollback_, compare_, render_). No mixed conventions or stylistic deviations.
Nine tools are well-scoped for a focused Linux hardening and reporting workflow, with each tool earning its place at a distinct lifecycle stage. No redundant or trivially thin tools.
The surface covers a full lifecycle: guidance, target listing, audit, plan creation, apply, rollback, comparison, and reporting. The only apparent gap is a register/remove target operation, since targets are described as operator-registered out of band.
Maintenance
Related MCP Connectors
Compliance frameworks (SOC 2, ISO 27001, CMMC, NIST, more) delivered to AI agents as MCP tools.
Cited, standards-aware compliance overlay for AI assistants (ISO, NIST, FedRAMP, IRAP), over MCP.
Tenant-scoped evidence intake and controlled AI context over MCP.
Remote MCP for Android CLI agent build gate, structured receipts, audit logs, and reviewer-ready evi
Related MCP Servers
- FlicenseAqualityDmaintenanceA secure, read-only MCP server for AI-powered system monitoring. It provides real-time OS metrics, config discovery, and safe log tailing to enable autonomous infrastructure audits without shell access risks.41-
- AlicenseNot gradedqualityBmaintenanceEnables LLMs and AI agents to perform defensive security posture assessments, privilege escalation surface audits, and post-quantum cryptography readiness checks through read-only diagnostic tools.MIT
- AlicenseNot gradedqualityCmaintenanceEnables authorized security auditing of AI-agent supply chains and agent-facing surfaces: deterministic local skill-bundle audits against eight attack patterns, secrets scanning, and scope-gated read-only recon of agent endpoints and MCP surfaces.MIT
- AlicenseNot gradedqualityCmaintenanceEnables enterprise MCP security auditing through 12 deterministic no-LLM tools for evidence-gated claims, skill/prompt supply-chain audits, token profiling, and server auth-mode checks.MIT