apsa
Audits Android apps and APK builds, analyzing DEX calls, resources, network configuration, exported components/providers, signing-block and v1 certificate evidence, and ELF hardening; also correlates Android security advisories and plans runtime scenarios for owned Android test apps.
Correlates Apple security advisories with audited mobile app inventories to identify version-affected vulnerabilities and support report reassessment.
Audits iOS apps and IPA builds, analyzing Mach-O headers, embedded entitlements, ATS exceptions, provisioning indicators, and PIE/canary/string evidence; supports runtime planning for iOS simulator apps but not physical devices.
Maps mobile app security checks to OWASP MASVS/MASTG guidance and provides OWASP coverage reporting, while not certifying compliance.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@apsascan ./app-release.apk and summarize the findings"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
APSA
Evidence-first security audits for Android and iOS.
This English README is the source of truth. The Korean README follows it.
APSA helps developers and security teams audit their own mobile apps. It inspects source code and APK/IPA builds, correlates public vulnerability information, and keeps evidence, coverage, and report history together. Use the CLI, terminal UI, or the same audit engine through MCP and reusable skills.
Pronounced “ap-sah”; Korean name 앱사. The name connects “app + audit” with App Security Audit. APSA combines the earlier Quaygate lint engine and Mobile Audit workflows in one package.
What it checks
Area | Available checks |
Source code | Java/Kotlin/Swift AST analysis; WebView and deep-link patterns; Manifest, Info.plist, storage, and dependency inspection |
Android builds | DEX calls and constant flow; resources and network configuration; exported components/providers; signing-block and v1 certificate evidence; ELF hardening |
iOS builds | Mach-O headers; limited entitlement/configuration checks (embedded XML entitlements, ATS exceptions, provisioning indicators); PIE, canary, and string evidence |
Public intelligence | Apple/Android advisories, CVE, CISA KEV, OWASP guidance, and OSV dependency correlation |
Reports and CI | SQLite history, comparison and reassessment, JSON/Markdown/SARIF export, coverage requirements, and expiring waivers |
Runtime | Prepared scenarios for owned Android test apps and iOS simulator apps; physical iOS devices are unsupported |
Model integration | stdio MCP tools/resources and packaged skills, without a required model provider or LLM API key |
Findings distinguish candidate, configuration-confirmed, version-affected, and runtime-confirmed evidence. coverage and warnings show what actually ran. A signing block does not prove signature authenticity, and an affected dependency version does not prove exploitability. A scan with no findings does not establish that the whole app is secure.
OWASP mappings describe relevant checks; APSA does not certify MASVS compliance or implement every MASTG test. Public advisories cannot reveal undisclosed zero-days. App files alone cannot establish a device's OS patch state. See OWASP coverage and the support matrix, both currently in Korean, for tested scope and limits.
Related MCP server: saglitzsecure-mcp
Install and run
Install uv on macOS or Linux. APSA targets CPython 3.11 and 3.12; not every host/Python combination has been tested (see the support matrix). The examples select Python 3.12, which uv can download if needed. Initial installation can use the network; scans can then use local inputs and cached intelligence.
Install the published package from PyPI:
uv tool install --python 3.12 apsa==1.0.3
apsa --version
apsa doctor --json
apsa demo --out ./apsa-demoThe package provides apsa and the compatibility aliases quaygate and mobile-audit. If the command is not found, run uv tool update-shell and open a new terminal. GitHub Releases provides the wheel, source distribution, checksums, and a verified bundle with locked requirements, attribution, and the build manifest.
For development or installation with the repository's locked dependencies, use a checkout:
git clone https://github.com/ictechgy/apsa.git
cd apsa
uv sync --locked --python 3.12
uv run --locked apsa doctor --json
uv run --locked apsa demo --out ./apsa-demodoctor checks parsers, optional device tools, and offline readiness. demo writes an intentionally vulnerable example and scans it; choose a new output directory.
To register the checkout's commands on your PATH:
uv tool install --editable . --force --python 3.12 --constraints requirements-release.txt
apsa --versionThis registers apsa and the compatibility aliases quaygate and mobile-audit. --force replaces existing tools with those command names. An editable installation depends on this checkout; keep it in place. If the command is not found, run uv tool update-shell and open a new terminal. The examples below assume apsa is on PATH. Without a global installation, prefix them with uv run --locked from the checkout.
apsa scan /path/to/owned/mobile-project
apsa scan /path/to/owned/app.apk
apsa scan /path/to/owned/app.ipa --sbom /path/to/build.cdx.json
apsa tuiUse a CycloneDX JSON SBOM from the actual build to improve dependency correlation. tui, or apsa without arguments, opens the terminal interface. Run apsa COMMAND --help for options.
Public intelligence and network use
A default scan reads local files and cached intelligence without uploading source code or builds. Public-feed collection and OSV queries use the network:
Command | Network behavior |
| Uses local inputs and cached intelligence |
| Fetches public vulnerability sources |
| Polls public sources and reassesses saved inventories; does not reread app files or implicitly query OSV |
| Sends discovered dependency names and versions to OSV |
| Also sends saved dependency names and versions to OSV |
apsa intel sync
apsa intel status
apsa intel watch --interval 900
apsa scan /path/to/owned/app --onlineWatch polls every 900 seconds by default, with a 60-second minimum; it is not a push stream. It runs until interrupted unless --cycles sets a finite number of polls. Reassessment saves a new snapshot when findings, coverage, or intelligence state changes. Inspect feed freshness, failures, and pending CVE processing with intel status. Fetching a CVE document and completing its processing are separate states. The default pending-item policy allows no backlog; use intel_max_pending to set an explicit allowance.
apsa reports reassess latest
apsa reports compare audit_BEFORE audit_AFTER
apsa reports export latest --format sarif --out audit.sarif
apsa reports verifyReassessment applies current intelligence to a saved inventory; run scan again for changed files or new static checks. If connected to an AI client, that client may send report metadata to its model provider. APSA's MCP context excludes source excerpts, full source files, and screenshot bytes; client-side data handling still depends on the client.
CI and background jobs
apsa policy init --out apsa.toml
apsa scan /path/to/owned/app --policy apsa.toml --out audit.json
apsa scan /path/to/owned/app --fail-on high --include-candidates
apsa scan /path/to/owned/app --background --json
apsa jobs status JOB_ID --jsonSeverity gates exclude candidate findings by default. Opt in with --include-candidates or the policy's allowed_statuses. Required rules accept only checked or not-applicable coverage; partial execution and missing required checks do not pass. Waivers need a finding ID, a reason, and an expiry date. A background job being completed means it finished; check its audit_incomplete flag and report before treating the audit as complete.
Exit code | Meaning for the unified CLI |
| Command completed or policy passed |
| Execution error |
| Invalid arguments |
| Incomplete audit or policy, failed report verification, or failed or partial intelligence synchronization |
| Findings exceeded the configured CI threshold |
| Interrupted |
--json emits an envelope containing ok, data or error, and exit_code; watch emits one JSON envelope per cycle (NDJSON). ok is true for codes 0 and 4; code 4 means evaluation succeeded but the CI threshold was exceeded. CI must check exit_code and the policy result. Reports may still be produced for codes 3 and 4. See the CI example and operations guide (Korean) for policies, backup, limits, and troubleshooting.
MCP and skills
apsa integrations --root /absolute/path/to/owned-apps
apsa mcp --root /absolute/path/to/owned-apps
apsa skill install
apsa context --report latest --jsonUse integrations to generate a configuration with the installed executable path, or a python -m apsa fallback when no apsa executable is found. The shape below is illustrative; replace both absolute paths:
{
"mcpServers": {
"apsa": {
"command": "/absolute/path/to/apsa",
"args": ["mcp", "--root", "/absolute/path/to/owned-apps"]
}
}
}MCP uses stdio and requires an explicit --root; repeat it for multiple roots. integrations defaults to the current directory when no root is supplied. Roots restrict audited targets and access to their reports and jobs. --allow-any-root explicitly removes that restriction. APSA does not change model-client configuration automatically; client authentication belongs to the client.
Purpose | MCP tools |
Audits |
|
Jobs |
|
Reports |
|
Intelligence |
|
Policy |
|
Runtime planning |
|
Resources include apsa://rules and apsa://reports/{report_id}. The previous quaygate:// and mobile-audit:// resource schemes remain compatible.
The default server exposes runtime planning. runtime_execute and runtime_start are registered only with --allow-runtime at server startup. Both preview a scenario by default; execute=true runs it. runtime_start starts a cancellable background device job. Runtime tests need an authorized, prepared test app; the default audit does not boot devices or install apps. See the operations guide (Korean) before running a scenario.
skill install copies the packaged APSA skill to ~/.codex/skills/apsa. To install directly into another model runtime's skill directory:
apsa skill install --name apsa --dest /path/to/runtime/skills/apsaUse --name quaygate or --name mobile-audit to update a skill installed under an older default name; custom edits are preserved unless --force is supplied. integrations reports skill status. For clients without MCP, context provides report context that excludes source excerpts, full source files, and screenshot bytes.
Compatibility and stored data
quaygate and mobile-audit invoke the same unified CLI. Python entry points python -m apsa, python -m quaygate, and python -m mobile_audit remain available. Existing reports and QG-* rule IDs retain their identity.
Data is stored by default in ~/.local/share/mobile-audit. Path precedence is --home → APSA_HOME → QUAYGATE_HOME → MOBILE_AUDIT_HOME → that default. Renaming does not copy the database or move history. Existing MCP configurations must include an authorized --root.
The legacy apk, ipa, and device subcommands retain quick lint output and exit codes 0/1/2; they do not save unified audit history or correlate vulnerability intelligence. Use scan for the full workflow. Older cached OSV records without severity need a fresh scan --online or intel watch --online; reassessment alone cannot recover missing scores.
Development, validation, and licensing
uv sync --locked --extra dev --python 3.12
make test benchmark
make export-release
make release RELEASE_OUT=dist/apsa-local-releaseChoose a new or empty release directory. Release verification requires uv 0.12.1 and builds wheel/sdist twice, compares their hashes, and checks a clean installation outside the checkout with offline source/APK scans, MCP, and skill installation. It writes hashes, an SBOM, dependency notices, and a release manifest without publishing. Initial dependency preparation can use the network. The same-host repeat check does not claim byte-identical builds across platforms.
Tagged releases use GitHub Actions to publish the verified distributions to PyPI after the supported CI matrix passes. See release publishing for the workflow and download contents.
Recorded product validation is in RELEASE_READINESS.md. The curated benchmark is a regression corpus, not a measure of production detection rates. Integration boundaries and the threat model are currently in Korean. Historical reviews remain tied to their original snapshots.
The source is publicly available on GitHub. LICENSE preserves the original Quaygate MIT notice. This publication does not declare an additional license for the combined product.
Available Tools
17 toolsaudit_reassessB
Re-evaluate a saved inventory against current cached intelligence and save a new report.
| Name | Required | Description | Default |
|---|---|---|---|
| report_id | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations cover the safety profile (readOnlyHint=false, destructiveHint=false, openWorldHint=false), and the description's 'save a new report' is consistent, adding the useful detail that a new artifact is produced rather than an in-place update. It does not disclose whether prior reports are overwritten, permission requirements, or how cached intelligence is refreshed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence that states the action, inputs, and outcome with zero filler. Every clause earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values needn't be explained, and the operation is simple with one parameter. However, the description omits prerequisite/permission context and any routing advice relative to the many sibling tools, leaving the definition only minimally complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% for the single report_id parameter, so the description must carry the burden. It refers to 'a saved inventory' and 'a new report' but never clarifies what report_id identifies or its required format, leaving the parameter's semantics under-specified.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and scope: re-evaluate a *saved* inventory against *current cached intelligence* and save a new report. The word 're-evaluate' plus 'saved inventory' distinguishes it from siblings like audit_scan or audit_start, which imply fresh scans, without naming them explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no explicit when-to-use, when-not-to-use, or alternative routing. The phrase 'saved inventory' weakly implies an existing report is prerequisite, but the agent receives no guidance on choosing this over reports_compare, audit_scan, or intelligence_sync.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
audit_scanA
Inspect local source/APK/IPA using cached intelligence and save an evidence report. No network or device mutation.
| Name | Required | Description | Default |
|---|---|---|---|
| sbom | No | ||
| target | Yes | ||
| device_info | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare readOnlyHint=false, destructiveHint=false, and openWorldHint=false, and the description is consistent with them while adding real context: the scan relies on cached intelligence (no network call) and persists an evidence report as a side effect. It does not say where the report is written or whether prior reports are overwritten, but the safety/offline profile is well disclosed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two dense sentences with the capability first and the safety constraint second; every clause carries information (scope, mechanism, output, safety). No filler or restatement of the tool name.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be described, and the offline/non-mutating profile is covered. What remains missing for correct invocation is parameter meaning (sbom, device_info) and where the evidence report lands, which matters because the tool's side effect is the deliverable.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% across three parameters, so the description must compensate and largely does not. 'Local source/APK/IPA' hints that target is an artifact path, but sbom and device_info are left entirely unexplained in both schema and description, including their expected format and why they would be supplied.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Inspect), a specific resource class (local source/APK/IPA), the mechanism (cached intelligence), and the outcome (saved evidence report). It implicitly separates itself from audit_start/audit_reassess by emphasizing 'local' and offline operation, but never names a sibling explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
'No network or device mutation' functions as an implicit boundary condition, steering the agent toward this tool for offline artifact inspection. However, there is no explicit when-to-use vs audit_start or audit_reassess, and no stated prerequisites (e.g., whether intelligence_sync must run first).
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
audit_startA
Start a persistent offline audit; use jobs_status for progress and jobs_cancel to stop it.
| Name | Required | Description | Default |
|---|---|---|---|
| sbom | No | ||
| target | Yes | ||
| device_info | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare readOnlyHint=false, destructiveHint=false, openWorldHint=false, covering the safety profile. The description adds value beyond that by revealing this starts a persistent background job that must be polled rather than returning results inline, a trait absent from the structured fields.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence that states the action first and then the follow-up routing. No filler, no redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Output schema exists, so return values need not be explained, and annotations carry safety. However, with three undocumented parameters at 0% schema coverage, the definition leaves an agent guessing about sbom and device_info, a meaningful gap for a kickoff tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and the description adds no information about any of the three parameters (sbom, target, device_info). With low coverage the description was expected to compensate, and it provides nothing on parameter meaning or format.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Start) and resource (audit) with a differentiating scope marker, 'persistent offline', that separates it from the one-shot audit_scan sibling. It does not open the schema to tell them apart, but the phrasing makes the async/background nature clear.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly routes the agent to jobs_status for progress and jobs_cancel to stop, which is useful lifecycle guidance. It does not address the more important selection question of audit_start vs audit_scan or audit_reassess, so it stops short of full alternative coverage.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
capabilitiesARead-only
Discover input types, evidence states, runtime requirements and available interfaces.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, destructiveHint=false, openWorldHint=false, so the safety profile is covered. The description adds the categories of information returned, but says nothing about auth requirements, idempotency, or whether it needs to be re-invoked, leaving the discovery semantics implicit.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence that lists the four discovery categories with no filler. Nothing is padded or repeated, making it appropriately sized for a zero-parameter tool.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be explained, and annotations cover the safety profile. For a no-param discovery tool the description is nearly sufficient, missing only a hint about when in the workflow it should be invoked.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Zero parameters, so there is no parameter semantics to document; the baseline for a no-param tool is 4. Schema coverage is vacuously complete, and no additional meaning is required from the description.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The verb 'Discover' plus an explicit enumeration (input types, evidence states, runtime requirements, available interfaces) makes it clear this is a self-describing discovery endpoint. However, it never states whose capabilities are being discovered (whole system vs. a subsystem) and gives no differentiation from siblings like runtime_devices or runtime_plan.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No when-to-use guidance is given. A discovery tool would benefit from saying it should be called first/once at startup or as a prerequisite to runtime_plan, audit_start, etc., but nothing routes the agent relative to the 16 siblings.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
dependency_checkB
Query OSV for inventoried package versions, refresh cached matches and save a reassessed report.
| Name | Required | Description | Default |
|---|---|---|---|
| report_id | No | latest |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare readOnlyHint=false, openWorldHint=true and destructiveHint=false; the description usefully adds that it makes external OSV queries, mutates a cache of matches, and persists a reassessed report. It stops short of stating auth requirements or overwrite semantics for the cached matches, so it is helpful but not exhaustive.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence that names the data source and the three effects with no filler. Efficient, though it packs three distinct actions without signaling their order or relative importance.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists so return values need not be described, and annotations cover externality and safety. However, the combination of a mutation tool with a 0%-covered parameter and no usage guidance leaves the description short of what an agent needs to call it confidently.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and the single report_id parameter (default 'latest') is never mentioned in the description. The text refers to saving a report but never explains that report_id selects which inventoried report is checked, leaving the key input ambiguous.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses specific verbs (query, refresh, save) tied to clear resources (OSV package matches, a reassessed report), so an agent understands this checks dependency versions against a vulnerability source. It does not differentiate from siblings like audit_reassess or intelligence_sync, which is the only clear gap.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no statement of when to invoke this versus alternatives such as audit_reassess (which also produces reassessed reports) or intelligence_sync. The agent is left to infer the workflow context entirely.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
intelligence_getA
Read an exact CVE; refresh=true fetches its official CNA record into cache.
| Name | Required | Description | Default |
|---|---|---|---|
| cve_id | Yes | ||
| refresh | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description usefully discloses that refresh=true mutates state by fetching the official CNA record into a cache, which explains the otherwise surprising readOnlyHint=false on a tool whose description starts with "Read". It does not cover error behavior for unknown CVE IDs or rate limits, so it falls short of 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with the primary action first and the optional flag's effect second. No filler, no redundancy, instantly scannable.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be described, and the non-obvious cache side effect of refresh is covered. The remaining gap is the accepted format/validation of cve_id, which an agent should not have to guess.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must carry parameter meaning. It explains refresh (fetches the CNA record into cache) but says nothing about cve_id format (e.g. CVE-YYYY-NNNN) or whether an unknown ID is an error, leaving one of two parameters undocumented.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ("Read an exact CVE") and the word "exact" implicitly distinguishes it from the lookup-by-query sibling intelligence_search. It is clear what the tool does, but it never names the sibling it differs from, so differentiation is inferred rather than stated.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is only implied: "exact CVE" suggests this is the direct-ID path versus intelligence_search, and refresh=true is given a condition, but there is no explicit "use this when / use X instead" guidance and no statement of when to set refresh.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
intelligence_searchBRead-only
Search cached CVEs and OWASP tests with provenance. Does not contact the network.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | ||
| query | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, openWorldHint=false and destructiveHint=false, so the safety profile is covered. 'Does not contact the network' largely restates openWorldHint=false, though 'cached' usefully signals potential staleness of results. No rate limits, freshness windows or result-shape notes are added.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences, with the core action and scope front-loaded and a single behavioral qualifier after it. No filler or repetition.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be explained, and the annotations cover safety. However, for a low-coverage search tool, the absence of any query-format guidance or result-scoping notes leaves a real gap for correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so both parameters carry their full burden in the schema titles alone. The description says nothing about whether 'query' is free text, a CVE ID, a keyword set, or how 'limit' interacts with results, leaving the agent to guess at query syntax for a search tool.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (search) and resource (cached CVEs and OWASP tests), with the 'cached' scope narrowing it versus the sibling intelligence_get and intelligence_sync. It is clear what the tool does, though it never explicitly names a sibling or the boundary between them.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no when-to-use or when-not-to-use guidance and no mention of alternatives among siblings like intelligence_get or intelligence_sync. The agent must infer that this is the free-text search path and that the others are direct-fetch or refresh paths.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
intelligence_syncC
Fetch official public mobile advisories and update local cache. Returns per-source failures and scope.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | ||
| sources | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=false, openWorldHint=true and destructiveHint=false, so the agent knows this reaches external sources and mutates local state without being destructive. The description usefully adds that it updates a local cache and reports per-source failures plus scope, but omits whether the sync is blocking, whether it overwrites existing cached advisories, and any auth or rate-limit constraints.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences with no filler, and the primary action is front-loaded ahead of the return-value note. It is efficient, though 'scope' is left unexplained and reads as an afterthought.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be restated, and the description does gesture at per-source failures. However, for a non-read-only tool with an open-world dependency and two completely undocumented parameters, the definition leaves too much unstated to call it complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and neither of the two parameters is described in the schema. The description never mentions 'limit' or 'sources', so an agent gets no help understanding what limit bounds or what source identifiers are valid — a real gap for a tool whose only inputs are these two fields.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb pair (fetch + update local cache) and a concrete resource (official public mobile advisories), which separates it from read-only siblings like intelligence_get and intelligence_search. It stops short of explicitly naming those siblings as the non-sync alternatives, so it is clear but not fully differentiated.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance on when to run a sync versus calling intelligence_search or intelligence_get directly, no mention of prerequisites, freshness needs, or how often it should be invoked. The agent must infer the trigger condition entirely from the name.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
jobs_cancelA
Request cancellation of an authorized job; poll jobs_status until terminal.
| Name | Required | Description | Default |
|---|---|---|---|
| job_id | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare readOnlyHint=false and destructiveHint=false but say nothing about timing. The description adds that this is a *request* (not synchronous) and that the job reaches a terminal state observable via jobs_status — genuinely useful context beyond the annotations. It stops short of noting reversibility or permission requirements.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One sentence, front-loaded with the action and followed by the essential next step. No wasted words.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Output schema exists, so return values need not be explained, and the async workflow is conveyed. For a mutation tool with no explicit permission or reversibility notes, a small gap remains.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0% for the single job_id parameter, so the schema gives no semantic help. The description only implies through 'authorized job' that the id must belong to a job the caller is permitted to cancel, adding minimal meaning.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Request cancellation of an authorized job') and implicitly contrasts with the polling tool it names, jobs_status. It doesn't distinguish itself from other job-related siblings like jobs_list, but the action is unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides clear follow-up guidance: 'poll jobs_status until terminal,' which tells the agent this is asynchronous and how to complete the workflow. It does not state when cancellation is inappropriate (e.g., already-terminal jobs) or name any exclusion.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
jobs_listC
List background jobs visible inside this server's configured roots.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description presents a pure enumeration operation ('List background jobs'), while the annotations declare readOnlyHint=false, meaning the tool is not guaranteed to be side-effect free. The description offers no explanation of what non-read-only behavior occurs, and no pagination or ordering behavior is disclosed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One short, front-loaded sentence with no filler, so it reads cleanly. It is terse to the point of under-specification, but there is no wasted or redundant text to penalize.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With an output schema present, return values need not be explained, and the tool's complexity is low (one optional parameter). However, the limit/pagination semantics and the tension with the non-read-only annotation leave real gaps for correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and the single 'limit' parameter (default 20) is never mentioned in the description, so its meaning, range, and pagination implications are undocumented anywhere. The description does not compensate for the coverage gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (List) and resource (background jobs) plus a scoping qualifier ('visible inside this server's configured roots'), which is more than a tautology. It does not explicitly distinguish itself from siblings jobs_status and jobs_cancel, but the verb choice makes the read/enumerate role inferable.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'inside this server's configured roots' implies the context in which listing makes sense, but there is no explicit when-to-use guidance and no mention of the sibling alternatives jobs_status or jobs_cancel for inspecting or cancelling a specific job.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
jobs_statusC
Read durable job state and progress. An interrupted worker never becomes a successful audit.
| Name | Required | Description | Default |
|---|---|---|---|
| job_id | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description asserts a pure read operation ("Read durable job state and progress"), but the annotations declare readOnlyHint=false, meaning the tool may modify state. That is a direct contradiction of the behavioral claim in the text. The remaining sentence adds no concrete behavior such as polling cadence, terminal state semantics, or error conditions.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences, front-loaded with the action, which is good. The second sentence is atmospheric and ambiguous rather than informative, so it does not clearly earn its place in a definition this terse.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be explained, but for a job-polling tool the description omits the essential context: how the job_id is obtained, how to interpret completion vs. failure, and how this differs from jobs_list. Combined with the annotation contradiction, this leaves the agent under-informed.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% for the single job_id parameter, and the description says nothing about its format, or where the id comes from (presumably audit_start or jobs_list). The name is largely self-explanatory, which prevents a 1, but the description does nothing to compensate for the missing schema documentation.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
"Read durable job state and progress" states a specific verb (Read) and resource (job state/progress), so the core purpose is unambiguous. It does not, however, distinguish itself from sibling job tools such as jobs_list or jobs_cancel, leaving the agent to infer the boundary from the name alone.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no explicit when-to-use, when-not-to-use, or alternative guidance. The agent must infer that this is called after audit_start to poll a job, and the cryptic line "an interrupted worker never becomes a successful audit" reads as an aphorism rather than a usage rule. None of jobs_list, jobs_cancel, or audit_scan are referenced as alternatives.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
policy_evaluateCRead-only
Evaluate team thresholds, explicit coverage requirements and expiring waivers without changing reports.
| Name | Required | Description | Default |
|---|---|---|---|
| report_id | No | latest | |
| baseline_id | No | ||
| policy_path | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, destructiveHint=false and openWorldHint=false, so 'without changing reports' is largely redundant with structured data. The description does add domain context by naming the three classes of policy checks performed (thresholds, coverage requirements, expiring waivers), which goes slightly beyond the annotations, but says nothing about required inputs, permission needs, or failure behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with no filler or repetition, and the non-mutating constraint is placed at the end where it reads naturally. It is efficient, though the packed enumeration of policy check types makes the sentence dense without adding usable operational detail.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists so return values need not be explained, but the tool has three parameters at 0% schema coverage and no usage guidance, so an agent cannot reliably construct a call — particularly the required policy_path. The description is too thin for a parameterized evaluation tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% for three parameters, so the description carries the full burden and fails to meet it. It never mentions policy_path (the sole required parameter), report_id, or baseline_id, leaving format, expected values, and the meaning of 'latest'/null defaults entirely undocumented.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Names a specific verb (evaluate) and enumerates the policy objects being checked (team thresholds, explicit coverage requirements, expiring waivers), which tells an agent this is a policy-compliance evaluation rather than a report mutation. However, it never distinguishes itself from plausible siblings like reports_compare or audit_scan, so the agent must infer where it fits.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no explicit when-to-use or when-not-to-use guidance, and no alternative tool is named. The only directional signal is the trailing phrase 'without changing reports', which implies a read-only, non-mutating context but does not say when this should be preferred over reports_compare or dependency_check.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
reports_compareBRead-only
Compare exact report IDs, retaining coverage differences and unproven remediation states.
| Name | Required | Description | Default |
|---|---|---|---|
| after | Yes | ||
| before | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, openWorldHint=false and destructiveHint=false, so the safety profile is covered. The description adds genuinely useful behavioral context about the comparison semantics ('retaining coverage differences and unproven remediation states'), which tells the agent what the diff preserves rather than normalizes away. It remains terse and jargon-heavy, and says nothing about cost or ordering of the two inputs.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with no filler; the core action leads and the retention clause follows. The dense phrasing ('unproven remediation states') is compact but requires domain familiarity.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be explained, and the safety profile is covered by annotations. Still, for a two-required-parameter comparison tool with zero schema coverage, the definition omits where the IDs originate and how the two inputs are ordered, leaving a real gap an agent would hit.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate, and it partially does by establishing that both parameters are report IDs ('exact report IDs'). However, it does not clarify the before/after ordering semantics, ID format, or whether the IDs must come from the same target, leaving the bare 'Before'/'After' schema titles to carry the rest.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Compare) and resource (report IDs), and 'exact report IDs' distinguishes it from the filter-based reports_list / reports_get siblings. It is clear what the tool does, though it never names a sibling explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no when-to-use guidance, no prerequisite stating where the two IDs come from, and no mention of alternatives such as reports_get for inspecting a single report. The agent must infer that this tool is selected whenever two known report IDs need a diff.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
reports_getBRead-only
Read grounded report context, or the exact evidence for one finding ID.
| Name | Required | Description | Default |
|---|---|---|---|
| report_id | No | latest | |
| finding_id | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, destructiveHint=false and openWorldHint=false, so the safety profile is covered. The description adds only the mode switch (context vs. finding evidence) on top of that; it discloses nothing about defaults, error behavior, or scoping that isn't in structured fields.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with no filler; the two retrieval modes are stated inline rather than buried. It is efficient, though arguably terse to the point of under-specification.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be described, and read-only annotations cover safety. What is missing is the meaning of the 'latest' report_id default and whether finding_id can be used without a report_id — usable, but not complete for a 2-param tool at 0% schema coverage.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must carry the load. It meaningfully ties finding_id to 'the exact evidence for one finding ID', but says nothing about report_id, its 'latest' default, or the null default for finding_id — leaving half the parameters unexplained outside the bare schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Read') and resource ('report context' / 'evidence'), and distinguishes two retrieval modes (whole report vs. one finding's evidence). It does not, however, differentiate itself from siblings like reports_list or reports_compare, which a 'get' tool sharing that namespace arguably needs.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied rather than stated: supplying a finding ID selects the evidence mode, omitting it reads report context. There is no explicit when/when-not guidance and no mention of the sibling tools an agent should consider instead (reports_list, reports_compare).
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
reports_listBRead-only
List saved audit IDs for exact follow-up reads.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, destructiveHint=false, and openWorldHint=false, so the safety profile is covered by structured data. The description adds that the output is 'saved audit IDs', which is useful context, but says nothing about pagination behavior or result set size despite the limit parameter.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with no filler, which is efficient. It may be slightly too terse to carry the tool's scope, but every word earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return-value detail is not required, and the safety profile is covered by annotations. However, for a list tool the description omits pagination/limit behavior and gives no relationship to reports_get or reports_compare, leaving gaps an agent would want filled.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and the description never mentions the 'limit' parameter or its default of 20. With one undocumented optional parameter, the description fails to compensate for the schema gap, leaving pagination semantics entirely unstated.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('List') and resource ('saved audit IDs') and hints at its role as a lookup step ('for exact follow-up reads'), which faintly distinguishes it from reports_get/reports_compare. It does not explicitly name a sibling or clarify the 'audit IDs vs reports' naming mismatch, so it falls short of full differentiation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'for exact follow-up reads' implies the workflow (list IDs, then read them individually), but there is no explicit when-to-use statement, no when-not, and no named alternative such as reports_get. Usage is only inferable.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
runtime_devicesARead-only
List adb devices and available iOS simulators without changing app state.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description adds the behavioral note 'without changing app state,' which reinforces the read-only nature but is somewhat redundant with the readOnlyHint annotation. It does not disclose other traits like output format details, though the annotations already cover safety and open-world aspects.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, efficient sentence with no unnecessary words, front-loading the core purpose clearly.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's simplicity (no parameters, read-only annotations, and an output schema), the description adequately conveys the core action. It could mention when to use it, but for a listing tool, the essential information is present.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has no parameters, so the baseline score is 4. The description does not need to explain any parameters, and the schema coverage is 100% for the empty object.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (List) and resources (adb devices and iOS simulators), making the tool's function unambiguous and easily distinguishable from siblings like runtime_plan.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Offers no guidance on when to use this tool versus alternatives, nor any prerequisites or context. It simply describes what it does without indicating appropriate scenarios or exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
runtime_planCRead-only
Generate an editable test scenario. Does not run it or log out the app automatically.
| Name | Required | Description | Default |
|---|---|---|---|
| package | No | ||
| platform | No | ||
| report_id | No | latest |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, destructiveHint=false, and openWorldHint=false. The description adds useful behavioral context by clarifying that the generated scenario is editable and will not be executed or trigger automatic logout. It does not cover permissions, persistence, or other operational details, but that is a modest gap given the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is very short and front-loads the core action in the first sentence, with the caveat immediately after. It is efficient with no filler, though its brevity means it omits important parameter guidance.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be explained here, and annotations cover the safety profile. Still, for a tool with three optional inputs, no schema descriptions, and ambiguous scope, the description should say more about what the scenario is based on and how package/platform/report_id shape it.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, and the description does not mention any of the three parameters (package, platform, report_id). It adds no meaning about what these inputs select or how they affect the generated plan, leaving the agent with only bare parameter names.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description gives a verb and object: 'Generate an editable test scenario.' However, 'test scenario' is vague without saying what domain or system it applies to, and no sibling tool is named or contrasted, so an agent must infer the tool's precise role among runtime and audit siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The second sentence provides an important usage boundary: it does not run the scenario or log out the app automatically. That implies the tool is for planning only, but it does not state when to choose it over alternatives such as audit_start or runtime_devices, nor does it describe prerequisites.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
17 tool updates
v1.0.3- First observed
audit_reassess - First observed
audit_scan - First observed
audit_start - First observed
capabilities - First observed
dependency_check - First observed
intelligence_get - First observed
intelligence_search - First observed
intelligence_sync - First observed
jobs_cancel - First observed
jobs_list - First observed
jobs_status - First observed
policy_evaluate - First observed
reports_compare - First observed
reports_get - First observed
reports_list - First observed
runtime_devices - First observed
runtime_plan
TDQS
Scored across 17 tools
Most tools target clearly distinct resources and actions, and descriptions clarify boundaries well. The main risk is mild overlap among audit_scan, audit_start, and audit_reassess, plus reports_get/list/compare, but each has a defined role.
Tool names consistently use snake_case and domain prefixes such as reports_, jobs_, intelligence_, and audit_. A few names are noun-oriented rather than verb_noun, but the overall convention is predictable.
At 17 tools the server is slightly above the ideal 3-15 range, but the surface covers several legitimate areas: auditing, intelligence, jobs, reports, runtime, and policy. Each tool appears to earn its place without obvious redundancy.
The set covers core audit lifecycle operations, cached intelligence access, report reading/comparison, background job control, dependency checks, and runtime planning. Minor gaps remain, such as policy CRUD or report deletion/export, but these are not fatal for the apparent domain.
Maintenance
Related MCP Connectors
Remote MCP for Android CLI agent build gate, structured receipts, audit logs, and reviewer-ready evi
MCP server for static security analysis of Android source code
Create App Store screenshots, icons, ASO copy, localization, and revisions via hosted MCP.
Enrich, search, assess, and manage threat intelligence through 80+ typed MCP tools.
Related MCP Servers
- AlicenseNot gradedqualityDmaintenanceEnables authorized Android security testing with static and dynamic analysis, Frida instrumentation, storage inspection, and traffic interception via MCP tools.1MIT
- AlicenseNot gradedqualityBmaintenanceEnables security scanning of source code and Apple configuration files, providing evidence-based findings, explanations, remediation plans, threat modeling, and knowledge base access through MCP tools.MIT
- FlicenseBqualityCmaintenanceEnables local-first static and optional controlled dynamic analysis of Android APK artifacts, with provenance-aware downloads, Markdown reports, and structured security findings via MCP stdio tools.9-
- AlicenseAqualityAmaintenanceEnables MCP clients to run deterministic AST safety audits, secret scanning, and sandboxed test execution on codebases, with human approval required before any remediation action.5MIT