Skip to main content
Glama
luiacuaniello

PerspectiveGraph

PerspectiveGraph joins what you already run - Trivy, Semgrep, Cloud Custodian, Falco, plus your AWS and Kubernetes state - into one graph of your real environment, and asks a single question of it: can someone get from the internet, through privilege that is too broad, to something worth stealing?

On a pull request it asks that question before the merge: the check goes red only when this change opens a route, and the fix comes back as its own pull request. Open source (Apache 2.0), runs on your infrastructure, collects no telemetry.

PerspectiveGraph: from the day's exploitable routes to a generated fix

Twelve seconds of make demo: what is exploitable now → the ranked routes → one route's kill chain and the fix it generates → whether the scores can be trusted. Sample scanner output and seeded verdicts, not a real environment.

A score here is what the model concludes from the evidence it was given, not a measured frequency: nothing has been calibrated against field data yet, and the engine says so itself rather than rounding up. What is measured, and what is not.

Check your own account in 30 seconds

No deployment, no Docker, nothing ingested. One static binary asks AWS's own policy evaluator which of your roles can reach administrator - applying the service control policies, permission boundaries and condition keys that a policy reader on its own does not see:

# macOS (Apple silicon); swap darwin_arm64 for linux_amd64, linux_arm64 or darwin_amd64
curl -sSL https://github.com/luiacuaniello/perspectivegraph/releases/latest/download/perspectivegraph_darwin_arm64.tar.gz | tar xz
./perspectivegraph redteam -roles -region eu-west-1

It is read-only and free: each check is one iam:SimulatePrincipalPolicy dry run, which creates nothing and costs nothing, and the only permissions it needs are that and iam:ListRoles, both inside SecurityAudit. Windows builds are on the releases page; every binary is signed with cosign and carries SLSA provenance, and two commands verify both before you run anything.

Add -compare and it also runs the engine over the same account, exiting non-zero where the two disagree. That is how the engine's first real false positive was found, and how it stays fixed. If it disagrees on yours, report it: no report is more useful. From here, how to evaluate this walks to a verdict on your own estate in stages that each end in an answer.

Related MCP server: cloud-pathfinder

See the whole engine in 90 seconds

make demo

Pulls the published, cosign-signed images, feeds them sample Trivy / Semgrep / Custodian / Falco / Kubernetes / IAM / SSO output, and prints the top attack path with its generated fix. Dashboard on http://localhost:3000, make down to tear it down. It needs Docker, jq and curl, compiles nothing, and takes about 23 seconds from an empty image cache; make demo-build is the same demo built from your working tree.

On Kubernetes, the chart is an official package on Artifact Hub:

helm install perspectivegraph oci://ghcr.io/luiacuaniello/charts/perspectivegraph \
  --version 1.27.0 # x-release-please-version

The chart and the three images it runs (ghcr.io/luiacuaniello/perspectivegraph, -dashboard and -postgres) are signed with cosign keyless and carry an SPDX SBOM and SLSA provenance: verify them rather than taking the supply chain on trust.

The dashboard opens on the decision, not the inventory: what is being exploited now, the fewest changes that remove the most risk, and how far the numbers can be trusted. Routes are ranked by triage priority - what the route reaches, whether runtime confirmed it, how exposed the entry is - so a lower-scoring route can outrank a higher-scoring one.

The day's decision surface

Attack path detail

Score accuracy

Why the route is P1, one probability with its range, and every hop with what it lets the attacker do and where its probability came from.

Whether the engine's own scores held up against recorded outcomes - and whether there are enough of them to say.

The screenshots are make demo with seeded verdicts: make seed-validation records synthetic outcomes on 8 of the 14 routes, the way a BAS run tests part of an estate, which is why the calibration panel has something to show - and why it still says not enough outcomes yet: 8 is far short of the 30 a verdict needs (leaning "calibrated on average": the engine predicted 58% where 50% held up). A fresh install and the public demo report insufficient data instead, until real outcomes exist. The public demo runs on one free VM, so treat it as best-effort.

Why?

A scanner reports that a container carries a critical CVE. It cannot report that the container sits behind an internet-facing load balancer, runs with a role that reads the production database, and is therefore the one finding out of ten thousand worth fixing this week. That needs the other tools' output in the same graph, which is what this builds. A developer gets a check that goes red only when their change opens a real route; a security team gets a short ranked list of attack paths instead of a flat list of findings.

Block the pull request that opens the path

No deployment required. The runner reads your estate read-only, ingests this pull request's scan, and answers in-process with the same engine:

- uses: luiacuaniello/perspectivegraph@v1
  with:
    mode: local
    aws-region: eu-west-1     # read-only; give the job an OIDC role with SecurityAudit
    report: trivy.json

The check goes red when this change opens a route to a sensitive asset - or makes one likelier - not when it adds a critical CVE: a critical on a host nothing routes to does not fail the build, and a medium on a container that now reaches the production database does. Routes that were there before the change do not count, even when they run through what it touches: the scan is applied to a copy of the estate and compared with it, and nothing is written. It also has a third outcome, because a pipeline whose scan never arrived must not get the same green tick as one that is clean:

Verdict

Exit

Meaning

clean

0

The engine analysed this change: it opens or worsens no critical path

blocked

1

It opens or worsens critical attack paths - the check names them

unknown

2

Nobody analysed it. The scan, the ingest or the SHA is wrong

Outside GitHub Actions it is one command, and it installs as a Trivy plugin too:

perspectivegraph gate -local -aws-region eu-west-1 -report trivy.json -slug owner/name -sha "$COMMIT_SHA"

trivy plugin install github.com/luiacuaniello/perspectivegraph
trivy image -f json myapp:pr-42 | trivy perspectivegraph gate -local -aws-region eu-west-1 -report -
WARNING

Fork pull requests get no secrets, so the gate fails closed as unknown on them. Do not work around it with pull_request_target: that runs your secrets - in local mode, cloud credentials - against the contributor's code.

On a public repository, a blocked check prints the route (real asset names, the CVE, the sensitive asset) into a public job log. Use soft-fail and post the detail somewhere private.

The manual covers the rest: pointing the action at a deployed engine, gating rendered manifests, a pre-collected estate, rolling the gate out, and every input in action.yml.

Let an agent query it

A language model cannot enumerate thousands of edges reliably or run Dijkstra, and asked for "the attack paths in my account" it will invent plausible ones. So the engine speaks MCP: the agent asks, and reasons over answers it could not have made up.

perspectivegraph mcp --api http://localhost:8080   # or https://demo.a3thinker.it, no credential needed

Eight tools, every one read-only and declared so on the wire. The one worth the integration is simulate_fix: it re-runs the simulation with the given edges cut and reports what actually changes. The server is in the official MCP Registry and on Glama; the tools and the client configuration are in the manual.

PerspectiveGraph MCP server on Glama

Project status & maturity

OpenSSF Best Practices Artifact Hub Go

The engine and its public API are complete, documented and tested. The scores are not yet calibrated. Nobody has run this over a real estate, tested the paths it surfaced and fed the verdicts back: the machinery for that loop is built and tested, the loop is not closed. So read a score as what the model believes and how sure it says it is, not as a measured frequency. Use it to find and cut routes; don't put its percentage in front of a board. Positioning spells out what is and isn't claimed.

What is measured today, as of v1.27.0. make bench-cloudgoat grades the engine in CI on four CloudGoat-shaped scenarios:

Scenario

Expects

Result

ec2_ssrf

a path

found it, invented none

iam_privesc_by_attachment

a path (leaked-credential origin)

found it, invented none

ec2_private_subnet_no_path

no path (open SG, private subnet)

produced none

iam_privesc_denied_by_guardrail

no path (explicit Deny wins)

produced none

Precision and recall are 1.00 on all four: four known shapes, two of them negative controls, so a regression gate rather than a measure of accuracy on your estate. On real AWS, make reachability-lab-aws checks exposure the same way for free, and make redteam-aws grades the engine's escalation claims against AWS's own policy evaluator. That grading has already caught one false positive, a permissions boundary the engine ignored; it is fixed, and make boundary-lab-aws fails whenever the engine and AWS disagree.

  • Clouds. AWS is live and verified against a real account, cross-account AssumeRole included. Azure is fixtures only, and there is no GCP connector.

  • Interface. The GraphQL schema is frozen and drift-guarded, so a breaking change comes with a major version, never in a patch: API stability.

  • Deployment. The backend is distroless and non-root on a read-only root filesystem, the images are pinned by digest, and under Helm every workload meets the restricted Pod Security Standard, asserted in CI. The compose defaults are open on purpose; PG_ENV=production refuses to start unless API and ingest are authenticated. A real rollout needs more (external PostgreSQL with AGE, TLS, backups, TRUSTED_PROXY_CIDRS behind a proxy): the operations runbook lists it.

  • Support. The newest release only, with a clock on security fixes (Critical 7 days, High 30): SUPPORT.md.

  • Telemetry. None. Out of the box it opens no outbound connection; GitHub, the AI assistant and the KEV/EPSS feeds each stay off until you configure them.

  • Scope. The reachable attack-path question, in the developer workflow. It is not a scanner, a CNAPP or a compliance product, and replaces none of them.

  • How it is written. By a human working with Claude (Anthropic): the design decisions and what ships are the maintainer's, and much of the implementation and its tests came out of that collaboration. Check it rather than trust it: make test, make bench-cloudgoat, govulncheck ./..., and CONTRIBUTING.

Documentation

License

Apache License 2.0.

Available Tools

8 tools
explain_attack_pathA
Read-onlyIdempotent

Give the full kill chain for one route: every hop, the relationship type, that hop's probability, where the probability came from (kev/epss/runtime are observed evidence; cvss/severity/heuristic are estimates), and the MITRE ATT&CK technique. Use this before explaining or acting on a route - the hop provenance is what tells you which parts of the story are evidence and which are assumption.

ParametersJSON Schema
NameRequiredDescriptionDefault
path_idYesThe id from list_attack_paths.

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, so the safety profile is covered. The description adds valuable behavioral context: it reveals that the tool distinguishes observed evidence (kev/epss/runtime) from estimates (cvss/severity/heuristic), which is a key behavioral trait an agent needs to interpret results correctly. It doesn't describe output format, but with no output schema and read-only semantics, the provenance explanation is the most important behavioral disclosure.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two sentences, front-loaded with the core purpose and the full list of what's included, followed by a usage directive. Every sentence earns its place; no filler or repetition of schema details.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a read-only, single-parameter tool, the description is nearly complete. It explains what the tool returns, why to use it, and how to interpret the provenance field. The only minor gap is that it doesn't explicitly state the output format or whether the response includes remediation steps, but the description's focus on the kill chain and provenance is sufficient for an agent to select and invoke it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% and the single parameter path_id is described as 'The id from list_attack_paths.' The description doesn't add much beyond that, but the parameter is simple and self-explanatory. Baseline 3 is appropriate since the schema fully documents the parameter.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('explain') and resource ('attack path'), and enumerates exactly what the full kill chain includes: every hop, relationship type, probability, provenance, and MITRE technique. It clearly distinguishes itself from sibling tools like list_attack_paths by focusing on a single route's detailed breakdown.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly says 'Use this before explaining or acting on a route' and explains why: the hop provenance tells you which parts are evidence and which are assumption. This gives clear when-to-use guidance and implicitly contrasts with list_attack_paths (which lists paths) and routes_to_target (which finds routes).

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_postureA
Read-onlyIdempotent

Summarize the environment: how many attack routes are open, how many are runtime-confirmed, how many assets and relationships are mapped, and which sensitive assets an attacker can currently reach. Start here to orient before enumerating anything.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, so the safety profile is covered. The description adds value by specifying what the summary actually surfaces – counts and reachable sensitive assets – which is behavioral information beyond the annotations. It does not contradict any annotation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is exactly two sentences, front-loaded with the core action and then listing the specifics. Every word earns its place – no filler, no repetition. This is a model of concise, structured documentation.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple, parameterless tool with no output schema, the description is complete. It explains exactly what information the summary provides (attack route counts, asset and relationship counts, reachable sensitive assets) and when to use it. An agent would need nothing else to call this tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has zero parameters, so the description has no parameter semantics to explain. Per the rubric, a zero-parameter tool gets a baseline of 4 because the schema (empty) fully covers it. The description does not need to add anything here.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('Summarize') and a clear resource ('the environment'), and enumerates exactly what the summary includes: counts of open and runtime-confirmed attack routes, mapped assets and relationships, and reachable sensitive assets. This clearly distinguishes it from sibling tools like list_attack_paths or explain_attack_path, which are focused on specific enumeration or explanation rather than an overall posture snapshot.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The phrase 'Start here to orient before enumerating anything' explicitly tells the agent when to invoke this tool (first, before any enumeration). This provides unambiguous usage context, even though it doesn't name alternatives – the siblings are implicitly the enumeration tools that follow. This is clear guidance on sequencing.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_score_trustA
Read-onlyIdempotent

Report how well the engine's probabilities have matched reality, measured against recorded red-team or BAS outcomes: the verdict (well-calibrated / calibrated-on-average / overconfident / underconfident / insufficient-data; calibrated-on-average means only the average matches, so no individual score may be quoted as a probability), the predicted-versus-observed rates, and what to do about the gap. Call this before quoting any score as a probability. If it reports insufficient-data, the numbers are expert estimates and must be presented as the model's own estimate, not as odds. It also reports discrimination - whether the score, and separately the triage Priority order, put confirmed paths above refuted ones (AUC with a 95% interval). Do not present the ranking as evidence of which path is most dangerous unless priorityDiscrimination reads 'discriminates'; 'insufficient-data' or 'indistinguishable-from-chance' means the order has not been shown to beat a coin, and must be said so.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already mark this as read-only and non-destructive, and the description goes well beyond them by explaining the calibration verdicts, the difference between overall average calibration and per-score probability claims, the AUC-based discrimination output, and the interpretation constraints on the Priority order.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is information-dense and front-loaded with the core purpose, but it is a single long run-on paragraph with many embedded caveats. Structuring the verdicts, rates, and discrimination guidance into separate sentences or bullets would improve scannability.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema and no parameters, the description carries full responsibility for explaining the return shape. It covers the verdict values, the predicted-versus-observed rates, the discrimination AUC and interval, and the required downstream presentation caveats. Nothing essential is missing for an agent to invoke and interpret the tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so the schema has nothing to document. The description uses the no-parameter baseline appropriately by focusing entirely on what the report contains and how to interpret it.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('report') and resource ('how well the engine's probabilities have matched reality'), and enumerates the concrete outputs: verdict, predicted-versus-observed rates, and discrimination. It is clearly distinguishable from siblings like get_posture or list_attack_paths.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives an explicit when-to-use instruction: 'Call this before quoting any score as a probability.' It also provides when-not guidance about not presenting rankings as evidence unless priorityDiscrimination says 'discriminates', and how to handle insufficient-data results.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_attack_pathsA
Read-onlyIdempotent

List the ranked routes from internet exposure to a sensitive asset, highest triage priority first. score is the modelled end-to-end exploit probability and priority (0-100, banded P1/P2/P3) is the triage order that also weighs corroboration and target sensitivity. IMPORTANT: these probabilities are expert estimates, not measurements calibrated against real outcomes - call get_score_trust before presenting any of them as a probability, and prefer the ranking over the absolute values.

ParametersJSON Schema
NameRequiredDescriptionDefault
appNoOptional application scope; omit for the whole environment.
limitNoHow many routes to return, priority-first.

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnly, idempotent, and non-destructive behavior. The description goes beyond annotations by warning that probabilities are expert estimates, not calibrated measurements, and by defining the meaning of score and priority.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is compact: it states the purpose, defines the key output concepts, and adds a clearly separated important caveat. Every sentence contributes useful information without redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a list tool without an output schema, it explains the ranking, meaning of score, meaning of priority, and the trust caveat. It does not fully enumerate all possible route fields, but the essential behavior and call context are sufficiently covered.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema already describes both parameters fully, so the description does not need to repeat them. It also does not add parameter-specific semantics beyond the schema's own descriptions, so the baseline score of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific verb ('List') and resource ('ranked routes from internet exposure to a sensitive asset'), and states the ordering principle ('highest triage priority first'). It is clearly distinguishable from siblings like explain_attack_path or get_score_trust.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives explicit operational guidance: call get_score_trust before presenting probabilities and prefer the ranking over absolute values. It does not explicitly contrast this tool with sibling alternatives such as routes_to_target or explain_attack_path, but the usage context is clear enough.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_fixesA
Read-onlyIdempotent

Return the remediation plan: the fewest changes that remove the most risk, choke points first, each with the share of critical-path risk it eliminates and how many routes it cuts. This is usually the right answer to 'what should we do' - a hundred routes typically collapse into a handful of changes.

ParametersJSON Schema
NameRequiredDescriptionDefault
appNoOptional application scope; omit for the whole environment.

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, and destructiveHint=false, so the safety profile is covered. The description adds behavioral context beyond annotations: output includes per-change share of critical-path risk eliminated and number of routes cut, plus prioritization by choke points.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is compact: the first sentence front-loads the core purpose and output details, and the second provides practical guidance without redundancy. Every sentence earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple read-only tool with one optional parameter and no output schema, the description sufficiently explains what the tool returns and how it behaves. No critical calling information is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% for the single optional 'app' parameter, so the schema already explains its meaning. The description adds no parameter-specific detail, but the baseline of 3 is appropriate because the schema fully handles parameter semantics.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('Return') and resource ('remediation plan'), and clearly defines the output as the fewest changes removing the most risk, ordered by choke points. It is conceptually distinct from siblings like list_attack_paths and explain_attack_path, which focus on attack paths rather than prioritized remediation actions.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives clear usage context: it is 'usually the right answer to what should we do', implying it is the go-to tool for deciding on remediation actions. It does not explicitly mention when not to use it or name alternative tools, but the use case is clearly conveyed.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

routes_to_targetA
Read-onlyIdempotent

Enumerate the k best distinct routes that reach a named sensitive asset. Answers 'how many ways in are there, and do they share a choke point' - which single-path views hide. Cutting a hop that every route traverses removes them all; cutting one that appears in a single route removes one.

ParametersJSON Schema
NameRequiredDescriptionDefault
kNoHow many distinct routes to return, best first. Fewer routes come back when fewer exist.
fromNoOptional entry-point name to start from.
targetYesSensitive asset name, e.g. 'account-admin (effective)'.

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, covering the safety profile. The description adds meaningful behavioral context beyond annotations by explaining the choke point concept and the effect of cutting hops, which informs the agent's interpretation of results. It does not contradict annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two sentences with no filler. The core purpose is front-loaded, and the second sentence provides the analytical value proposition. Every word contributes to understanding the tool's function and purpose.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the moderate complexity (3 parameters, no output schema) and the annotations covering safety, the description is fairly complete. It explains the tool's purpose and the meaning of its output (routes, choke points) without explicitly detailing return structure, which is acceptable given the tool's conceptual nature and the absence of an output schema.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema fully documents each parameter (k, from, target). The tool description references 'k best distinct routes' and 'named sensitive asset', which aligns with the schema but does not add additional semantic detail beyond what the schema already provides. Baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a precise action: 'Enumerate the k best distinct routes that reach a named sensitive asset.' It clearly identifies the resource (routes to a target) and the specific value it provides (answering 'how many ways in' and choke point analysis), distinguishing it from single-path views.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly contrasts with 'single-path views' and explains when this tool is valuable (multi-route choke point analysis), implying that single-path tools (like list_attack_paths or explain_attack_path) are for different use cases. However, it does not name alternatives or provide explicit exclusion conditions, so it stops short of full guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

search_assetsA
Read-onlyIdempotent

Full-text search across indexed assets and findings by name, CVE id, or keyword. Use it to resolve a name a human mentioned into the node ids the other tools take.

ParametersJSON Schema
NameRequiredDescriptionDefault
sizeNoMaximum number of matches to return, best match first.
queryYesText matched against asset and finding names, labels, ids, severities and CWEs, e.g. 'log4j', 'PII', 'CVE-2021-44228'.

TDQS

A4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, and destructiveHint, covering the safety profile. The description adds context about searching 'indexed assets and findings' and returning node ids, but it does not describe the exact response shape or empty-match behavior. This is adequate but not rich.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences with no filler. The core operation and matching scope are front-loaded in the first sentence, and the practical usage guidance is in the second. Every clause earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a read-only two-parameter search tool, the description covers scope, purpose, and how results should be consumed ('node ids the other tools take'). With no output schema, slightly more detail about the result structure could help, but the description provides enough for correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the query and size parameters are already well documented with examples and constraints. The description's phrase 'by name, CVE id, or keyword' is a slight restatement of the schema's field list rather than new semantic information. Baseline 3 applies since the schema does the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: 'Full-text search across indexed assets and findings by name, CVE id, or keyword.' It clearly differentiates this tool from the attack-path and posture siblings by emphasizing the lookup/translation role, and the second sentence reinforces the specific job of resolving human-mentioned names into node ids.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly tells an agent when to use this tool: 'Use it to resolve a name a human mentioned into the node ids the other tools take.' It gives clear usage context but does not name alternatives or state when-not-to-use conditions, though no search sibling exists to exclude.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

simulate_fixA
Read-onlyIdempotent

Ask what actually happens if specific relationships are cut: re-runs the whole simulation with those edges removed and reports how many routes disappear and how the compromise probability moves. This is the tool to reach for before recommending a change - it is a deterministic counterfactual over the real graph, not an estimate, so it settles 'would this help' instead of arguing about it.

ParametersJSON Schema
NameRequiredDescriptionDefault
cutsYesRelationships to remove. Use node ids from explain_attack_path steps (from/to).

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, and destructiveHint, so the safety profile is covered. The description adds meaningful behavioral detail beyond annotations: it re-runs the whole simulation, is deterministic, and reports both route count changes and probability movement. No contradiction with annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences with no filler. The core behavior is front-loaded, the output is specified, and the usage guidance is integrated naturally without redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description covers what the tool does, what it reports, and when to use it. There is no output schema, but the description explains the key outputs. It could mention cost or performance implications of re-running the full simulation, but that is a minor gap given the annotations already cover safety.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3, but the description adds value by explaining that cuts are relationships removed and that the simulation runs with those edges removed. It also reinforces the schema's instruction to use node ids from explain_attack_path steps, making parameter usage clearer.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: 'Ask what actually happens if specific relationships are cut' and describes the concrete outputs (routes disappearing, compromise probability moving). It clearly differentiates itself from siblings by emphasizing this is a deterministic counterfactual over the real graph, not an estimate.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly says 'This is the tool to reach for before recommending a change' and frames it as settling 'would this help' instead of arguing. It does not name a specific sibling alternative, but the deterministic-versus-estimate contrast gives usable context for choosing it.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

TDQS

A4.2/5.0

Scored across 8 tools

Disambiguation4/5

Most tools have clearly distinct purposes: explain_attack_path drills into one route, list_attack_paths ranks all routes, routes_to_target enumerates routes to a specific asset, and simulate_fix tests changes. The only potential confusion is between list_attack_paths and routes_to_target, but their descriptions clarify one is a global ranked list while the other is target-specific choke-point analysis.

Naming Consistency4/5

The dominant pattern is verb_noun (explain_attack_path, get_posture, list_attack_paths, search_assets, simulate_fix). One outlier, routes_to_target, is a noun phrase rather than verb_noun, and there is minor singular/plural variation (attack_path vs attack_paths, fix vs fixes), but overall the naming is predictable and readable.

Tool Count5/5

Eight tools is well-scoped for a security analysis server. Each tool earns its place: orientation, enumeration, deep-dive, trust calibration, remediation planning, and counterfactual simulation. No tool feels redundant or missing to make the set bloated or thin.

Completeness4/5

The tool surface covers the core lifecycle: orient (get_posture, search_assets), enumerate (list_attack_paths, routes_to_target), explain (explain_attack_path), validate score trust (get_score_trust), and remediate (list_fixes, simulate_fix). A minor gap is the lack of a dedicated asset-detail or relationship-detail tool, but search_assets and explain_attack_path compensate sufficiently for common workflows.

Maintenance

ActivityActive
ResponsivenessSlow

Related MCP Connectors

Related MCP Servers

  • A
    license
    A
    quality
    C
    maintenance
    A local-first AWS security tool that uses graph theory to discover attack paths (e.g., Internet → Role → DB) and prioritize remediations. It allows agents to perform read-only security audits and generate Terraform fixes without data exfiltration.
    15
    46 PyPI
    4
    Apache 2.0
  • A
    license
    A
    quality
    A
    maintenance
    Enables MCP-aware agents to read and explore Oracle Cloud Infrastructure — listing, fetching, and searching across 39 resource types spanning Identity, Compute, Block Volume, Networking, Object Storage, and OKE. Runs locally over stdio using credentials from ~/.oci/config, with fail-closed safety gates and audit logging governing future write operations.
    4
    Apache 2.0