Skip to main content
Glama

aidefense_evaluate_program

Read-onlyIdempotent

Get the AI Defense Matrix evaluation playbook for assessing an AI security program: per-cell prompts, gap-inventory template, and a workflow that walks each asset class first and rolls findings up to the Govern column. Supports mode='gate' for binary deployment-gate decisions (returns the deployment-gate workflow plus gate-tier prompts only) and consumerPattern for scoping to consumed-vs-built AI deployments. The AI applies these prompts against your program documentation locally, and no program details leave your client. This server never requests your program docs or product roadmap and instructs your AI to keep them local—the matrix, framework alignments, and playbooks flow to your AI for local analysis.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
modeNoassessment (default): full program assessment — all maturity tiers. gate: binary deployment-gate decision — returns the deploymentGateWorkflow plus only gate-tier prompts (drops 90-day and mature prompts).
assetNoOptional: focus the prompts on one asset class.
assetsNoOptional: focus the prompts on multiple asset classes (e.g., for a deployment that touches orchestration + runtime data + agent identities). Takes precedence over `asset` if both are set.
functionNoOptional: focus the prompts on one NIST CSF function.
frameworkNoOptional: scope cellPrompts to those whose 'sources' field cites the named framework. Accepts a bare slug ('iso-42001') for any prompt citing that framework, or a 'framework:concept-id' form ('mitre-atlas:AML.T0051') to match an exact technique. Composes with mode, consumer_pattern, asset, and assets.
consumer_patternNoconsumed: organization consumes a third-party model (GPT-4 via API) — drops AI-Workload Platforms, Training Data, AI-Generated Code rows, and AI Model identify/protect/detect/respond/recover (keeps ai-model.govern). built: organization hosts/trains its own model — all rows in scope. hybrid (default behavior when omitted): all rows in scope.

TDQS

A4.5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already mark the operation as read-only, idempotent, and non-destructive. The description goes further by disclosing that the AI applies prompts locally, no program details leave the client, the server never requests program docs or roadmap, and gate mode returns only the deployment-gate workflow plus gate-tier prompts. This is valuable behavioral context beyond annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is moderately sized but each sentence earns its place: core purpose first, then mode/pattern options, then privacy/processing details. No redundant phrasing or filler; the structure is logical and front-loaded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Despite no output schema, the description tells the caller what to expect: per-cell prompts, gap-inventory template, workflow, and gate-mode variant. It covers the full range of use cases and important constraints (local processing, no data exfiltration). For a tool with six optional parameters, this is well-rounded and complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so baseline is 3. The description adds some meaning for mode='gate' (binary deployment-gate returns gate workflow plus prompts only) and consumer pattern scoping. However, it refers to 'consumerPattern' instead of the schema's 'consumer_pattern', which is a naming inconsistency that could confuse an agent. Since the schema already documents parameters fully, the description neither significantly adds nor detracts, except for this mismatch.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource: 'Get the AI Defense Matrix evaluation playbook' for assessing an AI security program. It lists concrete contents (per-cell prompts, gap-inventory template, workflow) and distinct behavior for mode='gate'. This clearly distinguishes it from sibling tools like 'aidefense_get_matrix' and domain-specific template services.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives clear context for when to use the tool (assessing an AI security program, binary deployment-gate decisions via mode='gate', scoping via consumer_pattern). It does not explicitly name alternative tools or state 'when not to use', but the specificity of the use case is strong enough for an agent to make a selection.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Try in Browser

Glama MCP Gateway

Add one secure layer between your agents and this server.

TDQS

A4/5.0
Disambiguation4/5

Each tool has a distinct domain prefix (assessment, cti, ir, malware, vuln, product) that makes its purpose clear, and overlapping tools like get_template and get_brief_template are explicitly described as standalone variants. However, the large number of similarly structured tools (get_guidelines, load_context, get_frameworks) across six domains could still cause an agent to reach for the wrong domain's tool without close attention.

Naming Consistency5/5

Tool names follow a highly consistent snake_case pattern: domain_get_* (e.g., ir_get_template, cti_get_guidelines), domain_load_context, domain_review_report, and domain_get_cross_server_routes. The few exceptions like product_compare_context and search_zeltser still match the general verb-first or domain-first style, so the overall naming scheme is predictable and readable.

Tool Count2/5

With 53 tools, the server is far too large for its stated purpose of 'website search.' The vast majority of tools are not about search but rather about writing guidelines, templates, and scoring rubrics for multiple report domains (assessment, CTI, IR, malware, vuln, product) plus an AI defense matrix. The count is inflated by repeating the same set of ~7 tools for six different domains, making it feel bloated and hard to navigate.

Completeness4/5

For the broad coaching/report-writing scope, coverage is thorough: every domain has a template, guidelines, context loader, review criteria, and frameworks, plus generic writing guidance and rating tools. The actual search capability is minimal (search_zeltser, get_article, get_index_info) but adequate for the core task; a notable gap is the lack of a tool to list or browse all articles, which would make discovery easier.

Resources