Skip to main content
Glama
Mipiti
by Mipiti

Mipiti MCP Server

MCP (Model Context Protocol) server for Mipiti — security posture platform.

Lets AI coding agents (Claude Code, Claude Desktop, Cursor, etc.) generate and manage threat models, controls, assumptions, compliance mapping, and evidence programmatically.

The Mipiti backend hosts an MCP server at https://api.mipiti.io/mcp. No installation needed — just configure your MCP client to connect.

Claude Code (quickstart)

claude mcp add --transport http Mipiti https://api.mipiti.io/mcp

You'll be prompted to log in via your browser (OAuth). That's it.

OAuth (manual config)

MCP clients with OAuth support (Claude Code, Claude Desktop, Cursor) automatically prompt you to log in via your browser. Add to your project's .mcp.json:

{
  "mcpServers": {
    "mipiti": {
      "type": "http",
      "url": "https://api.mipiti.io/mcp"
    }
  }
}

On first connection, your MCP client opens a browser window where you approve access with your Mipiti account. Tokens refresh automatically.

API Key

For clients without OAuth support, or headless/CI environments, create an API key in Settings:

{
  "mcpServers": {
    "mipiti": {
      "type": "http",
      "url": "https://api.mipiti.io/mcp",
      "headers": {
        "X-API-Key": "your-api-key"
      }
    }
  }
}

Related MCP server: threatmodel-mcp

Standalone Package (Alternative)

If you prefer running the MCP server locally (e.g., for development or self-hosted instances), install the mipiti-mcp package. This is a thin HTTP client that calls the Mipiti API.

pip install mipiti-mcp
# Or run directly with uvx
uvx mipiti-mcp

Environment Variables

Variable

Required

Default

Description

MIPITI_API_KEY

Yes

—

Your Mipiti API key

MIPITI_API_URL

No

https://api.mipiti.io

API base URL

SERVER_VERSION

Yes

—

Identifier for the running server's MCP surface (instructions, tool docstrings, schemas, behavior). Sent on every tool call. Clients invalidate cached MCP guidance when this changes. For local runs, any sentinel string is fine ("local", "dev"). For deployed runs, use a value that changes when this package's source changes (commit SHA is typical).

Claude Code (standalone)

{
  "mcpServers": {
    "mipiti": {
      "command": "uvx",
      "args": ["mipiti-mcp"],
      "env": {
        "MIPITI_API_KEY": "your-api-key",
        "SERVER_VERSION": "local"
      }
    }
  }
}

Tools (128)

Threat Modeling

Tool

Description

generate_threat_model

Generate a complete threat model from a feature description. Runs a multi-step AI pipeline producing trust boundaries, assets, attackers, control objectives, and assumptions. Progress reported automatically via MCP protocol — the tool blocks until complete. Optional provenance_* params record where the description came from at creation (for a repository: provenance_kind="code" + provenance_repo_url + provenance_commit_sha).

update_threat_model

Change a model's metadata: name (no new version; titles are unique within a workspace, case-insensitive), parent_id or clear_parent (its place on the recursive composition tree; no new version), and the provenance_* values (where its description came from: code with a commit SHA means the code is authoritative and the model follows it, anything else means the description is intent and the code is measured against it; bumps the version). Changes apply in that order; a failure names the ones already applied.

refine_threat_model

Refine an existing threat model based on an instruction. Creates a new version. Only affected entity types are modified — unaffected entities are preserved server-side. An instruction that cannot be applied (a targeted change naming an entity the model does not have) writes nothing and returns {model_id, changed: false, message}.

query_threat_model

Ask a question about an existing threat model. It only answers; it never changes the model.

get_threat_model

Get the full details of a specific threat model (trust boundaries, assets, attackers, assumptions). Use include_cos=True to include control objectives.

list_threat_models

List all saved threat models with IDs, titles, versions, and creation dates. Supports source filter and include_assessment_summary=True to inline per-model posture counts in one call (avoids N+1 looping assess_model).

delete_threat_model

Permanently delete a model and all its data.

export_report (scope="model")

Export as PDF, HTML, or CSV.

export_report (scope="model", format="archive")

Export the self-contained JSON audit archive of the model's current state (latest version, controls, live assertions with CI verdicts, findings, decisions in force, attestations, sufficiency signatures). Independently verifiable: the verdicts in it are the origin's record of what it claimed, which is what a third party checks against the signatures.

import_threat_model_archive

Restore an audit archive into a target workspace as version 1 of a new model. Fresh model_id per import; title collisions auto-suffix. It queues no judgement: the result carries the estimate, and judge_objectives queues it on request. The restored model arrives unverified — the origin's assertion verdicts and run-attested flags are not credited in the importing workspace, which earns them by running verification against code it can reach.

Entity CRUD

Tool

Description

add_asset / edit_asset / remove_entity (entity_type="asset")

Targeted single-entity changes for assets. Creates a new version.

add_attacker / edit_attacker / remove_entity (entity_type="attacker")

Same for attackers. surface_extent (whole: the attacker's operations range over any entry of the interface it reaches; point: one named entry) is an operator declaration: supplying it attests it and requires change_reason on either tool, and an attested whole makes the objectives that attacker anchors for-all obligations. A create declares only whole; narrowing is an edit_attacker call, checked against the objectives the attacker anchors.

revalidate_entity_quality

Judge every live asset's and attacker's quality warning again, in the background. Creates no version; returns {accepted, queued, model}, and the refreshed warnings appear on the next read of the model.

get_entity

Read one entity of any kind. An attacker also carries surface_extent and surface_extent_source, which says whether a person attested it.

Trust Boundaries

Tool

Description

get_threat_model

Returns existing trust boundaries (along with assets, attackers, assumptions). Review current boundaries before adding or modifying.

add_trust_boundary / edit_trust_boundary / remove_entity (entity_type="trust_boundary")

CRUD for trust boundaries. Defines where trust transitions occur in the system architecture. Attackers are positioned at boundaries; COs are annotated with boundary reachability. Changes auto-generate boundary assumptions for newly unreachable COs.

Controls

Tool

Description

get_controls

List controls with current status. Use summary_only=True for a compact response (id, description, status, verification_status, assertion_count, co_ids, assumption_groups, attestation_dependency).

get_control_objectives

List COs with which controls cover each one. Pair with get_reachability_verdicts for per-CO composer reachability state.

update_control_status

Mark implemented or not_implemented. Requires at least one assertion first.

refine_control

Modify a control's description with justification. Platform evaluates whether the mitigation group still covers the COs. An accepted refinement keeps the control's assertions and judges them again against the new description; nothing is superseded.

regenerate_controls

Propose a regeneration of the controls; it starts nothing. Supports mode="per_co" and co_ids to target specific COs.

get_control_generation_status

The model's proposed control build (proposal, with its estimate and the model_version and set_revision a start must name) and the last build started: its status, progress, and the next action (hint).

start_control_build

Start the proposed build. Generating or refining a model, editing an entity and regenerate_controls each propose one and start nothing. Call once for the proposal and a fresh estimate, then with confirm_estimate=True and the model_version and set_revision reviewed. A started build holds the model until it publishes its result in one step; meanwhile other writers of its controls are refused and reads show the last published controls.

discard_control_build

Drop a held build (queued, deferred, paused or blocked): what it staged is discarded, the published controls are untouched, and the build is proposed again. A running build is paused first.

list_control_revisions / undo_model_change

Every change to a version's controls, with its author. undo_model_change(target="controls") undoes the latest one (latest first, no redo); target="version" replaces the latest model version with a copy of the latest earlier version not already discarded, keeping the replaced version in the history as discarded.

pause_control_generation

Pause a model's control build (for example one started by mistake). A running build stops at its next step; everything done so far is kept unpublished, nothing new is started or billed, and nothing resumes it except resume_control_generation; discard_control_build drops it instead. A paused model can then be deleted as usual.

resume_control_generation

Resume control generation that was paused (get_control_generation_status reports paused), or retry one that stopped before finishing (blocked): a service it depends on was unavailable, or some new controls could not be checked for duplicates and were held back. A paused run resumes at once; for a blocked one the services are checked first, so a retry during an outage costs nothing. Either way only the unfinished work runs, billed to the original generation.

strengthen_controls

Work on the objectives whose mitigation groups the background judge found do not cover them. Generation stops after drafting and judging unless the workspace strengthens automatically; get_control_generation_status reports the diagnosis. Call once for the estimate (nothing starts, nothing is charged), then with confirm_estimate=True and the model_version and set_revision the estimate returned to start a background run. A gap only the environment can close is answered with an assumption, never a control: an accepted one is bound into the group, and otherwise a proposal waits in the review queue for a person.

judge_objectives

Judge every objective that has no judgement for its current controls and none queued: the diagnosis's not_judged count (judging counts the ones already queued; wait for those). Call once for the estimate (nothing is queued, nothing is charged), then with confirm_estimate=True to queue; any credits it consumes are metered as each judgement runs. Objectives with no mitigation group come back in ungrouped and are not judged. Not a repair: a judgement can come back insufficient. Refusals (409 generation in progress, 402 balance, 503 unavailable) come back as data.

import_controls

Import controls from JSON or free text, auto-mapped to COs and deduplicated. The imported controls await their judgement: the groups they join credit nothing until it is asked for.

judge_imported_controls

Estimate, and with confirm_estimate=True queue, the judgement of the imported controls awaiting one.

delete_control

Soft-delete with justification. Blocked if it's the only control covering a CO.

check_control_gaps

AI-powered gap analysis across all controls.

get_mitigation_groups / set_mitigation_groups

Inspect and modify how controls are grouped into mitigation paths for a CO (AND within groups, OR across groups). Platform AI-evaluates whether proposed changes preserve CO coverage.

set_control_objective_cal

Set per-CO ISO/SAE 21434 Cybersecurity Assurance Level (1-4). Persisted on the control_objectives identity side-table; survives soft-delete + revival; no new model version.

Assumptions and Attestation

Tool

Description

get_threat_model

Returns existing assumptions (along with assets, attackers, trust boundaries). Review current assumptions before adding or modifying.

add_assumption

Add an assumption, optionally linking it to COs via linked_co_ids.

edit_assumption

Update description and/or linked COs.

remove_entity (entity_type="assumption")

Soft-delete (preserved for audit). Its CO links are cleared and its attestations retired.

restore_entity (entity_type="assumption")

Restore a soft-deleted assumption to active. Its CO links are not restored (set them with edit_assumption), and it must be attested again.

submit_attestation

Record that a responsible party affirmed an assumption holds. Provide attested_by, statement, expires_at. A claim, never a proof over every site: it can cover an existential clause and never a for-all one. Attesting accepts the assumption, a judgment: a program is refused with 403 and an escalation_id unless the workspace delegates assumption_accepted to it. Editing the assumption's description retires the attestation.

list_attestations

Attestation history for an assumption.

set_control_assumption_groups

Declaratively set a control's assumption group structure: mark it externally handled by a single assumption (shorthand), clear the groups (the control's status is not changed), or express compound cases with multiple groups (within a group = AND, across groups = OR; e.g. "AWS KMS + quarterly review"). Attested groups count as active for mitigation group completeness.

get_control_assumption_groups

Inspect the current assumption group structure on a control. Groups express alternative sets of external claims (within = AND, across = OR).

convert_assumption_to_controls

Retire the assumption linkage and propose the control build its COs owe; start_control_build starts it.

Assertions and Evidence

Tool

Description

get_assertion_types

The catalogue as data: every type, what it proves, its soundness class, its params (an array-valued param carries its item schema), and the class vocabulary. Read-only.

submit_assertions

Submit typed, machine-verifiable claims about system properties (30 assertion types) for a control, an assumption or a functional test (name one of control_id, assumption_id, functional_test_id). Each object for a control may carry covers: the objective id (CO-NN) or clause ids (cls_…) it proves; a declared binding survives review, an undeclared one is inferred and capped below sound credit.

list_assertions / delete_assertion

List or delete assertions for a control.

edit_evidence

Attach (action="add") or detach (action="remove") auxiliary metadata (docs, links). Evidence is contextual — only assertions prove implementation.

get_verification_report

Shows verified, partially verified, and unverified controls with sufficiency details.

get_sufficiency

Quick check: do the assertions of one control (or one functional test, with functional_test_id) collectively cover all aspects? For the per-clause work list read get_control_work_order: where the order names a required class for a clause, required_evidence carries the class, the clause id to bind evidence to, and a submission skeleton to fill in. A claim that carries a soundness_tier reports its weakest clause's tier.

get_scan_prompt

Returns targeted prompts for scanning the codebase against not_implemented controls.

get_review_queue

The workspace review queue, ranked: escalation, proposal (including assumption proposals a strengthening run raised), unaccepted_assumption (an assumption something depends on that is not accepted), open_assumption, stale_control (implemented/verified controls not checked in 90+ days). Escalations and proposals are decided with decide_proposal; an unaccepted assumption is accepted with submit_attestation. Start here for periodic maintenance.

submit_findings / list_findings / update_finding

Report and track negative findings (gap discovery).

remediate_finding

Without apply, read-only: a structured diff of the changes the remediation would make, shaped by the finding's kind (e.g. for structural_duplicate_controls: which controls would be kept, which dropped, the union of CO mappings + framework refs that would land on the survivor). With apply=True and a non-empty justification (one-line operator rationale, recorded on the audit trail), it commits them. The agent is responsible for the preview-then-apply norm — surface the diff and get explicit confirmation before applying.

Evidence soundness classes

Every assertion type declares the class of the fact it reports, and the class bounds what a passing verdict can establish. get_assertion_types returns it per type; the platform and the CI verifier hold their own tables equal to the catalogue's.

Class

A pass establishes

Types

presence

A named construct, configuration value, dependency, file or pattern occurrence exists in the tree. Existence, not behaviour; a test file existing is presence.

function_exists, class_exists, test_exists, the configuration, dependency, semantic and RTL structure types

under_approximating_scan

A syntactic scan over a scope with no false-positive guarantee. A clean result proves the absence of the syntactic form only.

pattern_matches, pattern_absent, no_plaintext_secret

existential_witness

A signed statement that a named execution ran and passed at this commit. Proves the path it drove and nothing beyond it.

test_attested

sound_over_approximation

Every site in a declared scope that can violate the property was enumerated, and each is a declared safe form or a reviewed exception. Sound modulo the declared sink list.

sink_default_deny

by_construction

The sink accepts only a declared boundary type, and every construction site of that type is default-denied.

typed_boundary

A clause that ranges over every entry of a surface (every endpoint, every query, every frame) is credited only by one of the two sound classes bound to it with covers; a test proves only the path it drove, and an attestation is a responsible party's claim. Which types carrying those classes a platform takes is a read, not an assumption: get_control_work_order's assertion_contract.sound_types names them, and where it names none the acts that remain are to scope the asset to the component the attacker actually reaches, attest a point extent with its reason, or record a risk acceptance or a not-applicable disposition. The two sound types take a declared scope, the sinks through which the property could be violated (a call, a constructor, a macro, a store to a named target such as an HDL assignment, or a module instantiation), a reviewed allowlist, and the property in one sentence; sink_default_deny adds the accepted safe_forms, typed_boundary the boundary_type and its constructors. Hardware sources are covered by the same rule.

Agent work orders & delegation

Tool

Description

get_control_work_order

The ticket for implementing one control: scan brief, what counts as proof (assertion contract), acceptance criteria, steps, reconcile rules, what this agent may decide on its own, open proposals, and the model's provenance. Call before implementing a control. Read-only.

reconcile_model

Reconcile the model with the code: pass the paths changed since the recorded commit and your observations (mechanism_named, component_present, component_absent, forbidden_behavior). The platform decides the consequence of each; proposals are never applied on the agent's word, except a component change on a code-derived model, which is applied and queued for a person's review.

create_proposal

Raise a change of scope or design (add_component, remove_component, design_change), or an assumption a gap needs that only the environment can meet. Raising is not deciding: a person (or an agent under a delegation rule) decides it with decide_proposal; design changes are never applied automatically.

list_proposals

Proposals and escalations on a model with their status (proposed / applied_pending_review open; accepted / rejected / reverted / superseded closed). A refused judgment (403 with escalation_id) appears as a decision_request; poll here until a person resolves it. Read-only.

decide_proposal

Accept or reject a proposal. A judgment: refused with 403 and an escalation_id unless the workspace's delegation policy names the decision for this agent at the proposal's tier. Do not retry a refusal. Accepting an assumption proposal accepts the assumption (its own decision, assumption_accepted), attested until expires_at; rejecting one keeps the precondition from being proposed again.

get_design_leverage

What eliminating each attacker position or asset by design would remove from the matrix, ranked by critical then high at-risk objectives removed. include_design_moves=True authors a concrete design_move per row; turn one into a design_change proposal with create_proposal. Read-only.

list_decisions

The model's decision ledger: every judgment recorded on it (finding dismissed / remediated, risk accepted, not-applicable declared, proposal accepted / rejected / reverted, assumption accepted, escalation resolved), newest first, with who decided and whether it was within the delegation policy. Append-only; nothing edits it. Call before raising a proposal or asking for a judgment, so you do not propose what a person rejected or ask again for what was already decided. Read-only.

Assurance

Tool

Description

assess_model

Deterministic assessment of all COs. Returns mitigated/at_risk/unassessed with risk_reason (missing_controls, pending_attestation, expired_attestation, coverage_gap, insufficient_by_design). For per-CO reachability state call get_reachability_verdicts.

get_findings_risks

Workspace-scoped triage dashboard: open findings, active risk acceptances, and at-risk COs across every model the workspace can access. Entry point when asked "what's open?".

get_risk_view (scope="model")

Per-model Prioritized Risk View: one row per live CO with derived risk tier, asset impact, attacker likelihood, control coverage, and open-finding count.

get_risk_view (scope="tag")

Cross-model variant of get_risk_view (scope="model"): same shape, aggregated across every member of a tag (model_id + model_title attached per row), and delegation-aware.

get_remediation_leverage

Per-model remediation plan: the not-yet-satisfied controls ranked by how many COs each one closes, plus a greedy minimal fix order (summary / ranked / greedy_plan). Use to prioritize which controls to implement first for the shortest path to coverage.

list_risk_acceptances

All risk acceptances on a model — risks explicitly accepted instead of mitigated. Includes CO id, owner, justification, status, review deadline.

create_co_disposition

Record that a control objective does not apply to this system (owner, justification, review deadline). The sibling of a risk acceptance: an acceptance says the exposure is real and is being carried, a disposition says the objective does not apply here at all. The objective stays in the matrix and in every coverage count, reported in its own class — what is suppressed is work (no controls generated, no coverage gap raised), never the accounting.

list_co_dispositions

Every signed judgment on a model's objectives, both kinds. Expired and revoked entries are included: a lapsed decision is part of the audit trail. Optional kind filter.

recompute_verdicts

mode="quote" (the default) returns the informational cost estimate and enqueues nothing (carries computed_at + the pricing rate_version). mode="recompute" force-enqueues a fresh evaluation of every control's coverage verdict and every live CO's group-sufficiency verdict, bypassing the quiet-period batching; actuals are metered as evaluation runs. mode="retry_parked" re-runs only the verdicts a transient failure parked, of every kind. Each carries a spend status object — exhausted means the work is queued and resumes automatically, never dropped.

judge_objective

Have ONE control objective's mitigation group judged — the remedy for an objective reading awaiting_judgement, where a group is built and nothing has decided whether it covers the objective. Prefer it over recompute_verdicts, which sweeps the whole model. Runs in the background and may consume credits. Not a repair: the judgement can come back insufficient, moving the objective to coverage_gap / insufficient_by_design. Refusals (409 generation in progress, 503 unavailable, 402 balance) come back as data.

Functional Conformance

Proves a feature does what it was specified to do (Capability × Condition), verified by the same assertion + CI engine as security controls.

Tool

Description

generate_functional_objectives

Derive capabilities (behaviours the feature must deliver), Given-When-Then functional objectives (walking each capability against a taxonomy of operating conditions), and a concrete implementable test per objective — so the agent implements the tests rather than deciding what to test. Requires a Pro plan; billable. refresh=true re-derives.

get_capabilities

Read the capability decomposition: every capability, or one by capability_id.

get_functional_objectives

Read the functional objectives (the test plan).

get_functional_coverage

Per-objective + per-test state (verified / covered / failing / untested), the Capabilities × Conditions matrix, and applicable / missing-objective / not-applicable cell accounting. gaps_only=True returns just the actionable gaps: applicable conditions with no objective yet, plus objectives that are failing or untested.

get_scan_prompt (kind="functional")

The agent brief: per not-yet-verified test, its implementation brief and the objectives it proves; plus objectives with no test and applicable conditions with no objective.

add_functional_test

Manually register an extra test satisfying one or more objectives (generation already specifies the tests; a manual test survives regeneration).

submit_assertions (functional_test_id=…)

Submit evidence assertions for a functional test (verified in CI, same as control evidence).

Composition (recursive-tree effective model)

Views over the effective model — own entities composed with everything inherited from ancestor threat models on the recursive tree. Available where the deployment enables composition; where it does not, read tools return a stable empty body with flag_enabled: false and the write tool returns 503.

Tool

Description

get_composition (view="overview")

Index: counts + tree metadata (parent_id, ancestor_chain, depth, child_ids) + structural warnings. Cheapest call — use first to learn whether composition is enabled and orient on the tree.

get_composition (view="entities")

Effective entity set keyed by kind (trust boundaries, components, assets, attackers, attack paths). Each entry carries provenance (own vs inherited) and a fully-qualified id for cross-model references.

get_composition (view="objectives")

Effective COs tagged with origin (own / cross / inherited). Pair with view="coverage" and get_reachability_verdicts (composed=True).

get_composition (view="coverage")

Per-CO coverage with credited inheritance: own_credit, inherited_credit, and the list of contributing controls (with the owning model id, origin, verification status, mitigation group). This is what drives the composition coverage view, not per-model get_verification_report.

get_reachability_verdicts (composed=True)

Per-CO reachability verdicts over the composed effective topology — same kinds (reachable / unreachable / indeterminate) as get_reachability_verdicts, but evaluated against the merged tree. Use on child models when ancestor topology matters.

get_composition (view="attack_paths")

Effective AttackPath set + lifted missing/dangling suggestions computed against the composed reach surface.

list_reconciliation_candidates

Paginated reconciliation candidates between this model and its ancestors. Tier certain is a deterministic match safe to auto-apply; tier heuristic is fuzzy and needs review. With disposition="rejected", the persisted rejections instead, oldest first, each with the surrogate id an unreject names.

decide_reconciliation_candidate

Mutating. decision="apply" records that the descendant's own entity is the inherited one: the own entity stays in the model and the composed view leaves it out, so the inherited entity is canonical (the record is dropped when the pair stops matching); the server re-validates against current live state and refuses a heuristic-tier candidate unless confirm_heuristic=True; bumps the model version and returns {model, controls_carried, controls_orphaned, orphaned_control_ids}. decision="reject" persists "these are NOT duplicates" at org scope, so the detector leaves the pair out of the active queue; idempotent on (model_id, kind, own_qid, inherited_qid), no new version, returns the record (keep its id). decision="unreject" removes a rejection by rejection_id, returning {ok: true}.

lift_composition_entity

Mutating. Promote a shared-anchor entity from two sibling descendants to their lowest common ancestor: each source's copy is soft-deleted and the inherited entity becomes canonical for every descendant of the LCA. Server re-detects field-level and attached-state conflicts against current live state; pass field_resolutions / attached_state_resolutions keyed by the conflict keys returned in the 400 detail. Server also runs an over-application gate against the LCA's descendant set; pass acknowledged_third_party_subtrees to acknowledge extra reach or skip_overapplication_gate=true to override after explicit operator confirmation. Bumps version on the LCA + both source descendants; returns {lift_id, lca_model, descendant_a_model, descendant_b_model, applied_migrations, lift_event} — the lift_event block matches the audit pack's lift_history entry.

split_composition_entity

Mutating. Inverse of lift_composition_entity: push an ancestor-owned entity down to one or more target descendants and soft-delete the ancestor's copy. A new local id is minted on each target; attached state (assertions, jira mappings, risk acceptances) on the ancestor's entity is duplicated to every target. Bumps version on the ancestor + every target descendant; returns {split_id, ancestor_model, descendant_models, applied_duplications, split_event} — the split_event block matches the audit pack's split_history entry.

undo_composition_event

Undo a prior lift or split (event_type="lift" or "split"). By default (dry_run=True) read-only: the inverse plan or the divergence refusal, as {plan, refusal} with exactly one non-null — surface it to the operator. With dry_run=False, mutating: re-runs the divergence detector and refuses with 409 + detail.refusal.reasons when state has materially evolved since the forward event; on success persists the inverse across every affected model and emits a lift_undone / split_undone activity event citing original_event_id, so the audit pack chains undo to its forward. Returns {undone_event_id, original_event_id, applied_state_ops, models}, models being {lca_model, source_descendant_models} for a lift and {ancestor_model, descendant_models} for a split.

Cross-model dependencies (delegation)

Declared reliance edges (distinct from the parent/composition tree, which is containment): a model depends on a control implemented in another model — for systems built on shared services (auth, logging, shared data) rather than sub-parts. The target is always a provider control (credit terminates at a proven mechanism). Reliance is workspace-scoped: a consumer can only delegate to provider models in the same workspace (these tools don't see models across workspace boundaries). Available where the deployment enables the model tree; whether reliance carries credit is also a deployment setting.

Tool

Description

declare_foundation

Mark a shared-service model as a foundation that advertises specific controls (provides) other models can delegate to. A capability advertises a control, never an objective.

manage_reliance

One edge at a time. action="create" declares a dependency: delegated (consumer has no local control for an objective; provider handles it — pass source_objective_id) or relied_upon (consumer keeps its own control but its validity depends on the provider's — pass source_control_id); it enters draft and runs LLM semantic validation. action="confirm" promotes a draft edge to active — the credit-soundness gate, refused unless validation returned valid (a partial/mode-mismatch is never silently credited). action="delete" removes an edge.

list_reliance

A model's dependency edges (as consumer) plus who relies on it (as provider — the blast radius before changing its controls).

attach_foundation

Without selections, read-only: which of a consumer's objectives each foundation capability covers (scored). With the chosen subset as selections, bulk-creates draft delegation edges for those (objective, provider control) pairs; each runs LLM validation and none credits until confirmed.

Tags (grouping)

Overlapping, semantics-free grouping of models (the Affiliation primitive) — for audit scopes, products, ad-hoc selections, or portfolios. A model may carry many tags; a tag never affects posture or credit. The group tools act on tags.

Tool

Description

create_group / delete_group

Create or remove a tag (deleting affects the grouping only, not the member models).

add_model_to_group / remove_model_from_group

Manage membership; a model can belong to many tags at once. A model added to a tag takes on the frameworks the tag selected.

list_groups / get_group / list_model_groups

Browse the workspace's tags, read one with its members, or list a model's tags.

get_group_dependencies

The reliance edges among a tag's members, each with its status and whether it credits its objective.

get_risk_view (scope="tag")

Aggregate per-CO risk across a tag's members. Delegation-aware (a CO mitigated via a verified cross-model delegation reads as covered).

select_compliance_frameworks (scope="tag")

Make a tag a compliance/audit scope: select frameworks for the tag, propagated to its members and to every model added later.

get_compliance_report (scope="tag")

Cross-model compliance coverage report scoped to a tag's members.

export_report (scope="tag")

Signed auditor HTML for a tag: every member's report, after the reliance edges among the members with their status.

Compliance

Tool

Description

list_compliance_frameworks

Available frameworks: the built-ins (among them OWASP ASVS, ISO 27001, SOC 2, NIST CSF, IEC 62443, ISO/SAE 21434, PCI DSS, GDPR) and any imported.

import_compliance_framework

Import a customer-specific framework (JSON: name, requirements, optional level_definitions).

select_compliance_frameworks

Select frameworks for a model, or for a tag (scope="tag").

get_compliance_report

Coverage report for a selected framework.

auto_map_controls

AI-powered semantic mapping of controls to framework requirements.

map_control_to_requirement

Manual control-to-requirement mapping.

auto_remediate_compliance

LLM-powered gap closure — proposes new assets, attackers, and controls for uncovered framework requirements.

Components

Tool

Description

add_component / edit_component / remove_entity (entity_type="component")

Components bridge trust boundaries (security architecture) to repositories (code organization). Component(id, name, repo_url, path, trust_boundary_ids) scopes controls to the codebase that implements them. Used for multi-repo systems and per-repo threat models. edit_component also accepts optional per-component level grades: target_sl (IEC 62443 Security Level, 1-4), eal (Common Criteria Evaluation Assurance Level, 1-7), fips_level (FIPS 140-3 Security Level, 1-4).

Organizations

Tool

Description

update_organization

Set per-organization level grades: target_ml (IEC 62443-4-1 Maturity Level, 1-5), csf_tier (NIST CSF Tier, 1-4). Admin-only. Use clear_target_ml / clear_csf_tier to explicitly reset to NULL.

Setup and Operations

Tool

Description

get_setup_status

Check which onboarding steps are done.

complete_setup_step

Mark an onboarding step as done (mcp_configured, mipiti_verify_installed, ci_secret_added, ci_pipeline_added).

CWE Classification

Tool

Description

get_cwe_catalog

Get the platform's CWE reference catalog status (current MITRE version, entry count). Reports enabled: false when not turned on for this instance.

get_model_cwe_tags

List CWE weakness classifications tagged onto a model's control objectives, with a staleness marker for tags whose catalog entry has since been deprecated, redefined, or removed.

classify_model_cwe

Classify a model's control objectives against the platform CWE catalog. Grounded — the model may only select from the catalog's current candidates, and every returned id is re-validated before storage.

Development

git clone https://github.com/Mipiti/mipiti-mcp.git
cd mipiti-mcp
pip install -e ".[dev]"
python -m pytest -v

Local Testing with Claude Desktop

{
  "mcpServers": {
    "mipiti": {
      "command": "uv",
      "args": ["run", "--directory", "/path/to/mipiti-mcp", "mipiti-mcp"],
      "env": {
        "MIPITI_API_KEY": "your-key",
        "SERVER_VERSION": "local"
      }
    }
  }
}

License

Proprietary. Copyright (c) 2026 Mipiti, Inc. All rights reserved. See LICENSE for details.

Available Tools

128 tools
add_assetAdd AssetA

Add a new asset to a threat model. Creates a new version.

Authoring contract: name the data or resource being protected and the security property at stake (Confidentiality / Integrity / Availability / Usage), not a mechanism or control ("per-organization key-wrapping material", not "KMS encryption"). An asset phrased as a mechanism is flagged with a quality_warning and yields under-specified control objectives. An asset that does not apply is recorded with a non-applicability assumption or create_co_disposition; there is no status to set.

The platform reasons the factor decomposition and composes impact with the prompt generation uses, so factors are calibrated alike; override one afterwards with edit_asset and a change_reason. component_ids links the asset to the deployable units that hold it, which feeds reachability (several for a multi-instance asset, e.g. a session token on client and cache).

A proposal matching a soft-deleted asset is gated: it either restores that asset (auto_restored: true, restored_asset_id, discarded_fields; its CO tombstones revive) or is refused as similar ({accepted: false, classification: "similar", candidate_restore_id}, nothing saved). A normal create returns {model, controls_carried, ...}. 503: an evaluator is unavailable, retry with backoff; 502: it answered malformed, retry.

ParametersJSON Schema
NameRequiredDescriptionDefault
nameYesAsset name (required).
notesNoOptional notes.
model_idYesID of the threat model.
descriptionNoOptional description (recommended — feeds the factor-reasoning prompt).
component_idsNoComma-separated component IDs scoping the asset (e.g., "CMP1,CMP2"). Empty / omitted = unscoped. Validated against components declared on the model.
server_versionYes
security_propertiesNoComma-separated properties, e.g. "C,I,A" (default: "C").

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does so richly: it discloses the versioning side effect, quality_warning flagging for mechanism-phrased assets, the soft-deleted-asset gating behavior (auto_restored, restored_asset_id, discarded_fields, or a {accepted: false, classification: "similar"} refusal), and 503/502 retry semantics. This is behavior an agent cannot infer from the schema.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The purpose is front-loaded and the contract/error details are organized under clear paragraphs. It is dense and fairly long, with a couple of asides (e.g., the factor-decomposition rationale) that could be trimmed, but nearly every sentence carries actionable information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 7-parameter mutation tool with an output schema, the description covers the mutation contract, the authoring requirement, edge cases (soft-deleted matches), and error/retry behavior. An output schema exists, so its decision to also sketch return shapes is a bonus rather than a gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 86%, so the schema already documents most parameters and the baseline is 3. The description adds meaning beyond the schema by explaining that component_ids 'feeds reachability' for multi-instance assets, and by giving the authoring contract that shapes what 'name' should contain. That is genuine added value, though it doesn't cover every parameter.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The first sentence gives a specific verb+resource ('Add a new asset to a threat model') and immediately notes the side effect ('Creates a new version'). It also routes the agent to the sibling edit_asset for post-hoc changes, so the tool is distinguishable from its nearest neighbor without opening a schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives clear context on when to use this tool (creating a new asset, and using create_co_disposition / non-applicability assumptions when an asset doesn't apply) and explicitly names edit_asset for overrides. It stops short of an explicit 'use add_asset when X, not when Y' clause, so a 4 rather than a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

add_assumptionAdd AssumptionA

Add an assumption. Creates a new model version.

Assumptions represent security properties outside the system owner's trust boundary. When linked to COs and attested, they mitigate those COs in the assessment.

Optionally attach a structured exclusion predicate (the exclusion_* params). The reachability composer matches active

  • attested assumptions with predicates against COs deterministically — class-3 (deterministic computation) evidence in addition to the operator-attested class-1 evidence. Pass any subset of the fields; unspecified fields default to wildcard ("*"). When exclusion_co_ids is non-empty, it takes precedence over the match fields.

Use this to resolve a CO whose composer verdict is indeterminate because no structural primitive backs an operator non-applicability claim: set exclusion_co_ids=<co_id> (and optionally the attacker/asset/property fields), and the composer will derive unreachable / reason: assumption_excludes on subsequent loads, with the assumption's structured predicate as the audit-trail cause.

ParametersJSON Schema
NameRequiredDescriptionDefault
model_idYesID of the threat model.
descriptionYesWhat is assumed (e.g., "Customer restricts CI runner egress").
linked_co_idsNoOptional comma-separated CO IDs this assumption covers.
server_versionYes
assumption_typeNo"external" (default, allows manual attestation) or "non_applicability" (requires CI verification, no manual attestation).external
exclusion_co_idsNoComma-separated CO IDs the predicate matches explicitly. When non-empty, overrides the match fields.
exclusion_asset_idNo"*" or concrete asset ID.
exclusion_attacker_idNoPredicate match — "*" wildcard (default when any other exclusion_* param is set) or concrete attacker ID.
exclusion_property_matchNo"C" | "I" | "A" | "U" | "*".
exclusion_attacker_vectorNoOne of "Network" | "Adjacent" | "Local" | "Physical" | "*".
exclusion_asset_component_idNo"*" or concrete component ID.

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It discloses the mutation side effect, the attestation/mitigation lifecycle, deterministic class-3 evidence behavior, wildcard defaults, and precedence rules. This is strong behavioral context, though it does not cover operational concerns like permissions or reversibility.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is front-loaded with the core action and side effect, then organizes the advanced exclusion predicate behavior into a focused section. It is fairly long, but the content earns its place given the tool's complexity and 11 parameters.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a complex tool with no annotations, an output schema, and high schema coverage, the description covers the essential operational context: side effects, matching behavior, defaults, precedence, and a concrete use case. It does not spell out every assumption lifecycle step, but the schema and output schema cover the structured details.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 91%, setting a high baseline. The description adds meaningful semantics beyond the schema: it explains that any subset of exclusion_* fields can be passed, unspecified fields default to wildcard, exclusion_co_ids takes precedence, and the predicate feeds deterministic composer evidence. This exceeds the baseline.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource: "Add an assumption." and clearly states the side effect "Creates a new model version." This distinguishes add_assumption from related operations like edit_assumption without relying on the title alone.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives a concrete, decision-relevant scenario: use this to resolve a CO with an "indeterminate" composer verdict by setting exclusion_co_ids. It provides clear context for when the exclusion_* predicate mechanism is appropriate, though it does not explicitly name alternatives or state when not to use the tool.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

add_attackerAdd AttackerA

Add a new attacker to a threat model. Creates a new version.

Authoring contract: capability names the operations the attacker can perform from its position and what they achieve — not just the access or vantage point. Phrase it as "From [position], the attacker can [concrete operations] …" (e.g. "From the network path between the API server and the database, the attacker can read and alter requests and responses to exfiltrate data in transit or inject forged responses"). A capability that states only access is flagged with a quality_warning and the control objectives derived from it may be under-specified.

The caller supplies identity-bearing fields (capability, position, archetype, trust_boundary_ids); the backend LLM-reasons the factor decomposition. Override any factor post-create via edit_attacker with a change_reason. Mirror of add_asset semantics.

Three outcomes (normal create / auto-restore / similar-rejection) mirror add_asset. 503 on factor-reasoning or restore-candidate evaluator outage, 502 on malformed restore-candidate response.

surface_extent says how much of the reached interface this attacker's operations range over. An attacker ranging over the whole interface makes the objectives it appears in for-all obligations, which only a sound witness (typed_boundary / sink_default_deny) can credit. Declaring it here is an operator statement about the attacker's reach, recorded attested with its change_reason, so a create takes the two together. Only whole is declarable on a create: narrowing to one named entry is a statement about the objectives the attacker anchors, and a create has none yet — add the attacker, then narrow it with edit_attacker and a change_reason, where the narrowing is checked against the assets those objectives defend. There is no attacker status to set.

ParametersJSON Schema
NameRequiredDescriptionDefault
model_idYesID of the threat model.
positionNoPosition/access level.
archetypeNoArchetype (e.g., "insider", "external").
capabilityYesAttacker capability description (required).
change_reasonNoRequired when ``surface_extent`` is supplied — documents the declaration for the audit trail.
server_versionYes
surface_extentNo``"whole"`` when, from its position, the attacker's operations range over ANY entry of the interface it reaches (any endpoint, request, row, file, message or frame). Recorded as attested by this call and requires ``change_reason``. Omit to leave it undeclared, which is the ordinary case.
trust_boundary_idsNoComma-separated trust boundary IDs.

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.9/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations present, the description carries the full transparency burden and does so thoroughly. It discloses that the call creates a new version, that the backend LLM-reasons the factor decomposition, that three outcomes mirror add_asset, that 503/502 errors can occur, and that surface_extent is recorded as an attested operator statement.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is long but information-dense, with important constraints front-loaded and clearly separated. It loses a point for some redundancy, such as mentioning the add_asset mirror twice, and for an extended rationale about for-all obligations that could be tighter.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a create operation with backend reasoning, side effects, error conditions, and an attestation parameter, the description covers all the non-obvious context an agent needs. It explains versioning, alternative routing, failure modes, the capability contract, and the special surface_extent workflow, while output format is covered by the output schema.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is high at 88%, but the description still adds significant meaning beyond the schema. It defines the capability authoring contract with a concrete pattern and quality_warning consequence, and it explains surface_extent's create-time restriction and attestation semantics in ways the schema does not.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Opens with a specific verb and resource: "Add a new attacker to a threat model." It also distinguishes itself from the sibling edit_attacker by noting that post-create factor overrides and surface narrowing happen there, and it names add_asset as its semantic mirror.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives explicit when-to-use guidance: use this tool to create an attacker, and use edit_attacker afterward to override factors or narrow surface_extent. It also tells the caller that only "whole" is declarable at create time and that narrowing belongs on a later edit, which prevents a common misuse.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

add_componentAdd ComponentA

Add a component to a threat model.

Components bridge security architecture to code organization. They map trust boundaries to repos so controls can be scoped to the codebase that implements them. They also drive the deterministic reachability composer's asset-boundary derivation: an asset's trust-boundary footprint is the union of its components' trust_boundary_ids.

Generation reads no components, so add or edit them after generate_threat_model, not before.

A component with empty repo_url is either speculative (your own code, not linked to a repo yet) or external (e.g. a third-party service, the customer's IdP, or other external infrastructure you call but don't own). The component's trust boundary tells them apart: bind an internal-zone component to its repo via edit_component; leave an external-zone component unbound — its component_unbound finding is a permanent external-dependency marker, not a gap to close. Binding by "some client code touches it" is wrong: client code for external dependencies lives in your repo too.

ParametersJSON Schema
NameRequiredDescriptionDefault
nameYesComponent name (e.g., "Backend API", "Auth Worker").
pathNoPath within repo for monorepos (e.g., "services/auth").
model_idYesID of the threat model.
repo_urlNoRepository URL (e.g., "github.com/org/backend"). Empty string is valid for speculative components — pass a real URL once you've identified the codebase.
server_versionYes
trust_boundary_idsNoComma-separated trust boundary IDs that this component spans (its deployment zone). Drives reach decisions for any asset scoped to this component.

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It discloses important behavioral traits: the tool's relationship to generation (must be called after generate_threat_model), the meaning of empty repo_url, and the permanent external-dependency marker behavior of component_unbound findings. It doesn't explicitly state side effects like whether it overwrites existing components or requires specific permissions, but the disclosed context is substantial.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is moderately long but every sentence earns its place. It front-loads the core purpose, then provides essential context about components, generation ordering, and the repo_url/trust boundary semantics. The structure is logical, though the final warning about client code could be seen as slightly tangential.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity (6 params, output schema present, many siblings), the description is quite complete. It covers the tool's role, when to use it, key parameter semantics, and common pitfalls. It doesn't describe the return value, but an output schema exists, so that's not required. It could mention prerequisites like needing a valid model_id, but that's likely implied by the schema.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 83%, so the schema already documents most parameters well. The description adds meaning beyond the schema by explaining the semantic distinction between speculative and external components (empty repo_url) and how trust_boundary_ids drive reach decisions. It doesn't add detail for every parameter, but the schema covers them adequately.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's function: 'Add a component to a threat model.' It goes beyond a simple definition by explaining what components are, how they relate to trust boundaries and repos, and how they drive the reachability composer. This distinguishes it from sibling tools like edit_component and add_trust_boundary.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides explicit usage guidance: 'Generation reads no components, so add or edit them after generate_threat_model, not before.' It also explains when a component with empty repo_url is speculative vs external, and explicitly warns against a common mistake ('Binding by "some client code touches it" is wrong'). This is strong, actionable guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

add_functional_testAdd Functional TestA

Hand-author a single functional test and map it to one or more objectives. Mutating.

Generation (generate_functional_objectives) already specifies the tests to implement, so use this only to register an extra test that generation did not produce; a manually-added test survives regeneration/refresh. For bulk-registering tests that already exist in your codebase, use import_functional_tests instead. This records the test at the status you claim — it does not run or verify anything; CI verification happens only when you attach TEST_EXISTS/TEST_ATTESTED evidence via submit_assertions with functional_test_id.

ParametersJSON Schema
NameRequiredDescriptionDefault
statusNonot_implemented | implemented | verified — an operator claim only; an independent CI run is what actually verifies the test. Defaults to not_implemented.not_implemented
model_idYesID of the threat model.
descriptionYesWhat the test proves.
component_idsNoComma-separated component ids the test exercises (optional).
server_versionYes
functional_objective_idsYesComma-separated objective ids the test satisfies (at least one required; get them from get_functional_objectives).

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and delivers: it declares the operation mutating, discloses a durability trait ('a manually-added test survives regeneration/refresh'), and draws the critical boundary that the recorded status is a claim only and is not run or verified here. That is exactly the kind of non-obvious behavior an agent needs.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences, each load-bearing: purpose first, then the disambiguation against the two sibling tools, then the verification caveat. No filler and the routing information is front-loaded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 6-param mutation tool with an output schema already present, nothing an agent needs is missing — required inputs are covered by the schema, and the description supplies the usage, persistence and verification context the schema cannot.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 83%, so the schema already documents status, functional_objective_ids, component_ids and the rest. The description reinforces the status-is-a-claim semantics and points to get_functional_objectives for ids, but adds little beyond the schema text, so the baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('Hand-author a single functional test') plus the mapping side-effect ('map it to one or more objectives') and flags 'Mutating.' An agent can distinguish this from import_functional_tests and generate_functional_objectives without opening any schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly says generation already specifies the tests, so use this only for an extra test generation did not produce, and names the alternative for bulk-registering existing tests (import_functional_tests). Both the when and the when-not are stated, plus a downstream note about CI verification via submit_assertions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

add_model_to_groupAdd Model To GroupA

Add a threat model to a tag. Mutating.

Links the model into the tag without moving or copying it; the model stays independently editable, and may belong to many tags. The model takes on the compliance frameworks the tag selected, marked as the tag's, so removing a framework from the tag removes it from the model again; a framework the model selected itself is left as it is.

ParametersJSON Schema
NameRequiredDescriptionDefault
group_idYesID of the tag.
model_idYesID of the threat model to add.
server_versionYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.9/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations the description carries the full burden, and it does substantial work: it discloses the operation is a link not a move/copy, that the model stays independently editable and may belong to many tags, and critically that tag-selected frameworks propagate and are removed again when dropped from the tag, while self-selected frameworks persist. Missing only auth/permission or failure-mode information.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with purpose and mutation flag in the first sentence, then a compact semantics paragraph. Every clause carries information; the framework-inheritance detail could be trimmed slightly but it is substantive rather than filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists so return values need not be explained. For a three-parameter mutation with no annotations, the description covers the important behavioral question (inheritance/removal semantics) but omits prerequisites and error conditions, leaving a small gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 67%; the schema documents group_id and model_id but leaves server_version bare. The description adds useful domain meaning by equating 'tag' with group_id, but never references model_id or server_version syntax, so it only partially compensates for the coverage gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Add a threat model to a tag') plus a mutation flag. Its sibling remove_model_from_group is the inverse operation, and this description is unambiguous about direction.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage context is implied rather than stated — nothing tells the agent when to prefer this over remove_model_from_group, list_model_groups, or create_group. The mutation flag gives a weak use signal but there are no explicit when/when-not conditions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

add_trust_boundaryAdd Trust BoundaryC

Add a trust boundary. Creates a new model version.

ParametersJSON Schema
NameRequiredDescriptionDefault
passesNoOptional comma-separated AttackVector values the boundary allows through (subset of "Network,Adjacent,Local,Physical"). Omit for the methodology default (passes-everything). Narrowing this set is what makes a boundary block specific attacker vectors in the deterministic reachability composer.
sealedNoOptional. Set True to declare the boundary has NO lateral ingress — the only way into its zone is crossing the perimeter (an air-gap / network-segmented enclave). On its own this is a suggestion: only an ATTESTED seal lets reachability decisively rule an asset unreachable instead of indeterminate, and the attestation is recorded with ``edit_trust_boundary`` (``seal_source="attested"`` with a ``change_reason``). Default False (assume a lateral pivot is possible). Set it only when the isolation is real and attestable.
crossesNoOptional comma-separated asset IDs that cross this boundary.
model_idYesID of the threat model.
descriptionYesWhat this boundary represents (e.g., "Public network to API server").
server_versionYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description must carry the full burden of behavioral disclosure. It only states 'Creates a new model version,' which is vague and does not explain side effects such as versioning implications, reversibility, permission requirements, or what happens to existing boundaries. The description fails to disclose any meaningful behavioral traits.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is extremely concise, consisting of two short sentences. It is front-loaded with the primary action and avoids redundancy. While it is arguably under-specified, it is efficiently structured with no wasted words.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with 6 parameters, an output schema, and a complex domain (trust boundaries in threat models), the description is insufficient. It does not clearly explain the model versioning behavior, any prerequisites, or the effect on the overall model. It also fails to reference the rich parameter semantics that the schema provides, leaving the agent without a clear operational picture.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema provides detailed descriptions for 5 of 6 parameters (83% coverage), so the description is not required to explain parameters. The description adds no additional meaning beyond what the schema already offers, earning the baseline score for high schema coverage.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the verb 'Add' and resource 'trust boundary', making the primary purpose obvious. However, it does not differentiate from the sibling edit_trust_boundary beyond the verb difference, and the added 'Creates a new model version' is ambiguous about whether that is a side effect or a separate action. Purpose is clear but not fully distinct.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool versus edit_trust_boundary or other model-modifying tools. The description does not mention prerequisites, the distinction between adding vs editing boundaries, or any context that would help an agent decide between alternatives.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

apply_control_changesetApply Control ChangesetA

Apply a batch of control operations atomically as ONE transaction.

Use this to reorganize a model's controls in a single step — for example to deduplicate controls (remap several onto the right objectives and delete the redundant ones at once), instead of many separate calls. All operations commit together or not at all.

Mapping-only: remap/delete/set_groups change objective mappings and retire controls but never re-author a control's description, so a kept or reused control keeps its status, evidence, and assertions. The orphan guard is evaluated on the FINAL state of the batch, so a delete paired with a covering remap or add in the same changeset is allowed; a changeset that would leave any previously-covered control objective uncovered is rejected as a whole and nothing is written.

ParametersJSON Schema
NameRequiredDescriptionDefault
opsYesJSON array of operation objects. Each object has an "op" of "remap", "delete", "add", or "set_groups": - remap: {"op": "remap", "control_id": "CTRL-03", "co_ids": ["CO1", "CO2"]} - delete: {"op": "delete", "control_id": "CTRL-09", "reason": "duplicate of CTRL-03"} - add: {"op": "add", "description": "...", "co_ids": ["CO5"], "mitigation_group": 1} - set_groups: {"op": "set_groups", "co_id": "CO1", "groups": {"1": ["CTRL-03"], "2": ["CTRL-04"]}}
model_idYesID of the threat model.
change_reasonYesWhy this reorganization is appropriate (min 10 chars). Recorded on every affected control's version history.
server_versionYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.4/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure and does so thoroughly. It explains atomic all-or-nothing commit, mapping-only semantics, preservation of status/evidence/assertions, and the orphan guard evaluated on the final state. This is exemplary transparency for a mutation tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is detailed but every sentence earns its place. It front-loads the core purpose, then adds necessary transaction and constraint details without repetition or fluff. The structure is well organized for a complex tool.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description fully addresses the tool's complexity: atomicity, operation types, orphan guard behavior, and mapping-only effects. An output schema exists, so return values are covered elsewhere. The only minor gap is server_version, which is neither described in the schema nor the description.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 75%, so most parameter meaning is already provided by the schema. The description adds valuable context about operation behaviors but does not significantly elaborate on individual parameter semantics beyond what the schema already includes. server_version remains undocumented, but the high schema coverage keeps this at baseline.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource: 'Apply a batch of control operations atomically as ONE transaction.' It clearly distinguishes this batch tool from single-operation siblings like remap_control and delete_control by emphasizing atomic, multi-operation changes.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly states when to use it: to reorganize controls in one step, such as deduplication, 'instead of many separate calls.' It does not explicitly name alternative tools or exclude single-operation cases, but the context is clear enough for an agent to choose correctly.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

assess_modelAssess ModelA

Run the deterministic assurance assessment over a threat model. Read-only — no LLM calls, no mutation.

Evaluates each control objective from its controls' implementation status and returns summary counts (mitigated / at_risk / unassessed) plus progressive metrics (defined / implemented / verified). For LLM-based reasoning about which COs are under-covered and what controls to add, use check_control_gaps instead.

Use summary_only=True to get just the counts without per-CO assessments.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNoMax control objectives to return (0 = all).
offsetNoSkip the first N control objectives.
statusNoOptional filter — "mitigated", "at_risk", or "unassessed".
model_idYesID of the threat model to assess.
summary_onlyNoIf True, return only summary counts (no per-CO details).
server_versionYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries full responsibility for behavioral disclosure. It explicitly states the tool is read-only, makes no LLM calls, performs no mutation, and is deterministic. It also describes what the assessment computes and returns, including summary counts and progressive metrics. This goes well beyond the structured schema and gives the agent a confident mental model of the tool's behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is well-structured and front-loaded: purpose first, behavioral safety second, output summary third, alternative routing fourth, and a practical usage tip last. Every sentence adds value and there is no redundant text repeating schema fields.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description is largely complete given the output schema exists and covers return values. It explains purpose, safety, result semantics, and the alternative tool. The only notable gap is that server_version is a required parameter with no schema description and no explanatory mention in the description, which is a small but real completeness gap for a required input.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 83%, so the schema already documents five of six parameters well. The description adds useful context for summary_only and reinforces the allowed status values. However, it does not help with server_version, the one parameter lacking schema documentation, and it adds no deeper semantics for limit, offset, or model_id beyond what the schema already provides.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: 'Run the deterministic assurance assessment over a threat model.' It clearly explains what the tool evaluates and returns, and distinguishes itself from check_control_gaps by naming that sibling explicitly. An agent can tell exactly what assess_model does and what it does not do.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives explicit routing guidance: for LLM-based reasoning about which control objectives are under-covered and what controls to add, use check_control_gaps instead. It also explains when to use summary_only=True. This is clear when-to-use and when-not-to-use guidance, not just implied context.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

assign_to_componentsAssign To ComponentsA

Replace an asset's or a control's component scope. Mutating.

Components are the canonical bridge between security architecture (trust boundaries) and code organization (repos). target_type selects what is being scoped:

  • "control" — replace a control's component scope. A control scoped to one or more components is visible to coding agents working in those repos (matched via Component.repo_url + Component.path); an unscoped control is visible everywhere. Use when wiring a previously unscoped control to the component(s) that implement it, adding a second component to a cross-cutting control (e.g. "all microservices enforce JWT validation"), or correcting a wrong assignment. target_id is the control ID (e.g. "CTRL-03").

  • "asset" — replace an asset's component scope. Linking assets to components flows boundary context into reachability derivation without giving Asset its own trust_boundary_ids. Multi-component is the right shape for a multi-instance asset (e.g., a session token on client + cache + DB — each component handles a distinct instance). target_id is the asset ID (e.g. "A1").

Both variants are mechanical / non-AI-gated and validate only that every referenced component exists on the model.

ParametersJSON Schema
NameRequiredDescriptionDefault
model_idYesID of the threat model.
target_idYesID of the asset or control to scope (must match ``target_type``).
target_typeYesEither "asset" or "control" — which entity to scope.
change_reasonYesWhy this scope is appropriate (min 10 chars). Captured in the version history.
component_idsYesComma-separated component IDs (e.g., "CMP1,CMP2"). Empty string = unscoped (a control becomes visible to every coding agent; an asset loses its explicit code-ownership binding). Every supplied ID must exist on the model.
server_versionYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does so well: it flags the operation as 'Mutating,' states both variants are 'mechanical / non-AI-gated,' and discloses the only validation rule. It also explains downstream effects like visibility to coding agents and boundary context flowing into reachability derivation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The core action and mutation flag are front-loaded, and the two variants are organized in clean bullets. Every sentence earns its place, with concrete examples and domain rationale that support correct invocation rather than padding.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Since an output schema exists, return-value documentation is not required. The description covers both target types, appropriate use cases, replacement semantics, unscoping via empty string, and validation rules, making it complete enough for an agent to call the tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is already high at 83%, and the description still adds meaningful value for target_type, target_id, and component_ids with examples, empty-string behavior, and multi-instance asset guidance. It leaves server_version undocumented in both schema and description, which prevents a 5.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource: 'Replace an asset's or a control's component scope.' The two target_type variants are fully spelled out with examples, so an agent can distinguish this from siblings like add_component or edit_component.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives explicit 'Use when' scenarios for the control variant: wiring an unscoped control, adding a second component to a cross-cutting control, and correcting a wrong assignment. It explains the asset variant's purpose for reachability, but it does not name when-not-to-use conditions or alternative sibling tools.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

associate_functional_testAssociate Functional TestA

Associate a functional test with one or more functional objectives.

Use this after suggest_functional_test_mappings, or to hand-map a test to the objectives it covers. The platform verifies each association is applicable before accepting it and returns any it declined under rejected_mappings.

ParametersJSON Schema
NameRequiredDescriptionDefault
model_idYesID of the threat model.
server_versionYes
functional_test_idYesThe functional test to associate.
functional_objective_idsYesComma-separated objective ids the test covers.

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the behavioral burden. It discloses that the platform verifies each association before accepting it and returns declined mappings under 'rejected_mappings'. This is useful behavioral context, though it does not cover side effects or idempotency.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three concise sentences, each earning its place: what it does, when to use it, and what behavior to expect. No redundancy or filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with four required parameters and an output schema, the description provides essential operational context, usage timing, and behavior. It does not explain 'server_version', but the output schema and remaining context make the tool sufficiently usable.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 75%, so the schema already documents most parameters. The description adds some context about the association intent and the rejection behavior, but it does not compensate for the undocumented 'server_version' parameter or add deeper meaning to the parameter values.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific operation: associating a functional test with one or more functional objectives. It names the related sibling tool suggest_functional_test_mappings, which helps distinguish this tool from other mapping-related tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives explicit usage timing: 'Use this after suggest_functional_test_mappings, or to hand-map a test to the objectives it covers.' This clearly explains when the tool is appropriate, though it does not enumerate exclusions or alternatives.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

attach_foundationAttach FoundationA

Delegate this model's objectives to a foundation's controls, in bulk.

Without selections it is read-only: it returns candidate (objective ↔ provider control) pairs with a match score, and nothing is created or credited. Show them to the operator.

With selections — a list of {"source_objective_id": ..., "provider_control_id": ...}, typically the confirmed subset of those candidates — it is mutating: each becomes a delegated draft edge that runs LLM validation and carries no credit until confirmed with manage_reliance(action="confirm"). Returns {created, failed}.

ParametersJSON Schema
NameRequiredDescriptionDefault
model_idYesthe consumer model.
selectionsNothe pairs to delegate; omit to get the candidates.
server_versionYes
foundation_model_idYesthe foundation to delegate to.

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and discharges it well: it discloses that the no-selections path is read-only and creates/credits nothing, that the selections path is mutating, that edges start as 'delegated' drafts, that LLM validation runs, that no credit attaches until confirmation, and that the return is {created, failed}. This is unusually rich behavioral disclosure for a mutation tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The dual-mode behavior is front-loaded and clearly partitioned by the 'Without selections' / 'With selections' structure, so every sentence earns its place. It runs a bit long, but the length is driven by necessary mutation and credit semantics rather than padding.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 4-parameter tool with a nested-ish selections payload and an output schema, the description covers the read-only vs mutating split, the credit/validation lifecycle, and the confirmation path, and it does not need to restate return values since an output schema exists. Nothing an agent needs to call it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 75% and the schema already labels model_id, foundation_model_id, and selections. The description adds real meaning beyond the schema for 'selections' by giving the exact pair shape ({source_objective_id, provider_control_id}) and describing the typical value (the confirmed subset of candidates). server_version remains unexplained in both places, keeping this short of a 5.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific verb and resource ('delegate this model's objectives to a foundation's controls, in bulk') and immediately distinguishes the tool from siblings like declare_foundation and manage_reliance by framing it as a bulk attachment operation. An agent can tell what it does without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives explicit when-to-use guidance for both modes: omit 'selections' to get read-only candidate pairs, pass 'selections' to create delegated edges. It also names the follow-up tool ('manage_reliance(action="confirm")') needed to carry credit, which is exactly the routing an agent needs.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

auto_map_controlsAuto Map ControlsA

LLM-map a model's existing controls to a framework's requirements. Requires PRO tier. Mutating: writes control-to-requirement mappings. Runs as a background job (typically 20-45s); this tool waits for completion and returns the result.

Sits between the manual map_control_to_requirement (one mapping at a time) and the full auto_remediate_compliance loop (which also excludes non-applicable requirements and proposes new entities for remaining gaps). auto_map_controls only creates mappings from controls that already exist — it never adds or excludes entities.

ParametersJSON Schema
NameRequiredDescriptionDefault
model_idYesID of the threat model.
control_idNoOptional single control ID to map; omit to map all of the model's controls.
framework_idYesID of the compliance framework.
server_versionYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure. It clearly states the operation is mutating, runs as a background job, waits for completion, returns results, requires PRO tier, and only maps existing controls. This is strong transparency for a complex asynchronous write operation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is front-loaded with the core action and constraint, followed by concise behavioral details and then sibling differentiation. Every sentence contributes useful information with minimal jargon or redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Despite the tool's complexity — async, mutating, tier-gated, with multiple sibling tools — the description covers what it does, how it behaves, its scope limitations, and its relationship to alternatives. The presence of an output schema means return-value detail is not required here. This is a complete and actionable description.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 75%, so the schema already documents model_id, control_id, and framework_id. The description adds context about 'existing controls' and 'never adds or excludes entities,' which indirectly clarifies scope, but it does not add meaningful per-parameter semantics beyond the schema, especially for server_version.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb and resource: 'LLM-map a model's existing controls to a framework's requirements.' It clearly distinguishes itself from siblings by naming map_control_to_requirement and auto_remediate_compliance and explaining how auto_map_controls differs from both.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicit usage placement is provided: it 'sits between' the manual single-mapping tool and the full auto-remediation loop. It also states exactly what it does not do — it never adds or excludes entities — so an agent can decide when to prefer it over alternatives.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

auto_remediate_complianceAuto Remediate ComplianceA

Automatically close compliance gaps for a framework. Requires PRO tier.

Three-phase loop: (1) auto-map existing controls to unmapped requirements, (2) exclude requirements for non-applicable taxonomy primitives, (3) suggest and apply new assets/attackers for remaining gaps.

Phase (3) routes every proposal whose name matches a soft-deleted asset/attacker through the same restore-candidate LLM gate add_asset uses, so reanimating a previously removed entity reinstates its stable ID and every CO tombstone + control tied to it (rather than spawning a duplicate fresh ID). The response distinguishes assets_added / attackers_added (genuinely new) from assets_restored / attackers_restored (revived soft- deletes) and lists restored_asset_ids / restored_attacker_ids. Proposals the gate classified as similar (or that fail-closed on an unavailable / malformed gate response) appear under skipped with a per-entry reason — the operator decides whether to restore manually or rephrase.

Converges automatically: stops when fully covered or when no further progress can be made.

This runs automatically when a framework is selected, but can be re-triggered manually if the model changes.

ParametersJSON Schema
NameRequiredDescriptionDefault
model_idYesID of the threat model.
framework_idYesID of the compliance framework.
server_versionYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.3/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and delivers exceptionally. It discloses the restore-candidate LLM gate shared with add_asset, the side effect of reinstating stable IDs along with CO tombstones and tied controls, fail-closed behavior when the gate is unavailable, per-entry skip reasons, response field distinctions (restored vs newly added), and automatic convergence.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is long but information-dense, with purpose and prerequisite front-loaded before the phase breakdown. The phase-3 paragraph is the most verbose section, yet nearly every clause carries operational significance for a tool with this side-effect complexity. Minor tightening of the restore-candidate explanation would be possible without losing value.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a complex mutation tool with zero annotations, the description covers prerequisites, the full operational loop, side effects on soft-deleted entities, failure modes, concurrency behavior, and trigger conditions. The output schema covers return values, so the response-field explanation is bonus. The one real gap is server_version, a required parameter whose meaning appears nowhere in the schema or description.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 67%: model_id and framework_id are documented in the schema, and the description adds contextual meaning by tying framework to the remediation target and noting that model changes justify re-triggering. However, server_version is a required parameter with no schema description and no description-side compensation, leaving its semantics entirely unknown.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The opening sentence states a specific verb and scope: "Automatically close compliance gaps for a framework." The three-phase loop then concretely defines what 'close' means (map controls, exclude non-applicable requirements, apply new assets/attackers), which clearly differentiates it from mapping-only siblings like auto_map_controls and check_control_gaps.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives clear invocation context: it runs automatically when a framework is selected, can be manually re-triggered when the model changes, and requires PRO tier. However, it never explicitly names alternatives or states when not to use this tool (e.g., when only control mapping is needed versus full remediation with asset/attacker creation), leaving some routing inference to the agent.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

check_control_gapsCheck Control GapsA

Analyze control coverage and surface control objectives that lack sufficient controls. Read-only (does not mutate the model); runs as a polled background job and uses LLM reasoning.

Complements the deterministic assess_model (which scores each CO's mitigated / at_risk / unassessed status from control implementation state) by reasoning about which COs are under-covered and where new controls are needed. Use this to decide what controls to add; use assess_model to score the current state.

ParametersJSON Schema
NameRequiredDescriptionDefault
model_idYesID of the threat model.
server_versionYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.4/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden, and it does so well: it explicitly states the tool is read-only, does not mutate the model, runs as a polled background job, and uses LLM reasoning. These are meaningful behavioral traits beyond what the schema alone reveals.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is compact and front-loaded: the first sentence states the core purpose, the second covers behavior, and the final paragraph distinguishes the tool from its sibling. No sentence is wasted.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description covers purpose, behavior, async execution, and the key alternative, and an output schema exists to cover return values. The main gap is the unexplained server_version parameter, and the description does not mention how to obtain or poll the background job's result.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is only 50%: model_id has a description but server_version does not. The description adds no parameter-level guidance at all, leaving server_version's meaning and format unexplained. It does not compensate for the low schema coverage.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses specific verbs ('Analyze', 'surface') tied to a clear resource: control coverage and control objectives lacking sufficient controls. It also explicitly differentiates itself from assess_model, making its purpose unmistakable.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives explicit guidance: used to decide what controls to add, while assess_model is used to score the current state. It also contrasts its LLM-based reasoning with assess_model's deterministic computation, so an agent can choose correctly between the two.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

classify_model_cweClassify Model CweA

Classify a model's control objectives against the platform CWE catalog.

Grounded: the model may only select from the catalog's current-version candidate ids, and every returned id is re-validated against the catalog before storage — a hallucinated or deprecated id is never persisted. Skips control objectives already tagged at the catalog's current version unless force is set. Returns a summary: {status, catalog_version, cos, classified, tags_written, skipped}. 404s if CWE classification is not enabled on this instance.

ParametersJSON Schema
NameRequiredDescriptionDefault
forceNore-classify control objectives even if already tagged at the catalog's current version (default false).
model_idYesID of the threat model to classify.
server_versionYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden, and it delivers: grounding constraints, re-validation before persistence, skip-if-already-tagged behavior, force semantics, the summary shape, and a 404 failure condition.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is compact and front-loaded with the core purpose, and each subsequent clause adds distinct behavioral information without redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description covers return summary, idempotence, and errors, and an output schema exists, so the main gap is the unexplained required 'server_version' parameter. It is otherwise sufficiently complete for a tool of this complexity.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 67% and the missing description is for the required 'server_version' parameter. The description reinforces 'force' and 'model_id' but adds no meaning for 'server_version', so the agent still cannot tell what value to supply.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The first sentence names a specific verb ('Classify'), a concrete resource (a model's control objectives), and the target catalog, which makes the tool's job unmistakable and distinguishes it from sibling reads like get_cwe_catalog and get_model_cwe_tags.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description makes clear this is the tool to use when classifying against the CWE catalog, and the skip/force behavior implies idempotent reclassification. However, it never explicitly states when not to use it or names an alternative for retrieving catalog IDs or existing tags.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

complete_setup_stepComplete Setup StepA

Mark one onboarding setup step as done. Mutating: updates the workspace onboarding checklist. Call after actually performing the corresponding setup action on the user's behalf.

Check current progress with get_setup_status first to avoid re-marking completed steps. An unrecognized step_id is rejected without any state change.

ParametersJSON Schema
NameRequiredDescriptionDefault
step_idYesThe step to mark complete, one of "mcp_configured", "mipiti_verify_installed", "ci_secret_added", "ci_pipeline_added".
server_versionYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden. It discloses that the operation is mutating, updates the checklist, and rejects unrecognized step_id values without changing state. This is strong behavioral disclosure, though it stops short of discussing permissions, idempotency on already-completed steps, or side effects beyond the checklist.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is compact and front-loaded, with the core purpose stated in the first sentence. The subsequent sentences provide actionable usage guidance and an important failure behavior without wasted words.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description covers purpose, mutation semantics, call timing, and invalid-input behavior, and an output schema exists to document return values. However, the required server_version parameter is left entirely unexplained, which is a meaningful gap for an agent trying to invoke the tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema describes step_id and its allowed values, but server_version has no schema description and the tool description does not explain it either. With only 50% schema coverage, the description should compensate for the undocumented required parameter, but it never mentions server_version at all.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the specific action ('Mark one onboarding setup step as done') and the resource it affects (workspace onboarding checklist). It also explicitly identifies the operation as mutating, which distinguishes it from read-only sibling tools like get_setup_status.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives explicit when-to-use guidance: call after performing the actual setup action. It also instructs the agent to check get_setup_status first to avoid re-marking completed steps, providing a clear workflow and an alternative tool reference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

convert_assumption_to_controlsConvert Assumption To ControlsA

Convert a violated or retired assumption to controls. Mutating.

Retires the assumption's CO linkage and, when any CO it covered is left with no control, proposes a control build for those COs: the result's proposal (null when nothing is owed) is started with start_control_build after review, and authors the controls then. Nothing is authored by this call. Use when an assumption is no longer valid and the system owner needs to implement controls instead.

Side effect on control-level linkage: this assumption is also removed from every assumption_groups entry on every control that referenced it. Any group left empty by the removal is dropped; a control's status is not changed, and a control left with no group is no longer backed by an assumption. Underlying assumptions are not deleted — only the linkages.

ParametersJSON Schema
NameRequiredDescriptionDefault
model_idYesID of the threat model.
assumption_idYesID of the assumption to convert.
server_versionYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.4/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and delivers extensively: it states nothing is authored here, describes the CO-linkage retirement, the proposal behavior (null when nothing owed), and detailed side effects on assumption_groups, including that empty groups are dropped, status is unchanged, and underlying assumptions are not deleted.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core action and mutability flag before the detail. The body is comparatively long but each sentence conveys a distinct behavioral fact (proposal handling, group side effects, non-deletion), so little is wasteful.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a complex mutating tool, it covers the operation, the deferred build workflow, and all linkage side effects thoroughly. An output schema exists, yet the description helpfully explains the key 'proposal' result field, leaving no material gap for correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 67%; model_id and assumption_id are documented in the schema while server_version is not. The description adds no parameter-level syntax or format detail, so it does not fully compensate for the coverage gap, though the remaining params are largely self-explanatory.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('Convert a violated or retired assumption to controls') and immediately flags its nature ('Mutating.'). It is clearly distinguishable from sibling tools like add_assumption, edit_assumption, and delete_assertion.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides explicit when-to-use guidance ('Use when an assumption is no longer valid and the system owner needs to implement controls instead') and directs the agent to the follow-up tool (start_control_build after review). It does not name when-not to use it or a direct alternative, but the context is clear.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

create_co_dispositionCreate Co DispositionA

Record that a control objective DOES NOT APPLY to this system — a signed, expiring judgment, not a dismissal.

The sibling of create_risk_acceptance, and the distinction between them is the claim being made. A risk acceptance says the exposure is real and we are carrying it. A disposition says this objective does not apply here at all — the asset is not handled the way the objective assumes, the attacker position does not exist in this deployment, the capability is not present.

The objective is not removed. It stays in the control-objective matrix, stays in every coverage count, and is reported in its own class alongside the owner and justification recorded here. That is the point: a reviewer can see the judgment and challenge it. An objective that simply vanished would be indistinguishable from one nobody modelled.

What it does change is work: no controls are generated for the objective and no coverage gap is raised against it, because an objective that does not apply is not a gap.

review_by is required and is not a formality — the claim stops applying on that date, and the objective returns to whatever posture its controls give it, gap included. Choose a date by which someone can realistically re-check that the claim still holds.

Use create_risk_acceptance instead when the objective DOES apply and the exposure is being carried deliberately. If an objective is only unaddressed rather than inapplicable, neither tool is right — add controls.

ParametersJSON Schema
NameRequiredDescriptionDefault
ownerYesWho owns the judgment (name / role). They answer for it.
model_idYesID of the threat model.
review_byYesISO 8601 date the claim expires (e.g. "2027-02-06T00:00:00Z").
justificationYesWhy the objective does not apply to this system.
server_versionYes
control_objective_idYesThe objective being declared not applicable.

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of disclosing behavioral traits. It does this thoroughly: the objective remains in the matrix, no controls are generated, no coverage gap is raised, and the claim expires at review_by. This goes far beyond what the schema alone communicates.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is front-loaded with the core claim and then structured into clear behavioral, temporal, and alternative-usage sections. It is longer than strictly necessary, but the length is justified by the nuance of the tool and the need to prevent confusion with create_risk_acceptance.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given 6 required parameters, no annotations, and an existing output schema, the description covers the tool's semantics, side effects, timing constraints, and sibling distinctions thoroughly. An agent has enough information to decide when to call it and what the call will do.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is high at 83%, so the baseline is already solid. The description adds meaningful semantics beyond the schema, especially for review_by: it is 'not a formality,' the claim stops applying on that date, and the date should be realistic for re-checking. It also reinforces that owner is the accountable party.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource: 'Record that a control objective DOES NOT APPLY to this system.' It also explicitly distinguishes this from create_risk_acceptance by defining the type of claim being made, so an agent can tell the tools apart without needing to compare schemas.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives explicit routing guidance: use create_risk_acceptance when the objective DOES apply and the risk is carried deliberately, and use neither tool when an objective is merely unaddressed rather than inapplicable. It also explains the operational consequences of using this tool, giving clear context for when it is appropriate.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

create_groupCreate GroupA

Create a tag, optionally with its first members. Mutating.

A tag groups models for viewing and reporting without asserting any relationship between them and without moving credit. Names are unique within the workspace (409 on a clash). Every model named in model_ids must be one the caller can access in this workspace.

ParametersJSON Schema
NameRequiredDescriptionDefault
nameYesthe tag name (unique within the workspace).
model_idsNooptional initial member model ids.
descriptionNooptional description.
server_versionYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, yet the description carries the full burden well: declares 'Mutating', warns uniqueness is per-workspace with 409 on a clash, and adds an access precondition ('every model named in model_ids must be one the caller can access'). That is rich behavioral detail beyond what the schema encodes.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the action and the core distinguishing property, then the necessary operational facts (uniqueness, 409, access check). Every sentence earns its place; no padding.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 4-param mutating tool with no annotations, the description covers mutation semantics, uniqueness/conflict behavior, and preconditions, and an output schema exists so return values need not be explained. It omits any mention of the server_version parameter or required-permission specifics, leaving a small gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 75% and already documents name/model_ids/description. The description reinforces semantics — uniqueness scope for name and the access precondition for model_ids — adding value the schema does not convey, though it says nothing about model_ids being nullable or server_version.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (Create) and resource (a tag/group) and immediately disambiguates the concept: 'A tag groups models for viewing and reporting without asserting any relationship.' This distinguishes create_group from relationship-asserting siblings like attach_foundation or add_model_to_group.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Clarifies intent — grouping for viewing/reporting, not relationship assertion — which implicitly tells an agent when this is the right tool versus modeling relationships. It does not explicitly name an alternative or when-not-to-use sibling, so it falls short of a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

create_proposalCreate ProposalA

Raise a proposal to change a model's scope or design. Call this when the code or your analysis says the model should gain or lose a component, or that an attacker position or asset should be removed by design; do not edit the model directly for those changes. Mutating: persists a proposal record.

A proposal is a change of scope or design. Raising one is not deciding it: a person (or an agent under a delegation rule that names the decision) decides it with decide_proposal. Design changes are never applied automatically. Poll list_proposals for the outcome.

ParametersJSON Schema
NameRequiredDescriptionDefault
kindYesOne of ``add_component`` (payload ``{name, repo_url?, path?, trust_boundary_ids?}``), ``remove_component`` (payload ``{component_id}``), ``design_change`` (payload ``{target_kind: "attacker"|"asset", target_id, design_move}``; take ``design_move`` from ``get_design_leverage``), ``assumption`` (payload ``{co_id, group_id, precondition, assumption_id?, gap?}``: a precondition about the environment that only something outside this system can meet, one declarative sentence; ``assumption_id`` names an existing assumption that states it). A precondition a person already rejected for that objective is refused.
payloadYesJSON object string with the fields for ``kind``.
evidenceNoOptional JSON object string, e.g. ``{paths: [], symbols: [], note: ""}``, pointing at what you saw.
model_idYesID of the threat model.
rationaleYesWhy this change is right (what in the code or design supports it).
server_versionYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the disclosure burden and does so well: it states this is mutating ('persists a proposal record'), that raising is not deciding, that design changes are never applied automatically, and that a person or delegated agent decides via decide_proposal. It stops short of stating explicit auth/permission requirements or any limits, so it is strong but not exhaustive.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two compact paragraphs, front-loaded with the action and the trigger condition, then the lifecycle note. Every sentence earns its place; the only minor cost is slight repetition between 'Raise a proposal...' and 'A proposal is a change of scope or design.'

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be explained, and the description covers the workflow end-to-end: what the tool does, when to call it, that it is not the decision step, that changes are not auto-applied, and how to poll for the outcome. Nothing an agent needs to call it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 83%, so the schema already documents kind, payload, evidence, model_id, rationale and their formats (including the per-kind payload shapes). The description adds conceptual framing for what a proposal is but no syntax or constraint detail beyond the schema, so the baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('Raise a proposal to change a model's scope or design') and explicitly scopes the kinds of changes covered (gain/lose a component, remove an attacker position or asset). It also distinguishes itself from direct model editing and from the decision step, so an agent can place it among siblings like decide_proposal and list_proposals.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives explicit when-to-use triggers ('when the code or your analysis says the model should gain or lose a component...'), an explicit when-not ('do not edit the model directly for those changes'), and names the alternatives for the follow-up step (decide_proposal, list_proposals). This is the full when/when-not/alternative triad.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

create_risk_acceptanceCreate Risk AcceptanceA

Record that an operator explicitly ACCEPTS the residual risk on a control objective instead of mitigating it — the write counterpart to list_risk_acceptances.

Use when a control objective's residual risk is a deliberate, documented decision rather than an unaddressed gap: the acceptance carries an owner, a justification, and a review deadline, and reads as active until it expires or is revoked. Prefer this over leaving a known-and-accepted risk implicit — it makes the decision auditable and forces a revisit by the deadline. An accepted objective is still surfaced (as accepted, not unaddressed) when triaging at-risk objectives.

ParametersJSON Schema
NameRequiredDescriptionDefault
ownerYesWho owns the acceptance (name / role).
model_idYesID of the threat model.
review_byYesISO 8601 date to revisit the acceptance (e.g. "2027-02-06T00:00:00Z").
justificationYesWhy the risk is accepted (the rationale of record).
server_versionYes
control_objective_idYesThe control objective whose residual risk is accepted.

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden and discharges it well: it discloses the lifecycle ('reads as active until it expires or is revoked'), the record contents (owner, justification, review deadline), and the triage effect (surfaced as accepted, not unaddressed). This adds real behavioral context beyond the bare mutation implied by 'create'. It stops short of covering permissions or revocation mechanics, which keeps it at a 4.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short paragraphs, front-loaded with the core purpose in the first sentence. Each subsequent sentence contributes a distinct fact — usage condition, lifecycle, triage effect — rather than restating schema fields. It is longer than a minimal definition but every sentence earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a mutation tool with six required parameters and zero annotations, the description covers purpose, when to use it, record contents, lifecycle, and downstream visibility; the output schema relieves it of explaining return values. The remaining gaps — who is authorized to accept and what server_version should contain — are minor against the strong selection and invocation guidance.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 83%, which is high, so the baseline is 3; the description adds collective framing by naming owner, justification, and review deadline as the components of an acceptance and linking review_by to expiry. It does not add per-parameter detail beyond the schema and leaves server_version unexplained, so it does not fully close that gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource — 'Record that an operator explicitly ACCEPTS the residual risk on a control objective instead of mitigating it' — and self-identifies as 'the write counterpart to list_risk_acceptances', which separates it from the read sibling. The behavior is unambiguous and goes well beyond a paraphrase of the title.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives an explicit trigger condition ('Use when a control objective's residual risk is a deliberate, documented decision rather than an unaddressed gap') and tells the agent to prefer this over leaving accepted risk implicit, citing auditability and the forced review deadline. It names the read counterpart but does not name an alternative tool for the unaddressed-gap case, so the when-not is implied rather than fully explicit.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

decide_proposalDecide ProposalA

Accept or reject a proposal. Call this only when the workspace's delegation policy names this decision for this agent at the proposal's tier (the delegation block of get_control_work_order says what you may decide). Mutating: closes the proposal and applies an accepted change.

This is a judgment. The call is refused with HTTP 403 and an escalation_id unless the delegation policy permits it; the refusal parks the decision for a person as a decision_request. Do not retry a refusal: report the escalation_id, poll list_proposals for the outcome, and continue other work.

Accepting an assumption proposal accepts the assumption: it is created if new, attested until expires_at and bound into the group that waited on it, which is then judged again. That is its own decision (assumption_accepted); a rule delegating proposal acceptance does not cover it. Rejecting one records the precondition as rejected for that objective, so it is not proposed again, and the objective keeps its gap.

ParametersJSON Schema
NameRequiredDescriptionDefault
noteNoOptional note recorded with the decision.
decisionYes``accept`` or ``reject``.
model_idYesID of the threat model.
expires_atNoISO 8601 date an accepted assumption lapses (e.g. "2027-03-29T00:00:00Z"). Applies to accepting an ``assumption`` proposal; omitted, the acceptance lasts a year.
proposal_idYesID of the proposal to decide.
server_versionYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations, so the description carries the full burden and does so: it declares the operation mutating ("closes the proposal and applies an accepted change"), describes the 403 + escalation_id refusal path, notes the decision is parked for a person as a decision_request, warns against retrying, and explains that accepting an assumption is itself a separate decision (assumption_accepted) not covered by a rule delegating proposal acceptance.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core action and the delegation gate, and every paragraph adds distinct information. It is dense and slightly long, but the length is justified by the authorization and refusal semantics.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be described. Combined with the delegation precondition, refusal handling, and per-kind acceptance semantics, an agent has everything needed to invoke this correctly and recover from a refusal.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 83%, so the schema already documents most fields. The description still adds real semantics beyond it: expires_at applies specifically to accepting an assumption proposal with a one-year default, and decision acceptance has downstream effects (attestation until expires_at, binding into the waiting group; rejection records the precondition as rejected so it is not re-proposed).

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Opens with a concrete verb pair and resource ("Accept or reject a proposal") and distinguishes itself from the sibling tools that create proposals (create_proposal) and read them back (list_proposals). The scope is unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

States explicitly when this may be called (only when the delegation policy names the decision at the proposal's tier, per get_control_work_order), what to do on refusal (report escalation_id, do not retry, poll list_proposals), and how acceptance differs by proposal kind. When/when-not and the follow-up path are all named.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

decide_reconciliation_candidateDecide Reconciliation CandidateA

Decide a reconciliation candidate from list_reconciliation_candidates. Mutating.

kind is assets, attackers or components; own_qid / inherited_qid are the pair's qualified ids ("child:A1", "parent:A1").

  • decision="apply" records that the descendant's own entity IS the inherited one: the own entity stays in the model and is left out of its composed view, so the inherited entity is canonical and credit keys on it. The record is dropped when the pair stops matching (an edit, a re-parent). A heuristic-tier candidate is refused unless confirm_heuristic=True acknowledges its structural divergence. The server re-validates the pair (400 when the model has moved since: refresh the list). Bumps the model version; returns {model, controls_carried, controls_orphaned, orphaned_control_ids}.

  • decision="reject" records "these are NOT duplicates" at org scope, so the pair leaves the active queue for everyone. Idempotent on the pair; no new version. Returns the record; keep its id.

  • decision="unreject" removes the rejection rejection_id (from list_reconciliation_candidates(disposition="rejected")), returning the pair to the queue. Returns {ok: true}.

503 where composition is not available on the backend.

ParametersJSON Schema
NameRequiredDescriptionDefault
kindNo
own_qidNo
decisionYes
model_idYes
rejection_idNo
inherited_qidNo
server_versionYes
confirm_heuristicNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description must carry the full burden, and it does: model version bumps, record dropping on edit/re-parent, idempotency of reject, server re-validation with a 400, and a 503 when composition is unavailable. These are exactly the side effects and failure modes an agent needs before calling a mutating tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with purpose and the mutating flag, then organized as per-decision bullets. Dense but every bullet carries operative information; minor redundancy in restating return payloads that the output schema already covers.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 3-way branching mutating tool with 8 undocumented parameters, the description supplies preconditions, side effects, error codes, and return shapes, more than enough for correct invocation. Nothing an agent needs is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With 0% schema coverage the description must document all 8 params; it explains kind, own_qid/inherited_qid with concrete qualified-id format examples, the decision enum, confirm_heuristic's role, and rejection_id's source. server_version and model_id are used but never explained, leaving a small gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('Decide') and resource ('reconciliation candidate'), explicitly ties it to the sibling tool list_reconciliation_candidates, and flags the mutating nature. An agent can immediately tell what this tool does and where its inputs come from.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Breaks out all three decision values with when-to-use semantics (apply = descendant IS inherited, reject = not duplicates at org scope, unreject = restore from rejected queue), names the alternative list call with the disposition filter for unreject, and states the confirm_heuristic precondition for heuristic-tier candidates.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

declare_foundationDeclare FoundationA

Mark a model as a shared foundation that advertises providable controls.

Mutating: records this model as a foundation and stores its advertised controls; other models can then delegate to them (see attach_foundation). A foundation is a shared service (auth, logging, a shared datastore) whose controls other models can rely on.

Each entry in provides advertises one of THIS model's controls as providable: {"control_id": "CTRL-07", "capability_label": "Validates session tokens", "description": "..."}. A capability always advertises a control (a proven mechanism), never an objective.

ParametersJSON Schema
NameRequiredDescriptionDefault
model_idYesID of the model to declare as a foundation.
providesYesList of advertised-control dicts. ``control_id`` is required per entry; ``capability_label`` and ``description`` describe what the control provides to consumers.
visibilityNoWho may delegate to this foundation. "workspace" (default) makes it discoverable to every model in the workspace; "explicit" limits it to models explicitly attached.workspace
server_versionYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.1/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It usefully discloses that the tool is mutating ('records this model as a foundation and stores its advertised controls') and defines what a foundation is, but says nothing about permissions, reversibility, idempotency, or whether re-declaring overwrites existing controls.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Purpose and mutation status are front-loaded, followed by supporting definition and the provides example. Every sentence is relevant, though the grammar and layering are slightly dense.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists so return values need not be described, and the description covers the mutation and the core provides semantics. For a mutating, no-annotation tool it is nearly complete, but missing any note on permissions, side effects on other models, or overwrite behavior.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 75%, and the description adds real meaning beyond the schema: it gives the exact shape of a provides entry and a key semantic constraint ('a capability always advertises a control (a proven mechanism), never an objective') that is not expressed in the schema. It does not elaborate on visibility or server_version, so not a 5.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('mark a model as a shared foundation that advertises providable controls') and clearly distinguishes this from the sibling attach_foundation, which is the consumer side of the relationship. An agent can tell what this tool does without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explains the delegation workflow and points to attach_foundation as the counterpart, giving clear context for when this tool applies. It lacks an explicit when-not to use it or any prerequisite/ordering guidance, so it falls short of the 5 bar.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

delete_assertionDelete AssertionA

Permanently delete a single assertion from a control or assumption. Mutating and destructive: the assertion record is removed, not soft-deleted, and its contribution to sufficiency/verification is dropped. It does NOT itself re-run verification; the deletion queues a background re-evaluation of the control's sufficiency, which a later get_sufficiency read reports once it lands.

Use to retract a claim that was submitted in error or that get_verification_report flagged as misaligned (off-topic for the control's current description). To add assertions use submit_assertions; to inspect them first use list_assertions. Only "own" assertions can be removed here — inherited assertions come from composed models and must be managed on their source model.

ParametersJSON Schema
NameRequiredDescriptionDefault
model_idYesID of the threat model.
control_idNoID of the control the assertion belongs to (omit if it belongs to an assumption).
assertion_idYesID of the assertion to delete.
assumption_idNoID of the assumption the assertion belongs to (omit if it belongs to a control).
server_versionYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden and does so richly: it discloses that deletion is permanent and not a soft-delete, that the assertion's contribution to sufficiency/verification is dropped, that it does NOT itself re-run verification, and that a background re-evaluation is queued and later surfaced via get_sufficiency. This is exactly the behavioral context an agent needs before a destructive call.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the destructive nature and scope, then the background-re-evaluation caveat, then routing/alternatives. Every sentence earns its place and no filler is present.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists so return values need not be described, yet the description still routes the agent to get_sufficiency for the queued re-evaluation result. For a destructive, ownership-constrained mutation tool, nothing an agent needs to call it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 80%, so control_id/assumption_id/model_id are already documented in the schema, and the description's 'from a control or assumption' largely restates that. It adds the ownership constraint (own vs inherited) which is behavioral rather than parameter-syntactic, so baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('Permanently delete') and a precisely scoped resource ('a single assertion from a control or assumption'), and immediately distinguishes itself from the sibling write tool submit_assertions and the read tool list_assertions. An agent can tell exactly what this does and what it is not.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives explicit when-to-use triggers ('retract a claim that was submitted in error or that get_verification_report flagged as misaligned'), names alternatives with their conditions (submit_assertions to add, list_assertions to inspect first), and adds a hard exclusion: only 'own' assertions can be removed, inherited ones must be managed at the source model.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

delete_controlDelete ControlA

Soft-delete a security control, optionally with a justification. Destructive (mutating): the control is retired, not permanently erased.

Blocks with HTTP 409 when the control is the ONLY control covering any control objective — removing it would leave that CO uncovered. Add a replacement control (or refine the threat model) before deleting.

ParametersJSON Schema
NameRequiredDescriptionDefault
reasonNoOptional justification recorded in the audit trail (recommended).
model_idYesID of the threat model.
control_idYesID of the control to delete.
server_versionYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of disclosing behavioral traits. It explicitly states that the operation is destructive/mutating, that it is a soft delete rather than a permanent erase, and that it can block with HTTP 409 under a specific coverage condition. This is substantial and goes well beyond a generic description, though it stops short of covering permissions or reversibility details.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is compact and front-loaded: the first sentence states the core action, the second clarifies destructiveness, and the third explains the blocking condition and required remedy. Every sentence contributes necessary operational information with no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's moderate complexity and the presence of an output schema, the description covers the essential operational context: what the operation does, its side effects, and the key failure mode. It does not explain server_version semantics or mention any other restrictions, but the core invocation context is sufficiently complete for an agent to use the tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 75%, so the schema already documents most parameters. The description adds context about the justification being optional and about the 409 condition, but it does not add meaningful detail for model_id, control_id, or server_version beyond what the schema provides. This is adequate given the high schema coverage.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource: 'Soft-delete a security control,' which immediately distinguishes this from other control-related tools. It also clarifies the semantic nuance ('retired, not permanently erased'), removing ambiguity about the lifecycle impact. This clearly separates it from hard-delete or update operations.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides clear context for when the tool will succeed or fail, including the HTTP 409 blocking condition and the prerequisite to add a replacement control or refine the threat model. It does not explicitly name alternative sibling tools or state when not to use the tool, but the practical guidance is strong enough to guide invocation correctly.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

delete_groupDelete GroupB

Delete a tag. Mutating; the member models are not affected.

The tag's framework selections, requirement exclusions and relevance data are deleted with it. Frameworks it propagated to its members stay selected on them.

ParametersJSON Schema
NameRequiredDescriptionDefault
group_idYesID of the tag to delete.
server_versionYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

B3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and largely does so: it declares the operation mutating, enumerates the collateral destruction (framework selections, requirement exclusions, relevance data), and explicitly bounds the blast radius (member models unaffected; propagated frameworks stay selected). It stops short of stating irreversibility or required permissions.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short sentences with the mutation status front-loaded and the destructive/who-is-affected split handled economically. No filler, though the opening 'tag' framing costs clarity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists so return values need no explanation, and the description usefully covers side effects for a no-annotation mutation. But it omits any handling of the undocumented server_version parameter and any irreversibility/permission note that a destructive tool warrants.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is only 50%, so the description should compensate but adds nothing about parameters. It never mentions group_id, and server_version is undocumented in both schema and description, leaving the agent without a value format or source for a required parameter.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The verb+resource pair is muddled: the tool is delete_group and its siblings all use 'group' (get_group, create_group, add_model_to_group), yet the description says 'Delete a tag.' An agent must infer that 'tag' and 'group' are the same entity, which weakens confident selection.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No when-to-use guidance, no prerequisites, and no routing to or away from siblings such as delete_threat_model, delete_control, or delete_assertion. The agent gets no signal about when deletion is appropriate versus other cleanup paths.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

delete_threat_modelDelete Threat ModelA

Delete a threat model and all associated data. Destructive and permanent — cannot be undone.

Mutating: removes the model along with every version, its controls, assertions, findings, attestations, and tag/reliance memberships. Reliance edges from other models that pointed at this one are invalidated, which can move those consumers' posture.

Confirm intent before calling. To keep a copy first, use export_report (scope="model", format="archive") (a self-contained, re-importable JSON archive). Returns {deleted: True, model_id}.

ParametersJSON Schema
NameRequiredDescriptionDefault
model_idYesID of the threat model to delete.
server_versionYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.3/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries full responsibility for behavioral disclosure. It clearly states the operation is destructive, permanent, and mutating; enumerates exactly what is removed; explains downstream effects on reliance edges and consumer posture; and documents the return value. This is exemplary transparency.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is front-loaded with the most critical fact ('Destructive and permanent'), then gives a concise but complete list of side effects, followed by usable guidance and the return shape. Every sentence earns its place; no filler or redundant restating of the name.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description is highly complete for a destructive operation: it covers permanence, affected data, indirect effects on other models, a backup path, and the response. The only notable gap is the unexplained required server_version parameter, which weakens end-to-end invocability.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 50%: model_id is described, but the required server_version parameter has no description and no enum or format guidance. The tool description adds nothing about either parameter, so an agent cannot confidently determine what server_version should be.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource: 'Delete a threat model and all associated data.' It sharply distinguishes this from sibling deletion tools by emphasizing the permanent, cascade-deleting scope, so an agent can tell it apart from delete_control, delete_assertion, delete_group, or remove_entity.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives clear operational guidance: confirm intent before calling and export an archive if a copy is needed. It does not explicitly enumerate when not to use this tool versus nearby siblings, but the destructive scope and backup advice effectively frame when it is appropriate.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

discard_control_buildDiscard Control BuildA

Discard a model's held control build. Mutating.

A build that is queued, deferred, paused or blocked holds the model without running. Discarding it drops everything it staged — the model keeps its published controls exactly as they were — releases the model, and proposes the build again so it can be started later. The credits it already consumed are not returned. A running build must be paused first (pause_control_generation, then wait for paused).

Returns one of:

  • {discarded: true, model_id, status: "discarded", proposal, status_detail}.

  • {discarded: false, http_status: 409, code: "pause_first", status} — the build is running; pause it first.

  • {discarded: false, http_status: 409, code: "not_held", status} — no build holds the model.

ParametersJSON Schema
NameRequiredDescriptionDefault
model_idYesID of the threat model whose held build to discard.
server_versionYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does so richly: it discloses what is destroyed (staged items), what is preserved (published controls untouched), side effects (model released, build re-proposed), cost (credits not returned), and a hard precondition (pause a running build first).

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the verb+resource and the mutating flag, then behavior, then return shapes. Every sentence is informative, though the enumerated return block is somewhat verbose given an output schema already exists.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a mutating tool with zero annotations, the description covers destruction semantics, side effects, credit cost, and preconditions, plus concrete success/error return codes. An agent has everything needed to call it correctly without opening the schema.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 50%: model_id is documented in the schema, server_version is not. The description adds no parameter-level meaning beyond the schema (it only refers to 'a model'), so it neither compensates for the coverage gap nor adds value. Baseline 3 for a two-param tool where the schema does the documented half.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('discard') and resource ('a model's held control build') and immediately qualifies the operation as mutating. It clearly separates itself from siblings like delete_control and start_control_build by scoping to a held (non-running) build.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly enumerates the applicable states (queued, deferred, paused, blocked) and the exclusion condition: a running build must be paused first, naming the alternative tool (pause_control_generation) and the wait condition. This is explicit when/when-not guidance with a named alternative.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

edit_assetEdit AssetA

Edit an existing asset. Only provided fields changed.

When changing identity fields, hold to the asset authoring contract: name the data/resource protected and its security property, not a mechanism — otherwise the result is flagged with a quality_warning (see add_asset). There is no status field to set.

The composed impact is server-derived from the factor fields; there is no way to set it directly. To change the rating, set factor values (the platform composes the new rating) and supply change_reason documenting the operator override of the LLM-generated factors. The reason is captured in the rating-revision audit trail.

LLM-gated on identity-bearing fields (name, description, security_properties). Factor and notes edits skip the gate.

Outcomes when identity fields change:

  • Accepted edit (LLM classifies as preserve) — normal envelope response.

  • Rejected edit (LLM classifies as replace / ambiguous) — {"accepted": False, ...}; nothing saved. Soft-delete + add-new instead.

Editing a soft-deleted asset is rejected — restore_entity (entity_type="asset") first. 503 on evaluator outage, 502 on malformed response, 400 when factor fields are sent without change_reason.

ParametersJSON Schema
NameRequiredDescriptionDefault
nameNoNew name (optional).
notesNoNew notes (optional).
asset_idYesID of the asset (e.g., "A1").
model_idYesID of the threat model.
descriptionNoNew description (optional).
blast_radiusNo"Isolated" | "Multiplicative" | "Cascading".
change_reasonNoRequired when any factor field is supplied — documents the operator override of LLM-generated factors for the audit trail.
recoverabilityNo"Trivial" | "Manageable" | "Permanent".
server_versionYes
usage_subscoreNo"None" | "Low" | "High".
impact_rationaleNoNew rationale (optional).
regulatory_scopeNo"None" | "Notification" | "Legal".
integrity_subscoreNo"None" | "Low" | "High".
security_propertiesNoComma-separated properties (optional).
availability_subscoreNo"None" | "Low" | "High".
confidentiality_subscoreNo"None" | "Low" | "High".

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries full behavioral burden and satisfies it: it discloses LLM gating on identity-bearing fields, the rejected-edit response shape, that nothing is saved on rejection, server derivation of impact, audit-trail capture, and specific 503/502/400 error conditions. This is far beyond a generic 'edit' statement.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is long but earns its length for a 16-parameter mutation tool: it is front-loaded with the core behavior, then uses bold labels, bullets, and error summaries to make the constraints scannable. There is no filler or repeated schema content.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity, the presence of an output schema, and no annotations, the description covers all decision-relevant behavior: partial-update semantics, LLM gating, audit trail, soft-delete handling, and error conditions. An agent has enough context to invoke it correctly without opening sibling definitions.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Despite 94% schema coverage, the description adds grouping semantics the schema alone does not convey: identity-bearing fields vs factor/notes fields, the no-status-field caveat, the requirement to send change_reason with factor fields, and the impossibility of setting impact directly. It clarifies the relationship between change_reason and the factor/rating parameters.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The opening sentence 'Edit an existing asset. Only provided fields changed' states a specific verb and resource with scope, and the rest clarifies it updates rather than creates (contrasting with add_asset). It is clearly distinct from sibling edit tools for other entity types.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives explicit routing and preconditions: soft-deleted assets require restore_entity first, rating changes require setting factor values plus change_reason, there is no status field to set, and rejected edits must be handled by soft-delete + add-new. It names alternatives (add_asset, restore_entity) and the conditions that select them.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

edit_assumptionEdit AssumptionB

Edit an assumption. Creates a new model version.

ParametersJSON Schema
NameRequiredDescriptionDefault
model_idYesID of the threat model.
descriptionNoNew description (omit to leave unchanged).
assumption_idYesID of the assumption to edit (e.g., "AS1").
linked_co_idsNoNew comma-separated CO IDs; replaces the existing linkage (omit to leave unchanged).
server_versionYes
clear_exclusionNoWhen True, removes the predicate entirely (the assumption becomes prose-only). Mutually exclusive with the exclusion_* params — if both are sent, the exclusion_* params win.
exclusion_co_idsNoComma-separated CO IDs the predicate matches explicitly; when non-empty, overrides the match fields. Supplying any exclusion_* param rewrites the whole predicate (unspecified fields default to "*").
exclusion_asset_idNo"*" or a concrete asset ID.
exclusion_attacker_idNoPredicate match — "*" wildcard or a concrete attacker ID.
exclusion_property_matchNo"C" | "I" | "A" | "U" | "*".
exclusion_attacker_vectorNoOne of "Network" | "Adjacent" | "Local" | "Physical" | "*".
exclusion_asset_component_idNo"*" or a concrete component ID.

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

B3.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full behavioral disclosure burden. It does disclose one meaningful side effect — editing creates a new model version — which is valuable versioning context that the schema does not state. However, it omits other behavioral traits like the predicate-rewriting semantics and the mutual exclusivity between clear_exclusion and exclusion_* params, both of which remain buried in the schema.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences with zero wasted words; the primary operation is front-loaded and the versioning side effect follows immediately. It is efficiently structured, though for a 12-parameter tool arguably under-specified rather than optimally sized.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The rich 92%-covered schema and the presence of an output schema carry most of the load, so an agent that reads the schema can likely invoke this tool correctly. However, the description alone offers no when-to-use guidance and does not explain the versioned workflow implied by the required server_version param. Adequate but with clear gaps.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 92%, so the baseline of 3 applies; the schema thoroughly documents the exclusion_* rewrite behavior, defaults, and mutual exclusivity. The description contributes no parameter-level meaning, and notably it does not clarify the one required param without a schema description (server_version), which is a minor missed opportunity.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

"Edit an assumption" is a specific verb-plus-resource statement that clearly identifies the operation and, by verb choice alone, distinguishes it from the sibling add_assumption. The second sentence, "Creates a new model version," adds precision about what an edit entails, though the description never explicitly contrasts it with sibling tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is given on when to edit an assumption versus adding one (add_assumption) or converting assumptions to controls (convert_assumption_to_controls). There are no prerequisites, exclusions, or contextual cues beyond the bare operation itself, leaving the agent to infer applicability.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

edit_attackerEdit AttackerA

Edit an existing attacker. Only provided fields changed.

When changing identity fields, hold to the attacker authoring contract: capability names the operations performable from the position ("From [position], the attacker can [operations] …"), not just access — otherwise the result is flagged with a quality_warning (see add_attacker).

The composed likelihood is server-derived from the factor fields; to change the rating, set factor values and supply change_reason for the audit trail.

LLM-gated on identity-bearing fields (capability, archetype, position). Factor and trust_boundary edits skip the gate.

503 on evaluator outage, 502 on malformed response, 400 when factor fields, surface_extent or attest_surface_extent are sent without change_reason.

Attesting surface_extent is a person's audited structural declaration, ledgered like a factor override and forking a model version: "whole" makes every objective the attacker appears in a for-all obligation. "point" is REFUSED where an asset on one of those objectives is implemented by several components and is not split-knowledge — reaching any one of them reaches the asset, so a narrowing to one named entry would not be true of it — and it never makes a clause whose own text is universal existential.

ParametersJSON Schema
NameRequiredDescriptionDefault
model_idYesID of the threat model.
positionNoNew position (optional).
archetypeNoNew archetype (optional).
capabilityNoNew capability (optional).
attacker_idYesID of the attacker (e.g., "T1").
attack_vectorNo"Network" | "Adjacent" | "Local" | "Physical".
change_reasonNoRequired when any factor field, ``surface_extent`` or ``attest_surface_extent`` is supplied — documents the operator override for the audit trail.
server_versionYes
surface_extentNo``"whole"`` (operations range over ANY entry of the interface reached) or ``"point"`` (one named entry). Supplying it attests it; requires ``change_reason``.
attest_positionNoOperator-attest the attacker's current position without changing it — e.g. to confirm a fully external attacker's empty crossed set so an objective blocked on an unpositioned attacker can be resolved. Pass ``true`` to attest.
user_interactionNo"None" | "Required".
attack_complexityNo"Low" | "High".
trust_boundary_idsNoComma-separated trust boundary IDs — the boundaries this attacker has crossed (its position). Replaces the existing set. Changing it operator-attests the position, which lets reachability trust it for a decisive verdict.
privileges_requiredNo"None" | "Low" | "High".
likelihood_rationaleNoNew rationale (optional).
attest_surface_extentNoRecord the extent already on the attacker as attested, without changing its value. Pass ``true`` to attest; requires ``change_reason``.
capability_prevalenceNo"Commodity" | "Targeted" | "Rare".

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries full behavioral burden and does so thoroughly: patch semantics, LLM gating, server-derived likelihood, quality_warning on contract violations, specific HTTP error conditions, and surface_extent attestation with forking and refusal rules. This goes far beyond a generic 'edit' description.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the essential edit semantics and dense with useful detail; no filler sentences. The surface_extent paragraph is long and intricate, but it encodes necessary refusal and forking semantics, so it earns its place despite being heavy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 17-parameter mutation tool with no annotations and an output schema present, the description covers all non-obvious behavior: patch semantics, gating, error responses, audit trail requirements, and attestation constraints. Nothing an agent needs to call it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 94%, but the description adds critical meaning beyond it: capability must name operations performable from the position, likelihood is server-derived and cannot be set directly, change_reason triggers and error conditions, and surface_extent attestation has deep structural consequences. Without this, several parameters would be ambiguous.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

First sentence, 'Edit an existing attacker. Only provided fields changed,' names the exact verb/resource and clarifies patch semantics. It is unambiguously distinct from sibling creation/editing tools and even references add_attacker for the authoring contract.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides clear operational context: edits an existing attacker, only supplied fields change, identity fields are LLM-gated, and change_reason is required for factor/surface_extent edits. It does not explicitly contrast with alternatives like add_attacker or edit_asset, but its scope is unmistakable and the 'see add_attacker' pointer helps.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

edit_componentEdit ComponentA

Edit a component's properties.

Per-component level grades are orthogonal axes — set whichever apply to the program the component is in scope for. Leave a field unset (None) to keep the current server-side value; backend treats absent fields as "unchanged".

ParametersJSON Schema
NameRequiredDescriptionDefault
ealNoCommon Criteria Evaluation Assurance Level (1-7). For components subject to CC certification.
nameNoNew name (empty = unchanged).
pathNoNew path (empty = unchanged).
model_idYesID of the threat model.
repo_urlNoNew repo URL (empty = unchanged).
target_slNoIEC 62443 target Security Level (1-4). For industrial / OT components that need a 62443 zone target.
fips_levelNoFIPS 140-3 Security Level (1-4) for the cryptographic module embedded in this component.
component_idYesID of the component (e.g., "CMP1").
server_versionYes
trust_boundary_idsNoNew trust boundary IDs (comma-separated, empty = unchanged).

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the behavioral burden. It meaningfully explains the non-obvious partial-update behavior ('backend treats absent fields as unchanged') and clarifies that grade fields are independent orthogonal axes. This goes well beyond a bare 'edit a component', though it does not cover permissions, validation, or side effects.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is compact and front-loaded with the core purpose, then adds only the essential behavioral nuances. The 'orthogonal axes' metaphor is slightly dense but earns its place by framing the partial-update model. It is not bloated.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description plus a 90%-covered schema is largely sufficient for calling the tool: required IDs are in the schema and partial-update semantics are stated. However, it lacks guidance on how this tool relates to sibling edit tools, and it leaves the required server_version parameter unexplained. The presence of an output schema mitigates some of the gap, but the overall guidance is not fully complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 90%, so the baseline is 3. The description adds valuable semantics beyond the schema by explaining that per-component grades are orthogonal and that leaving fields unset preserves server-side values. This clarifies how to handle the nullable grade fields and the empty-string fields in practice.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action ('Edit') and the resource ('a component's properties'), which distinguishes it from other edit_* siblings that target different resources. However, it does not explicitly differentiate itself from sibling tools by scope or use case, so it stops short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives solid operational guidance about partial updates ('absent fields are unchanged') and when to set grade fields ('set whichever apply'), but it does not explicitly say when to use this tool versus alternatives like edit_asset or edit_trust_boundary. Usage context is implied rather than stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

edit_evidenceEdit EvidenceA

Attach or detach an auxiliary evidence item (doc, link, artifact reference) on a control. Mutating.

Evidence is contextual metadata only: it does NOT count toward a control's implementation status, and removing it changes neither the status nor any assertion. Only assertions prove controls.

ParametersJSON Schema
NameRequiredDescriptionDefault
urlNoadd — optional file path or URL pointing at the artifact.
typeNoadd — "code", "test", "config", "document" or "link" (default "code").code
labelNoadd — human-readable description of the item (required).
actionYes"add" attaches an item; "remove" detaches one.
model_idYesID of the threat model.
control_idYesID of the control.
evidence_indexNoremove — zero-based position of the item in the control's ``evidence`` array (read it with ``get_controls(control_id=...)``; default 0, the first item).
server_versionYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and largely meets it: it declares the operation mutating and, more valuably, states the non-effect of removal on status and assertions - a semantic trait an agent could otherwise guess wrong. It still omits permission requirements, whether edits are auditable/undoable, and any failure modes.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Purpose is front-loaded in the first sentence, followed by a one-word mutation flag, then the semantic caveat. The caveat is slightly restated twice ('removing it changes neither the status nor any assertion' / 'Only assertions prove controls'), but both sentences carry distinct weight and nothing is padded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Eight parameters at 88% schema coverage plus an output schema mean structured fields handle details; the description supplies exactly the missing conceptual layer - that this is a mutation of contextual metadata with no effect on control implementation status - which is what an agent needs to decide whether to call it.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 88%, so the schema already documents the action enum, item type values, and the zero-based evidence_index. The description adds only a high-level gloss (attach vs detach, what an item is), which is marginal beyond the structured data, so baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a concrete verb pair (attach/detach), the resource (an auxiliary evidence item), the item kinds (doc, link, artifact reference), and the parent object (a control). The closing line distinguishes evidence from assertions, which is exactly the boundary that separates this tool from assertion-oriented siblings in the list.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The statement that evidence 'does NOT count toward a control's implementation status' and that 'only assertions prove controls' implicitly tells the agent when this tool is appropriate versus proof-bearing tools. However, there is no explicit when-to-use rule, no prerequisite call chain (e.g., read evidence_index via get_controls before removing), and no named alternative.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

edit_trust_boundaryEdit Trust BoundaryB

Edit a trust boundary. Creates a new model version.

ParametersJSON Schema
NameRequiredDescriptionDefault
tb_idYesID of the trust boundary (e.g., "TB1").
passesNoNew comma-separated AttackVector values the boundary allows through (subset of "Network,Adjacent,Local,Physical"). Use the empty string to set "blocks all"; omit to leave unchanged. Reach-relevant — narrowing or widening this set can flip CO verdicts.
sealedNoNew isolation flag. True declares NO lateral ingress (the only way in is crossing the perimeter — an air-gap / segmented enclave), which lets reachability decisively rule the boundary unreachable; False assumes a lateral pivot is possible. Reach-relevant — changing it can flip CO verdicts. Setting it records an operator attestation of the seal. Omit to leave unchanged.
crossesNoNew comma-separated asset IDs.
model_idYesID of the threat model.
descriptionNoNew description.
seal_sourceNo"attested" | "unattested". Only an operator-attested seal lets reachability decisively rule an objective unreachable past the boundary; an unattested (default/model-suggested) seal is treated as pivotable. Use "attested" to attest a boundary already marked sealed without re-toggling it; "unattested" retracts. An attested seal implies ``sealed``. Requires ``change_reason``.
change_reasonNoRequired when ``passes``, ``sealed``, or the seal attestation actually changes. Captured in the audit trail; documents why the boundary's vector filter, isolation claim, or attestation changed.
server_versionYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

B3.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the burden. It does disclose one important side effect—edits create a new model version—which is valuable behavioral context for a mutation tool. However, it does not mention the audit trail, change_reason requirements, attestation implications, or whether edits are reversible.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is only two sentences and front-loads the core purpose. The first sentence is somewhat redundant with the tool name, but the second earns its keep by revealing the versioning side effect.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 9-parameter mutation tool, this is barely adequate. The rich paramether schema and presence of an output schema compensate for much of the missing detail, but the description alone does not fully convey the implications of editing a trust boundary, such as versioning consequences and the audit/change_reason context.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 89%, so the schema already explains most parameter semantics. The tool description adds no parameter-level details, but with high schema coverage the baseline of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific verb and resource ('Edit a trust boundary') and adds a meaningful distinguishing behavior ('Creates a new model version'). It clearly identifies this as the editing counterpart to add_trust_boundary and other edit_* tools, though it does not explicitly name a sibling alternative.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool vs. add_trust_boundary or other editing tools. The verb 'Edit' implies use for existing trust boundaries, but no exclusions, alternatives, or context signals are given.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

export_reportExport ReportA

Export a threat model or a tag as a downloadable document. Read-only; no side effects on the source.

scope selects what is exported and how scope_id is read; format selects the representation:

  • scope="model" (scope_id = model id) supports format ∈ {csv, pdf, html, archive}:

    • csv — the model's current state rendered as CSV; returned inline as UTF-8 text in content.

    • pdf / html — rendered document returned base64-encoded in content_b64 (with content_type). Runs as a server-side job; progress is reported automatically while it completes, which may take time for large models.

    • archive — the self-contained, independently-verifiable JSON audit bundle of the model's current state: its latest version and controls, live assertions (with Tier 1 / Tier 2 verdicts and attested flags) and the runs behind them, open findings and those a person closed, risk acceptances and other decisions in force, assumption overrides, attestations, and instance sufficiency signatures; each control's per-clause evidence basis travels with it. Earlier versions, activity and chat are not in it. Those verdicts are the origin's record of what it claimed, which is what a third party checks against the signatures; an importing workspace credits what its own verification establishes (see import_threat_model_archive). Returned as {..., "envelope": <dict>}; feed the envelope to import_threat_model_archive to restore it into any workspace. Model scope only.

  • scope="tag" (scope_id = tag id) supports only format="html": the signed auditor report, every member model's report after the reliance edges among the members (each with its status and whether it credits its objective), in one HTML document returned inline in content. csv, pdf, and archive are rejected for tag scope.

ParametersJSON Schema
NameRequiredDescriptionDefault
scopeYesexport boundary — "model" or "tag".
formatNo"csv" (default), "pdf", "html", or "archive". Tag scope requires "html"; "archive" is model-only.csv
scope_idYesid of the model or tag selected by ``scope``.
server_versionYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does so: it declares read-only/no side effects, discloses that pdf/html run as a potentially slow server-side job with automatic progress reporting, specifies the exact return fields per format (content, content_b64 + content_type, envelope), and enumerates what the archive does NOT include (earlier versions, activity, chat).

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the purpose and safety note, then organized as a clean scope→format matrix. The archive bullet is dense and long, but every clause conveys behavior an agent needs (contents, exclusions, return shape, downstream tool); minor verbosity rather than padding.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, yet the description still pins down the format-dependent return keys, which is exactly the ambiguity an output schema alone would leave. For a read/export tool with two enums and a cross-field constraint, nothing an agent needs to call it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 75%, but the description adds substantial meaning beyond it: it explains that scope determines how scope_id is interpreted, which format values are legal per scope, and what each format's payload actually is. It documents the interaction the flat schema cannot express.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource (export a threat model or tag as a downloadable document) and immediately bounds scope to two entity types. It also names the counterpart tool (import_threat_model_archive), so an agent can place it in the workflow without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives explicit when-to-use mapping: which format applies to which scope, and explicitly states that csv/pdf/archive are rejected for tag scope. It routes to import_threat_model_archive for restoring an archive. It stops short of stating when to prefer a different export/report tool (e.g. get_compliance_report, get_verification_report), so not a full 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

generate_functional_objectivesGenerate Functional ObjectivesA

Derive capabilities, functional objectives, and the concrete tests to implement from the feature spec.

Capabilities are the behaviours the feature must deliver; each is walked against a taxonomy of operating conditions (nominal, boundary, invalid input, dependency failure, concurrency, …) to produce testable Given-When-Then objectives — and then a concrete, implementable test is specified for each objective (so the agent implements the tests rather than deciding what to test). Requires a Pro plan. Billable — may take some time. refresh=true re-derives from scratch, replacing prior generated (not manually authored) capabilities, objectives, and tests.

ParametersJSON Schema
NameRequiredDescriptionDefault
refreshNoRe-generate from scratch instead of serving cached output.
model_idYesID of the threat model.
server_versionYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full disclosure burden and does so well. It explains side effects: refresh=true re-derives from scratch and replaces prior generated content while preserving manually authored content. It also discloses billing, time cost, and the systematic taxonomy-driven process.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is front-loaded with the core purpose, then adds a precise definition of the internal concepts, and finishes with operational constraints. Every sentence earns its place, and the length is appropriate for the complexity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity and the presence of an output schema, the description covers the input requirements, workflow, output nature, side effects, and usage constraints. An agent has enough context to invoke the tool correctly and understand its consequences.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema describes model_id and refresh, and the description adds meaningful depth to refresh by specifying that it replaces prior generated (not manually authored) capabilities, objectives, and tests. However, server_version remains undocumented in both the schema and description, so coverage is not complete.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb, 'derive', and names the resource: capabilities, functional objectives, and concrete tests from a feature spec. It clearly distinguishes this generation tool from retrieval siblings like get_functional_objectives by emphasizing that it creates testable Given-When-Then objectives and implementable tests.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives useful context about when to use the tool (from a feature spec, requires Pro plan, billable, may take time) but does not explicitly contrast it with related alternatives like get_functional_objectives, add_functional_test, or import_functional_tests. Usage is implied rather than explicitly routed.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

generate_threat_modelGenerate Threat ModelA

Generate a complete threat model from a feature description.

Analyzes the feature using the Security Properties (Confidentiality, Integrity, Availability, Usage) methodology with capability-defined attackers. Produces trust boundaries, asset inventory, attacker inventory, control objective matrix, and assumptions.

Runs a multi-step AI pipeline. Progress is reported automatically.

Similar-model short-circuit: if the backend finds an existing model in the workspace whose feature description substantially overlaps with the new one, it does NOT generate a duplicate. This tool returns {"similar_models": [{"id", "title", "reason"}, ...], "suggestion": "..."} with the candidate IDs instead. The agent should then either:

  • Call refine_threat_model on one of the candidates to extend the existing model (usually the right answer — avoids duplicate modeling of the same system and preserves control/assertion history).

  • Retry this tool with force=True to bypass the check and create a genuinely new model anyway (e.g., when the similarity is superficial and the operator confirmed the new model is distinct).

The request names its purpose, so the platform always generates: it never reads the description as a question or a change to another model.

ParametersJSON Schema
NameRequiredDescriptionDefault
forceNoSkip the similar-model detection and always create a new model. Default False — the check fires unless the operator / agent has explicit reason to bypass it.
parent_idNoOptional ID of an existing model to wire the new model under as a child on the recursive composition tree. The child then inherits the parent's topology and participates in composition (delta / inherited control credit). Default None — the model is created flat.
provenance_refNoBranch or tag name at that commit (optional).
server_versionYes
provenance_kindNoWhere the description came from, one of ``code``, ``ticket``, ``document``, ``manual``, ``mixed``. Empty (default) records nothing. For an existing repository pass ``provenance_kind="code"`` with ``provenance_repo_url`` and ``provenance_commit_sha`` (the HEAD you gathered from): the code is then authoritative and the model follows it. Any other kind means the description is intent and the code is measured against it. The same record can be set later with ``update_threat_model``.
feature_descriptionYesDescription of the feature or system to threat model. Can be a few sentences or a detailed spec.
provenance_repo_urlNoRepository URL the description was gathered from (``code`` kind).
provenance_commit_shaNoCommit SHA the description was gathered at (``code`` kind).
provenance_source_refNoIdentifier of the ticket or document the description came from (``ticket`` / ``document`` kinds).
provenance_source_urlNoURL of that ticket or document.

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the behavioral burden and performs well: it discloses a multi-step AI pipeline, automatic progress reporting, the exact JSON shape returned on similar-model short-circuit, and the rule that the request is always treated as generation rather than a question or a change to another model. It does not state permission/auth requirements, rate limits, or whether created models are reversible, which keeps it from a 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The purpose is front-loaded in the first sentence, and the remaining sections are logically ordered: methodology, outputs, pipeline behavior, then similarity handling. It is longer than typical, but most sentences carry decision-relevant information; the closing paragraph about request intent is useful but slightly advisory.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the complex multi-step generation, 10 parameters, 90% schema coverage, and an output schema, the description covers the non-obvious behaviors an agent needs: pipeline execution, progress reporting, similarity short-circuit return shape, and next-step routing. It stops short of stating permissions or workspace side effects, but those are secondary for this tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 90%, so the schema already documents the 10 parameters in detail; per the rubric that sets a baseline of 3. The description adds a usage cue for force=True but does not add meaning for parent_id, provenance fields, or the other parameters beyond what the schema provides.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (generate) and resource (threat model) scoped to a feature description, and names the methodology used. Sibling differentiation is explicit: the short-circuit paragraph points to refine_threat_model as the alternative when an overlapping model exists, so the agent can distinguish this tool from its nearest neighbor without opening schemas.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives a clear when-not condition (backend finds a substantially overlapping model) and prescribes the alternatives: call refine_threat_model on a candidate, or retry with force=True to bypass. It also explains why refine is usually right (avoids duplicate modeling and preserves control/assertion history), which is exactly the kind of routing guidance the dimension rewards.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_assertion_typesGet Assertion TypesA

List the assertion types submit_assertions accepts, with their params.

Read-only. Returns the catalogue as structured data: every type, what it proves, its soundness class, which params it requires, which it accepts (an array-valued param carries its item_schema), and a worked example. soundness_classes defines the five classes by the fact a pass establishes, weakest to strongest — presence, under_approximating_scan, existential_witness, sound_over_approximation, by_construction — and sound_classes names the two that can credit a for-all clause. covers gives the accepted form of a binding declaration.

Call this before writing assertions. submit_assertions names the types and their required params in its own description, but descriptions are prose a client may present only in part, and a half-list reads exactly like a whole one. This returns data, so what you get back is the complete contract.

ParametersJSON Schema
NameRequiredDescriptionDefault
typesNoOptional comma-separated type names to return (e.g. "file_exists,pattern_absent"). Omit for all of them.
server_versionYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.4/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden. It states "Read-only" and describes the return payload in detail: every type, what it proves, its soundness class, required/optional params, item_schema for array params, a worked example, and the semantics of soundness_classes, sound_classes, and covers. This goes well beyond a generic list call and gives the agent a clear expectation of behavior and output structure, even without an output schema reference.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is longer than typical, but every sentence contributes meaningful detail: purpose, output semantics, and usage rationale. It is front-loaded with the core purpose and then expands. While slightly verbose, it avoids redundancy and earns its length through dense, relevant content.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool has an output schema, so return-value details could be omitted, but the description still explains the output's semantic structure, which is valuable. The main gap is the unexplained server_version parameter, but the tool itself is simple and well-contextualized. Overall, the description is sufficient for correct invocation, with only that minor parameter gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema has only 50% coverage, with server_version lacking any description. The tool description does not mention either of the tool's own parameters (types, server_version), so it adds no semantic value beyond the schema. Since the description is expected to compensate for schema gaps, the failure to explain server_version is a significant omission, leaving the agent without guidance on a required parameter.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a precise statement of purpose: "List the assertion types submit_assertions accepts, with their params." It clearly specifies the resource (assertion types) and the action (list), and immediately connects to the sibling tool submit_assertions, distinguishing itself as the authoritative catalogue. This is unambiguous and fully differentiates the tool from its siblings.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives explicit guidance: "Call this before writing assertions." It explains why this tool is preferred over relying on submit_assertions' prose description, noting that prose may be partially presented and a half-list reads like a whole one. This provides a clear when-to-use directive and justifies why the structured data is superior, covering the alternative explicitly.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_capabilitiesGet CapabilitiesA

A model's capabilities (behaviours the feature must deliver), or one of them. Read-only.

Without capability_id: every capability, each with its id, name/description and a summary of its component and asset bindings. With it: that capability with its bound components and assets.

ParametersJSON Schema
NameRequiredDescriptionDefault
model_idYesID of the threat model.
capability_idNoOptional; omit to list every capability.
server_versionYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden, and it does disclose a key trait: 'Read-only', which is the safety profile an annotation would normally supply. It also describes what each mode returns (ids, names, binding summaries). It omits auth/permission and error behavior, so it is not fully exhaustive.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short, front-loaded paragraphs with the list/single split clearly separated. Every sentence adds information and nothing is padded, though the parenthetical definition is slightly redundant with the title.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be exhaustively explained, and the two required params plus the optional one are adequately covered. The description is complete enough to call the tool correctly, with only minor gaps around permissions and alternatives.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 67%: model_id and capability_id are documented in the schema, while server_version is not. The description clarifies capability_id's effect on the result shape, but that largely restates the schema's own 'omit to list every capability' note, and server_version remains unexplained.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource (fetch a model's capabilities) and defines the term inline ('behaviours the feature must deliver'). It clearly separates the two modes (list all vs. single with bindings). It does not, however, distinguish itself from similarly-scoped siblings like get_control_objectives or get_functional_objectives.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives mode-selection guidance tied to capability_id (omit to list all, provide for one with bindings), which is useful implied usage. It offers no guidance on when to prefer this tool over sibling capability/objective tools, and no prerequisites or exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_compliance_reportGet Compliance ReportA

Compliance gap-analysis report for one framework at a chosen scope. Read-only; no side effects. Requires PRO tier.

Evaluates every framework requirement against the mapped controls in scope and classifies each as covered, partial, uncovered, unmapped, or excluded, then returns coverage counts plus per-requirement rows. The framework must first be activated at the same scope via select_compliance_frameworks (with the matching scope), otherwise there is nothing to report on.

scope selects the boundary and how scope_id is read:

  • "model" — a single threat model (scope_id = model id).

  • "tag" — rolled up across every member model of a tag (scope_id = tag id).

Filtering / pagination:

  • level — level filter for level-aware frameworks; returns only requirements at or below this level (e.g. 1 for L1 only). Omit (or 0) for all levels. Honored for all scopes.

  • status — one of "covered", "partial", "uncovered", "unmapped", "excluded"; empty = all statuses. Model scope only.

  • offset / limit — per-requirement row pagination; offset skips the first N rows, limit caps rows returned (0 = no explicit limit). Model scope only.

A tag report is neither paginated nor status-filtered; passing status, offset, or limit with scope="tag" raises an error rather than silently returning unfiltered rows.

ParametersJSON Schema
NameRequiredDescriptionDefault
levelNooptional level filter; omit for all levels.
limitNomax requirement rows to return, 0 = no explicit limit (model scope only).
scopeYesreport boundary — "model" or "tag".
offsetNoskip the first N requirement rows, pagination (model scope only). Default 0.
statusNooptional per-requirement status filter (model scope only).
scope_idYesid of the model or tag selected by ``scope``.
framework_idYesframework to report on (already selected at this scope; see ``list_compliance_frameworks``).
server_versionYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.9/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations present, the description carries the full burden and does so: read-only, no side effects, PRO tier requirement, the semantics of how scope_id is interpreted per scope, and a documented error path for invalid scope/filter combinations. This is well beyond what structured fields provide.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core purpose and safety profile, then cleanly sectioned into scope and filtering/pagination blocks. Slightly long, but nearly every line carries operational detail that affects a correct call.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists so return values need no explanation, yet the description still clarifies the shape (coverage counts plus per-requirement rows). All required params, scope behavior, and error conditions are covered for a complex 8-param, 4-required tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is already high (88%), but the description adds real meaning: how scope_id is read under each scope value, that level returns requirements at or below the given level, the exact status enum values, and that offset/limit/status are model-scope only.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource+scope: a read-only compliance gap-analysis report for one framework at a chosen scope, and enumerates the classification outcomes (covered/partial/uncovered/unmapped/excluded). This clearly differentiates it from siblings like check_control_gaps and get_verification_report.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly states the prerequisite that the framework must first be activated at the same scope via select_compliance_frameworks, names the follow-up tool (list_compliance_frameworks) for framework ids, and specifies the exact error behavior when status/offset/limit are passed with scope="tag".

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_compositionGet CompositionA

A model's composed view: its own entities with everything inherited from its ancestors on the recursive tree. Read-only.

view selects what is returned:

  • overview (default, ~1-2KB, read it first): {model_id, model_version, flag_enabled, tree: {parent_id, ancestor_chain, depth, child_ids}, counts: {entities, control_objectives, reconciliation_candidates}, warnings}.

  • entities: the effective entity set keyed by kind (trust boundaries, components, assets, attackers, attack paths), each entry {kind, qualified_id, owner_model_id, owner_title, origin, entity}. Paginated (page, page_size); kind (e.g. "attackers") keeps one kind.

  • objectives: effective COs {co_qid, asset_qid, attacker_qid, security_properties, origin}, where origin is own, cross (inherited, with a local asset or attacker) or inherited.

  • coverage: per effective CO {co_qid, is_covered, own_credit, inherited_credit, contributing_controls} — the composed figures, not the per-model get_verification_report. Paginated; origin filters by contributing-control origin.

  • attack_paths: {effective_paths, lattice_positions, authored_paths, suggestions: {missing_path, dangling_path}} against the composed topology.

Where composition is not available, every view returns its shape empty with flag_enabled: false rather than an error. Composed reachability is get_reachability_verdicts(composed=True).

ParametersJSON Schema
NameRequiredDescriptionDefault
kindNo
pageNo
viewNooverview
originNo
model_idYes
page_sizeNo
server_versionYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does well: it declares 'Read-only', specifies that unavailable composition returns an empty shape with flag_enabled=false rather than an error, notes pagination on the entities and coverage views, and gives an overview size hint (~1-2KB). It omits permission/auth prerequisites and error behavior for invalid model_id, so not exhaustive.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Purpose and read-only status are front-loaded, then the bulk is a per-view bullet list where every entry earns its place by telling the agent which view to pick. Given five distinct return modes, the length is proportionate and no sentence is filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 7-parameter tool with an output schema (so return values need not be described), the definition covers behavior, view selection, edge cases, and pagination well enough to call correctly. The only omissions are the undocumented server_version and model_id semantics.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate, and it largely does: it fully documents the view enum with the return shape of each mode, explains kind as a single-kind filter, origin as a contributing-control-origin filter, and notes page/page_size as pagination. model_id is only implied and server_version is never mentioned, leaving a small gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific resource and scope: the model's composed view combining its own entities with everything inherited from ancestors on the recursive tree. This distinguishes it from per-model read tools like get_verification_report and get_control_objectives, which it explicitly contrasts.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives clear sequencing guidance ('overview ... read it first') and routes the agent away from misuse by naming alternatives: coverage is the composed counterpart to per-model get_verification_report, and composed reachability belongs to get_reachability_verdicts(composed=True). It stops short of an explicit 'use this when / do not use this when' statement, but the routing is strong.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_control_assumption_groupsGet Control Assumption GroupsA

Get the current assumption group structure for a control.

Assumption groups define alternative sets of external claims that can satisfy a control:

  • Within a group: AND — all assumptions must be active and attested

  • Across groups: OR — any complete group is sufficient to mark the control as externally handled

ParametersJSON Schema
NameRequiredDescriptionDefault
model_idYesID of the threat model.
control_idYesID of the control (e.g., "CTRL-03").
server_versionYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the behavioral burden. It explicitly indicates a read-only operation ('Get') and adds meaningful domain behavior: assumptions within a group are ANDed, groups are ORed, and a complete group marks the control as externally handled. It does not discuss auth or error cases, but the getter nature and semantics are transparent enough for an agent.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is compact and front-loaded with the purpose, followed by a concise bulleted explanation of the AND/OR semantics. Every sentence adds value and no filler exists.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a read-only control query with an output schema, the description gives the essential domain model: assumption groups, external claims, AND within groups, OR across groups, and the effect on externally-handled status. It could additionally state that this is the read counterpart to set_control_assumption_groups, but the sibling names already make that connection.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema already describes model_id and control_id, but server_version is undocumented, and the description adds no parameter-level detail. With 67% schema coverage, the description does not need to restate everything, but it also does not compensate for the missing server_version context.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific verb ('Get') and a resource ('assumption group structure') scoped to a control, then defines exactly what that structure means with an AND/OR explanation. This clearly distinguishes it from write-oriented siblings like set_control_assumption_groups.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The word 'Get' and 'current' provide clear retrieval context, and the description implies when an agent would want to inspect the assumption group structure. It does not explicitly name alternatives or state when not to use it, so it falls short of full routing guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_control_generation_statusGet Control Generation StatusA

Read a model's control build: the one proposed, and the last one started. Read-only.

A build runs only when someone starts it; a write that owes controls PROPOSES one. proposal (or null) carries mode, objective_count, estimated_credits and the model_version / set_revision that start_control_build must name. Poll until terminal; hint names the next action.

status: queued | generating | deferred | pausing | paused | blocked | complete | failed | skipped | discarded | none. deferred waits for the daily budget reset. pausing is stopping; paused keeps its staged work until resume_control_generation (or discard_control_build). blocked carries code (dependency_unavailable or analysis_incomplete), message and retry_after_seconds: relay the message and retry with resume_control_generation, never regenerate_controls, which redoes and re-bills the work.

While running: ready_cos / target_cos count progress, never coverage; stage names the stage; elapsed_seconds is the time since the last progress (large means it may be stuck). selfheal_activity is a SAMPLE of what a strengthening round works on: read refining_total / authoring_total / set_aside_total for the counts; set_aside objectives wait for a person and are NOT a failure.

Once complete: covered_cos of judged_cos objectives would be mitigated by their controls; awaiting_judgement_cos have no answer yet. diagnosis counts covered, uncovered and undecided (what strengthen_controls works on), judging (queued: wait), not_judged (none queued: judge_objectives) and awaiting_assumption (in the review queue). strengthening says whether that pass has run. analysis_pending means the figures may still move; duration_seconds is the runtime.

ParametersJSON Schema
NameRequiredDescriptionDefault
model_idYes
server_versionYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.9/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does so well: it enumerates the status state machine, explains deferred vs pausing vs paused, states that paused retains staged work, surfaces blocked codes with retry_after_seconds, and warns that set_aside is not a failure and that large elapsed_seconds may mean a stuck build. It omits auth/rate-limit or polling-frequency guidance, which keeps it short of a 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Purpose is correctly front-loaded, but the bulk of the text is a field-by-field walkthrough of the response, which an output schema already covers. The state-semantics explanations add interpretive value, so it is not wasted, but it is longer than needed for a two-parameter polling tool.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a stateful build-tracking tool the behavioral coverage is thorough and an output schema exists, so return values do not need narrating here. The remaining gap is the complete silence on the two required input parameters.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% and both required params (model_id, server_version) are undocumented in the schema, so the description must compensate. It never explains what model_id or server_version are or their formats; the only version-adjacent text refers to returned fields (model_version, set_revision), not inputs.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The opening sentence names a specific verb and resource ('Read a model's control build') and immediately states scope ('the one proposed, and the last one started') plus the read-only nature. An agent can distinguish this from start_control_build, resume_control_generation and regenerate_controls without opening another schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives a concrete usage loop ('Poll until terminal; hint names the next action') and explicit alternative routing ('retry with resume_control_generation, never regenerate_controls', and judge_objectives for awaiting_judgement). It lacks a clean when-to-call/when-not framing at the top, but the routing content is strong.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_control_objectivesGet Control ObjectivesA

Get the control objective matrix, or one control objective. Read-only.

Two modes, selected by whether co_id is set:

  • Matrix mode (co_id omitted) — returns the model's COs, each with references to the controls that cover it. By default returns a compact summary (total count only); pass offset/limit to page through full CO records.

  • Single mode (co_id set) — returns that one CO's typed fields, the IDs of any controls that map to it, and the deterministic reachability verdict (the structural derivation that backs any reach claim on the CO). Tombstoned COs (removed: true) are returned with the flag set; the verdict is omitted because reach state is frozen at the removal version. offset/limit are ignored in this mode.

For pass/fail assurance scoring use assess_model.

ParametersJSON Schema
NameRequiredDescriptionDefault
co_idNoIf set, single mode — return this one control objective (e.g. ``CO3``) with its verdict. If omitted, matrix mode.
limitNoMatrix mode — max to return (0 = summary only, no per-CO records).
offsetNoMatrix mode — skip the first N control objectives.
model_idYesID of the threat model.
server_versionYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.9/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full behavioral burden, and it does so thoroughly: read-only semantics, mode-dependent return contents, tombstone handling, verdict omission, and ignored parameters are all disclosed. This goes well beyond a minimal 'gets data' phrasing.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is well-organized with a short opening summary followed by clear bullet-style mode breakdowns. Every sentence carries useful information, and the critical read-only and mode-selection cues are front-loaded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description fully covers the tool's behavioral complexity: mode selection, pagination, summaries vs full records, tombstone behavior, verdict semantics, and a pointer to the correct sibling. An output schema is present, so the description need not restate return values, and an agent has enough information to invoke this tool correctly in either mode.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 80%, so the schema already documents most parameters. The description adds meaningful context beyond the schema by explaining that co_id selects the mode, that limit=0 means summary-only, and that offset/limit are ignored in single mode.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly identifies the resource (control objective matrix or a single control objective) and the action (Get), and it explicitly frames the operation as read-only. It also distinguishes itself from the sibling assess_model by routing pass/fail assurance scoring to that tool.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It specifies exactly when to use each mode based on whether co_id is set, describes defaults for matrix mode, and states that offset/limit are ignored in single mode. It also explicitly names assess_model as the alternative for pass/fail assurance scoring.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_controlsGet ControlsA

A model's controls, or one control. Read-only.

Without control_id: {controls, total, returned}, the published set, filtered by status, co_id and component_id and paged by offset / limit. Orphaned controls (every mapped CO tombstoned) are left out unless include_orphaned=True, soft-deleted ones unless include_deleted=True. summary_only=True returns only id, description, status, verification_status, assertion_count, co_ids, assumption_groups and attestation_dependency. An empty list on a new model means its proposed build was never started (get_control_generation_status); while a build holds the model the list is the last published set, with a building marker.

With control_id: that control, with an orphaned flag; version reads it as of a model version (0 = latest).

An objective id on a control is an ATTACHMENT, not credit: only a member of a required mitigation group earns any, and a control that is defense-in-depth everywhere can be implemented and verified without moving an objective. Read get_mitigation_groups before evidence work.

status is the operator's (not_implemented / implemented / verified); verification_status is the EVIDENCE: verified (both tiers pass and the assertions cover the whole description), partially_verified (a tier failed, clauses are unproven, or an attestation EXPIRED — get_sufficiency says which; the fix is more assertions or a narrower description, never a recompute), pending, unverified (none submitted). A for-all clause is credited only by a sound type bound with covers; get_control_work_order lists what each clause needs. status="implemented" finds what still needs evidence. Only this model's own controls are listed; inherited ones are counted by assess_model.

ParametersJSON Schema
NameRequiredDescriptionDefault
co_idNo
limitNo
offsetNo
statusNo
versionNo
model_idYes
control_idNo
component_idNo
summary_onlyNo
server_versionYes
include_deletedNo
include_orphanedNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full disclosure burden and largely meets it: read-only nature, default exclusion of orphaned/soft-deleted rows via include_orphaned/include_deleted, the summary_only projection, the 'building' marker while a build holds the model, and version=0 meaning latest. It omits any permission/auth or rate-limit context, which keeps it short of a 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The purpose is front-loaded and the 'Without control_id / With control_id' split gives useful structure. The later paragraphs on objective attachment, mitigation groups, and clause crediting are dense and somewhat tangential to invoking this reader, so it runs longer than strictly needed.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 12-parameter domain-specific tool the coverage is strong: return shapes are sketched, filter semantics are defined, and follow-up tools are named for the edge cases it flags. An output schema exists, so the return-value detail is redundant but not harmful.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate, and it does: control_id as the mode switch, status with its enum values, co_id/component_id/offset/limit as filters and paging, include_orphaned/include_deleted defaults, summary_only, and version semantics (0 = latest). Only the self-evident required ids (model_id, server_version) go unmentioned.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The opening sentence states a specific resource and scope ('A model's controls, or one control') and immediately flags the operation as read-only. It cleanly separates the two modes (list without control_id, single entity with control_id), so an agent can distinguish it from neighbors like get_control_work_order or list_assertions.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives real operational routing: read get_mitigation_groups before evidence work, consult get_control_generation_status when the list is empty, and use status="implemented" to find what still needs evidence. It does not, however, explicitly say when to prefer this over sibling readers such as get_control_work_order or get_sufficiency.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_control_work_orderGet Control Work OrderB

The work order for one control: the ticket to read BEFORE implementing it. Read-only.

Returns {model_id, model_version, control, objectives, max_tier, scan_brief, assertion_contract, acceptance_criteria, required_evidence, steps, reconcile_rules, delegation, open_proposals, provenance}: where to look, what counts as proof (assertion_contract: the types by soundness class, the evidence and universal_rule, the sound_types the platform takes, what to submit with, when the control counts as verified), what to do when the code disagrees with the model, what this agent may decide alone (delegation), and whether code or description is authoritative (provenance).

Where the order names a required class for a clause, required_evidence holds one entry per such clause: its clause, the clause_id to put in covers, its quantifier, the required_class, what is missing, and a suggested_submission skeleton. The skeleton is a fill-in, not a submission: replace its <...> placeholders; one left unfilled is refused, by this client and again by the platform, since it would record a claim nothing backs. A for-all clause takes [by_construction, sound_over_approximation]: prefer typed_boundary (the type the sinks accept, its constructors), else sink_default_deny (the sinks, the safe forms, a reviewed allowlist). When the bound evidence is the wrong CLASS, missing says so. acceptance_criteria is GENERATED from those entries, so a clause that must hold at every site reads as such and no number of tests closes it.

A type's soundness is the class to branch on; behavioral is a compatibility field for readers written before the classes.

ParametersJSON Schema
NameRequiredDescriptionDefault
model_idYes
control_idYes
server_versionYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

B3.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden, and it does substantial work: it declares read-only behavior, describes the returned contract fields, and explains validation rules such as unfilled placeholders being refused by both client and platform. It omits auth requirements and error modes, but for a read-only tool this is a strong disclosure.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The purpose is front-loaded, but the description is a dense wall of domain-specific prose. Most sentences are relevant to the work-order contract, yet the lack of formatting and heavy nesting make it harder to scan than necessary.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so the description need not explain return values, though it still does at length. The critical gap is input guidance: with no annotations and no parameter descriptions, an agent receives no help understanding the three required inputs, leaving the definition incomplete for correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema has 0% description coverage for three required parameters: server_version, model_id, and control_id. The description does not explain what any of these inputs mean or what format they expect, so it adds no semantic value beyond the bare schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific resource and scope: 'The work order for one control,' and clarifies it is the ticket to read before implementing it. It distinguishes itself from mutation-oriented siblings by explicitly saying 'Read-only,' but it does not name any alternative tool for comparison.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives clear timing guidance: read this 'BEFORE implementing' the control. That is a useful usage context, but it provides no explicit exclusions or alternative tools for when this is not the right call.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_cwe_catalogGet Cwe CatalogA

Get the platform's CWE reference catalog status.

Returns {enabled, current_version, entry_count, versions}. When CWE classification is not turned on for this instance, enabled is false and the rest is empty — this is a normal informational response, not an error.

ParametersJSON Schema
NameRequiredDescriptionDefault
server_versionYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.6/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations supplied, the description carries the burden of behavioral disclosure. It does so by explaining the returned fields and, importantly, calling out the disabled-CWE edge case as a normal informational response rather than an error. It does not discuss authentication or side effects, but for a read-only status tool the main behavioral surprise is covered.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is compact and front-loaded: the first sentence states action and resource, the second gives the return shape, and the edge-case clarification earns its place. There is no filler or repetition.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

This is a simple status tool and the description covers return semantics and the disabled-case behavior, aided by the presence of an output schema. However, the complete absence of server_version guidance is a real gap: an agent cannot reliably know what to pass for a required field.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema has one required parameter, server_version, with 0% schema description coverage, and the description never mentions it or explains what values or formats are expected. The description adds no meaning beyond the bare property name.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific action ('Get') and resource ('platform's CWE reference catalog status'), and further clarifies what is returned (`{enabled, current_version, entry_count, versions}`), which distinguishes it from sibling CWE-related tools like get_model_cwe_tags or classify_model_cwe.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies the tool is used when an agent needs CWE catalog status, such as checking whether CWE classification is enabled, but it gives no explicit when-to-use guidance, exclusions, or alternatives. The sibling list contains related tools, but the description does not route between them.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_design_leverageGet Design LeverageA

Rank what eliminating each attacker position or asset BY DESIGN would remove from the matrix. Call this when deciding whether to change the design instead of implementing controls: it shows which single design change retires the most critical and high at-risk objectives. Read-only; no side effects.

Each row in ranked is an attacker or asset with the objectives its removal would take out of the matrix (objectives_removed, broken down by tier in removes), how many of those are currently at risk (removes_at_risk / removes_at_risk_by_tier), and the controls that would be retired. Rows are ranked by critical, then high, at-risk objectives removed. design_move (a concrete change of design that would eliminate the row) is filled only when include_design_moves is true. To act on a row, raise a design_change proposal with create_proposal; never apply a design change yourself.

ParametersJSON Schema
NameRequiredDescriptionDefault
topNoNumber of rows to return. Default 5.
model_idYesID of the threat model.
server_versionYes
include_design_movesNoAlso author a ``design_move`` per row. Default False.

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations to inherit safety traits, the description carries the full burden, and it does well: it declares 'Read-only; no side effects,' explains the ranking order, discloses that design_move is populated only when include_design_moves is true, and describes the per-row fields. It does not cover edge cases such as empty results, tie-breaking, or failure behavior, which keeps it one step short of excellent.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is rich but tightly organized: purpose, invocation context, output semantics, and the follow-up action rule each occupy clear, purposeful sentences. There is no filler, tautology, or unnecessary repetition, and the important 'never apply a design change yourself' rule is included without bloat.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a read-only analytical tool with an output schema, the description supplies the call context, ranking semantics, row field meanings, and the required next step via create_proposal. It does not define the matrix or critical/high at-risk terminology, and server_version remains undocumented, but overall it is well above the minimum needed for an agent to select and invoke the tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 75%, so model_id, top, and include_design_moves are already documented in the schema. The description adds useful context by clarifying that design_move is filled only when include_design_moves is true, but it does not explain server_version or provide additional parameter-level detail beyond the schema, placing it at the baseline 3.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The opening sentence states a specific action: ranking what eliminating each attacker position or asset BY DESIGN would remove from the matrix. It clearly distinguishes this from remediation-focused tools by framing the purpose as changing design instead of implementing controls, and it anchors the follow-up action to create_proposal.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly tells the agent when to call it: 'Call this when deciding whether to change the design instead of implementing controls.' It also gives clear post-call guidance by instructing the agent to raise a design_change proposal and never apply a design change itself. However, it does not name sibling alternatives or state explicit when-not-to-use conditions, so it stops short of full exclusion guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_entityGet EntityA

Get a single entity of any core type by ID. Read-only.

Dispatches on entity_type to the per-type read and returns that type's native record as-is (not wrapped in an array):

  • asset — the asset's typed fields. Soft-deleted assets carry deleted: true; the caller decides whether to surface them. entity_id e.g. A-01.

  • attacker — the attacker with its factor decomposition, its surface_extent (unset / point / whole) and surface_extent_source, which says whether a person attested it. Soft-deleted attackers carry deleted: true. entity_id e.g. T-03.

  • component — the component. Speculative components (repo_url="") are returned as-is: the empty repo IS the lifecycle state, not an error. entity_id e.g. CMP-01.

  • trust_boundary — the boundary incl. its passes set (closed-vocabulary subset of {Network, Adjacent, Local, Physical}). entity_id e.g. TB-Net.

  • assumption — the assumption with its override applied (mirrors list_assumptions' merge for one entity: typed fields, the structured exclusion predicate when present, and the override layer — status / justification / linked CO IDs). Soft-deleted assumptions carry deleted: true. entity_id e.g. AS-01.

ParametersJSON Schema
NameRequiredDescriptionDefault
model_idYesID of the threat model.
entity_idYesID of the entity to fetch.
entity_typeYesWhich entity to read — one of ``asset``, ``attacker``, ``component``, ``trust_boundary``, ``assumption``.
server_versionYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations present, the description carries the full behavioral burden and does so well: it discloses that records are returned natively (not wrapped in an array), that soft-deleted entities carry deleted: true rather than erroring, that speculative components with empty repo_url are a valid lifecycle state, and that assumptions return with their override merged.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The purpose and read-only status are front-loaded in the first two lines, then structured as a scannable per-type bullet list. Each bullet carries type-specific semantics rather than filler, though the overall body is long and slightly repetitive in its per-type 'entity_id e.g.' phrasing.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

All five dispatchable entity types are covered with their distinctive return semantics, and an output schema exists to document return shape. For a typed-getter tool, an agent has everything needed to select the right entity_type and interpret the response.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 75% and the schema already enumerates entity_type values, so the description is not strictly required here. It still adds value by giving concrete entity_id format examples per type (A-01, T-03, CMP-01, TB-Net, AS-01) and by tying each entity_type to the specific shape returned.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Starts with a specific verb+resource ('Get a single entity of any core type by ID') and immediately states read-only. It also names the dispatch mechanism (entity_type) and enumerates the five core types, clearly separating it from siblings like edit_asset, add_asset, and remove_entity.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Clearly implies single-entity lookup by ID and states the tool is read-only, which frames the usage context. However, it never explicitly contrasts itself with enumeration siblings (list_threat_models, list_assertions) or states when not to use it, so no explicit alternative routing is provided.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_findings_risksGet Findings RisksA

Workspace-scoped triage dashboard: open findings, active risk acceptances, and at-risk Control Objectives across every model the workspace can access.

Use this as the entry point when an operator asks "what's open?" or "what should I work on next?" — one round-trip returns all three categories with model context and risk dimensions (severity, status, risk_tier, owner, review_by) so the agent can triage without per-model fan-out. The endpoint is read-only and fast; it composes from existing per-model queries server-side.

Returns the envelope verbatim: {workspace_id, evaluated_at, models, findings, risk_acceptances, at_risk_cos, summary}. summary carries totals (open_findings, total_findings, active_risk_acceptances, total_risk_acceptances, at_risk_cos) for quick health-check responses.

ParametersJSON Schema
NameRequiredDescriptionDefault
server_versionYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.1/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden, and it delivers: it explicitly states the endpoint is read-only and fast, composes server-side from existing per-model queries, and returns the envelope verbatim. This gives an agent a clear mental model of the operation's safety and behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is well-structured: purpose first, usage trigger second, return envelope third. Each sentence adds substantive information, and there is no filler or redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description covers scope, contents, usage context, and return shape in enough detail for an agent to understand what the tool provides. The only notable missing context is the meaning of 'server_version' and how it relates to workspace-triage behavior.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The sole required parameter 'server_version' has no schema description and no coverage in the tool description. The description never mentions the parameter, its purpose, format, or allowed values, leaving the agent to guess what to pass. With 0% schema description coverage, the description needed to compensate and did not.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly identifies a specific resource ('workspace-scoped triage dashboard') and enumerates exactly what it returns: open findings, active risk acceptances, and at-risk Control Objectives across every model. This distinguishes it from siblings like list_findings and list_risk_acceptances by emphasizing the aggregate, multi-model scope.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives an explicit trigger ('what's open?' or 'what should I work on next?') and positions the tool as the entry point, with the benefit of no per-model fan-out. It does not explicitly name an alternative tool or state when not to use it, so it falls just short of a full 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_functional_coverageGet Functional CoverageA

A model's functional coverage report, or just its actionable gaps. Read-only.

By default the whole picture: per-objective state (verified / covered / failing / untested), the Capabilities × Conditions matrix, and the applicable / missing-objective / not-applicable cell accounting. With gaps_only=True only what needs action: applicable conditions with no objective yet, and objectives that are failing or have no passing test.

ParametersJSON Schema
NameRequiredDescriptionDefault
model_idYesID of the threat model.
gaps_onlyNoReturn only the actionable gaps (default False).
server_versionYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations present, the description carries the full burden and does well: it explicitly declares 'Read-only' and enumerates what each mode returns. It doesn't discuss permissions, cost, or latency, but for a read-only report the safety profile is adequately disclosed.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with purpose and the read-only flag, then the default behavior, then the gaps_only behavior. The matrix and cell-accounting enumeration is dense but each clause maps to real return content, so little is wasted.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists so return values needn't be spelled out, yet the description helpfully sketches them anyway. Parameter meaning is covered for the one non-obvious flag, making the definition complete enough to call correctly despite the two undocumented-by-description identifiers.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 67%, and the description meaningfully expands gaps_only beyond the schema's terse 'Return only the actionable gaps' by defining a gap as applicable conditions with no objective, failing objectives, or objectives lacking a passing test. This adds genuine semantic value, though server_version and model_id rely on the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific resource ('functional coverage report' for a model) and distinguishes its two modes (full picture vs. actionable gaps). It is clear, but it doesn't differentiate itself from near-siblings like get_functional_objectives, get_sufficiency, or get_verification_report, which also surface coverage-adjacent views.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is implied through the two-mode explanation: default for the whole picture, gaps_only=True when only actionable items matter. However, it never states when to choose this tool over the many sibling reporting/gap tools (check_control_gaps, get_sufficiency), leaving cross-tool selection to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_functional_objectivesGet Functional ObjectivesA

List a model's functional objectives, or fetch one by id. Read-only; no side effects.

A functional objective is a Capability × Condition test plan expressed as a Given-When-Then statement. functional_objective_id selects the behaviour:

  • omitted / empty string -> list every functional objective for the model (the full functional test plan).

  • a functional-objective id -> return just that one objective's detail, including its capability, condition, Given-When-Then statement, and current test state.

For pass/fail coverage state across all objectives use get_functional_coverage; for the actionable gaps pass it gaps_only=True.

ParametersJSON Schema
NameRequiredDescriptionDefault
model_idYesID of the threat model whose functional objective(s) to read.
server_versionYes
functional_objective_idNoOptional. Omit (or pass "") to list every objective; pass an id (from a prior list call) to fetch that one.

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the burden and does state 'Read-only; no side effects,' which is the key safety signal. It also discloses what the single-object fetch returns (capability, condition, Given-When-Then, test state), though it says nothing about auth requirements or pagination on the list path.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core action and the read-only note, then a bulleted enumeration of the id-selection behavior. The definition sentence for 'functional objective' earns its place, but the prose is slightly longer than strictly needed.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a read tool with an output schema, the description needn't explain return values, and it covers the list-vs-fetch duality plus the sibling alternative. The only omissions are auth/pagination behavior and any treatment of server_version.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 67%, and the description adds real meaning beyond the schema by spelling out the empty-string/omitted sentinel behavior and where the id comes from ('from a prior list call'). model_id is only implicitly covered, and server_version is never discussed.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('List a model's functional objectives, or fetch one by id') and explicitly defines what a functional objective is, which lets an agent distinguish it from generic objective tools. It also names the sibling get_functional_coverage and the exact condition that routes there.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives explicit when-to-use branching: omitted/empty id -> list the full test plan, an id -> fetch one objective's detail. It also names the alternative tool (get_functional_coverage) and the parameter (gaps_only=True) that selects it.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_functional_satisfaction_groupsGet Functional Satisfaction GroupsA

Read the satisfaction-group structure for a functional objective. Read-only; no side effects.

A satisfaction group is a set of functional tests that together satisfy the objective: AND within a group (every test in the group must be verified), OR across groups (any one complete group satisfies the objective). Returns the current numbered groups plus any tests associated with the objective but not placed in a group.

Use before set_functional_satisfaction_groups to see the current structure, or to trace why an objective is / isn't satisfied. This is the functional analog of get_control_assumption_groups / get_mitigation_groups.

ParametersJSON Schema
NameRequiredDescriptionDefault
model_idYesID of the threat model.
server_versionYes
functional_objective_idYesThe objective whose groups to read.

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries full burden. It discloses read-only/no side effects, explains the AND/OR group semantics, and states what the return includes (numbered groups plus unplaced tests). This is valuable behavioral context beyond a bare 'get'.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise and well-structured: front-loaded purpose, then semantics, then usage guidance. Every sentence adds unique value with no repetition or fluff.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

There is an output schema, so return details are covered elsewhere. The description sufficiently covers the tool's role, semantics, read-only nature, and relationship to siblings. Minor gap: no mention of error conditions or prerequisites, but not critical for this simple get operation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 67% (server_version lacks a description). The description reinforces the meaning of functional_objective_id by explaining what a satisfaction group is, but doesn't add new details about server_version or model_id beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states a specific verb ('Read') and resource ('satisfaction-group structure for a functional objective'). It distinguishes itself from siblings by explicitly positioning itself as the functional analog of get_control_assumption_groups / get_mitigation_groups and by implying the read counterpart to set_functional_satisfaction_groups.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It provides explicit usage context: 'Use before set_functional_satisfaction_groups to see the current structure, or to trace why an objective is / isn't satisfied.' It also names the sibling alternatives, making the selection decision clear.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_groupGet GroupA

Get one tag by ID, with its member model ids. Read-only; no side effects.

Returns {id, workspace_id, name, description, created_at, model_ids}. Discover tag IDs with list_groups.

ParametersJSON Schema
NameRequiredDescriptionDefault
group_idYesID of the tag.
server_versionYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations supplied, the description carries the full burden and does state the safety profile ('Read-only; no side effects'), which is exactly the behavioral fact an agent needs before calling. It stops short of permissions or error behavior (e.g., what happens for an unknown ID), so it is solid but not exhaustive.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short sentences, each earning its place: purpose, safety profile, return shape, and ID discovery. The scoping verb phrase is front-loaded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so restating the return fields is redundant but harmless, and the description adds the safety profile and the ID-discovery path an agent needs. The main remaining gap is that server_version is unexplained and the group/tag naming is never reconciled.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 50%: group_id is already documented in the schema, and server_version has no description anywhere. The description's 'by ID' adds no syntax or format detail beyond the schema, so baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Get one tag by ID, with its member model ids'), and the phrase 'one ... by ID' distinguishes it from the sibling list_groups, which it names. The only wrinkle is the group/tag terminology mismatch between the tool name (get_group) and the description ('tag'), which adds minor ambiguity about whether these are the same entity.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly tells the agent how to obtain a valid ID ('Discover tag IDs with list_groups'), which routes it to the right sibling for the prerequisite step. There is no explicit when-not guidance, but for a simple fetch-by-ID tool the context is clear.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_group_dependenciesGet Group DependenciesA

The reliance edges among a tag's member models. Read-only; no side effects.

Returns {tag_id, tag_name, models, edges, total}: one row per edge a member declared on another model's control (manage_reliance / attach_foundation), with both models' titles, the mode (delegated or relied_upon), the source objective or control, the provider control, its status (draft, active, broken, rejected), validation_verdict and credit_state (whether it credits its objective now). Empty when no member relies on another.

Use it to review the cross-model dependencies of a product or audit scope before an auditor export, or to find broken edges. A model's own edges, in both directions, are list_reliance.

ParametersJSON Schema
NameRequiredDescriptionDefault
group_idYesID of the tag.
server_versionYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden and does so: it declares read-only/no side effects, describes the empty-result case ('Empty when no member relies on another'), and explains edge direction via the mode values (delegated vs relied_upon) plus which calls create edges.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the purpose and the read-only guarantee, then the return shape, then usage. The return-field enumeration is dense but each field (mode, status, validation_verdict, credit_state) carries semantic weight; slightly longer than strictly necessary given an output schema exists.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a read-only audit/query tool with an output schema present, the description covers purpose, invocation context, edge semantics, and enumeration values. Nothing an agent needs in order to call it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 50%: group_id is documented as 'ID of the tag' but server_version is undocumented. The description adds context by using tag/member-model terminology consistently, tying the unspecified group_id to the tag concept, but it never explains server_version or parameter format. Baseline 3 is appropriate given partial coverage.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific resource and scope: 'the reliance edges among a tag's member models,' and explicitly contrasts with the sibling list_reliance, which covers a single model's own edges in both directions. An agent can distinguish the two without opening either schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives concrete when-to-use cases ('review the cross-model dependencies of a product or audit scope before an auditor export, or find broken edges') and names the alternative (list_reliance) with the condition that selects it. No inference required.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_mitigation_groupsGet Mitigation GroupsA

Get the current mitigation group structure for a control objective.

Returns the grouped view of controls for this CO with details (id, description, status) for each control:

  • groups: numbered groups (within=AND, across=OR)

  • defense_in_depth: tracked but not required for mitigation

  • unmapped: model controls not mapped to this CO (available for assignment)

Use cases:

  • Before set_mitigation_groups to see the current structure

  • When reviewing a CO's assessment to understand why it is at_risk or mitigated

  • When deciding which unmapped controls to assign to a CO

ParametersJSON Schema
NameRequiredDescriptionDefault
co_idYesID of the control objective (e.g., "CO5").
model_idYesID of the threat model.
server_versionYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden and does a solid job: it discloses the output categories (groups, defense_in_depth, unmapped), the semantics of grouping (within=AND, across=OR), and that unmapped controls are available for assignment. It does not explicitly state that the operation is read-only, but 'Get the current structure' and the use case 'Before set_mitigation_groups' make this clear.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is well structured and front-loaded: purpose first, then return details as a compact list, then concrete use cases. Every segment earns its place and the overall length is appropriate for the tool's complexity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description covers the key semantics an agent needs: what the return groups mean, why defense_in_depth is tracked but not required, and how unmapped controls are used. With an output schema present, the lack of detailed return-field explanations is acceptable, though the description could more fully explain the within/across AND/OR logic.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 67%, with co_id and model_id already documented in the input schema. The description adds no additional meaning for the parameters themselves and does not explain server_version, which is undocumented even though it is required. This is adequate but not improved beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool retrieves the current mitigation group structure for a control objective and specifies the returned grouped view with details. It is not a tautology and is distinguishable from generic 'get' tools, though it does not explicitly contrast itself with similar group-fetching siblings like get_control_assumption_groups or get_functional_satisfaction_groups.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

A dedicated 'Use cases' section explicitly tells the agent when to call this tool: before set_mitigation_groups, when reviewing a CO's assessment, and when deciding which unmapped controls to assign. It gives clear context but does not state when not to use it or name alternative group-fetching tools.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_model_cwe_tagsGet Model Cwe TagsA

List CWE weakness classifications tagged onto a model's control objectives.

Each tag's name/description are resolved from the platform's CWE catalog, never model-authored. A tag whose CWE id has since been deprecated, redefined, or removed by MITRE carries a stale reason (missing / deprecated / changed) — re-run classify_model_cwe to refresh it. 404s if CWE classification is not enabled on this instance.

ParametersJSON Schema
NameRequiredDescriptionDefault
model_idYesID of the threat model to inspect.
server_versionYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries full responsibility for behavior disclosure. It does so thoroughly: tags are resolved from the platform CWE catalog and never model-authored, stale tags are characterized with specific reasons (missing/deprecated/changed), and a 404 is documented when the feature is disabled. This is strong behavioral context for a read-only list tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is front-loaded with the core purpose, and every subsequent sentence adds meaningful behavioral or error context. It is compact, avoids repetition of the tool name, and contains no filler or redundant restatement of the input schema.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return-value details do not need to be repeated. The description covers purpose, data provenance, stale semantics, refresh action, and the relevant 404 error condition. The main gap is the undocumented required server_version parameter, which prevents full completeness for an agent trying to invoke the tool correctly without further context.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 50%, with server_version left completely undocumented in both the schema and the description. The description indirectly clarifies model_id by discussing 'a model's control objectives,' but it adds nothing to explain the required server_version parameter. With low schema coverage, the description should compensate for the gap and does not.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource: 'List CWE weakness classifications tagged onto a model's control objectives.' This clearly states what the tool does and separates it from the related classify_model_cwe sibling by referencing that tool for refreshing stale tags. However, it does not explicitly contrast itself with get_cwe_catalog, so it stops short of full sibling differentiation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly names classify_model_cwe as the action to re-run when a tag is stale, which gives the agent a condition-based routing choice. It also notes the 404 failure mode when CWE classification is not enabled. It does not provide broader when-to-use versus when-not-to-use guidance relative to other catalog-related siblings.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_reachability_verdictsGet Reachability VerdictsA

Per-CO reachability verdicts, over this model alone or the composed tree. Read-only; derived each time, never stored.

composed=False (default): derived from this model's own structure (components, asset.component_ids, trust_boundary.passes, each attacker's trust_boundary_ids and vector, assumption exclusion predicates), deterministic, the derivation an auditor re-runs. co_id returns one verdict (404 if absent or tombstoned). Returns {model_id, model_version, verdicts: [{co_id, kind, reason, narration, boundary_id?, assumption_id?}]}; kind is reachable, unreachable or indeterminate.

composed=True: the same derivation over the model with everything it inherits from its ancestors, for a child on the composition tree. Paginated (page, page_size); kind_filter keeps one kind; co_id is ignored. Returns {model_id, flag_enabled, verdicts: [{co_qid, asset_qid, attacker_qid, kind, reason}], total, page, page_size}, empty with flag_enabled: false where composition is not available.

An indeterminate verdict names the missing structure: attacker_unpositioned (edit_attacker with trust_boundary_ids), asset_unbounded (assign_to_components(target_type="asset")), no_shared_boundary (reposition the attacker, rescope the asset, or an add_assumption exclusion), missing_entity (restore it, or remove the CO). model_coherence_report presents the same gaps as findings.

ParametersJSON Schema
NameRequiredDescriptionDefault
pageNo
co_idNo
composedNo
model_idYes
page_sizeNo
kind_filterNo
server_versionYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full behavioral burden and does so richly: it discloses read-only status, that verdicts are derived each time and never stored, determinism, 404 behavior for absent or tombstoned COs, pagination behavior, `flag_enabled: false` when composition is unavailable, and the meaning of each `kind`. It also enumerates the indeterminate-verdict reasons and points to the corrective actions for each.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is front-loaded with the core purpose and then organized by mode, which is appropriate for a complex two-mode tool. It is long but mostly information-dense, with each section explaining distinct behavior. Some return-shape detail may overlap with the output schema, but the structure remains readable and useful.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 7-parameter tool with no annotations, 0% schema coverage, and compositional behavior, the description is unusually complete. It covers both modes, parameter interactions, return behavior, error conditions, and the follow-up actions for indeterminate verdicts. An agent has enough context to invoke the tool correctly in either mode.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate, and it explains five of the seven parameters in meaningful behavioral terms: `composed`, `co_id`, `page`, `page_size`, and `kind_filter`. It does not explain `server_version` or `model_id`, though `model_id` is likely self-evident. The default behavior for `composed` and the interaction between `co_id` and composed pagination are especially well covered.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: per-CO reachability verdicts over a model alone or the composed tree. It immediately distinguishes the two scopes (`composed=False` vs `composed=True`) and notes the read-only derived nature of the result. An agent can tell what this tool returns without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives explicit mode-specific context: `composed=False` derives from the model's own structure, while `composed=True` is for a child on the composition tree. It also explains that `co_id` returns one verdict in the default mode and is ignored in composed mode, and that `kind_filter` keeps one verdict kind. It stops short of naming a preferred alternative beyond noting that `model_coherence_report` presents the same gaps as findings.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_remediation_leverageGet Remediation LeverageA

Remediation-leverage plan for a model: which controls to implement first to close the most control objectives with the least work.

Returns the model's not-yet-satisfied controls ranked by how many control objectives each one closes (ranked), plus a greedy minimal fix order — the sequence of controls that reaches the most mitigated objectives with the fewest controls (greedy_plan) — and a summary of the collapse (total objectives, currently mitigated, how many controls the plan needs). Use to prioritize implementation work: a single call tells the agent which controls give the highest leverage, so it can tackle the shortest path to coverage instead of fixing objectives one at a time. Read-only.

Composed models: each entry in ranked and greedy_plan also carries its owning model — owner_model_id and owner_model_title — and an inherited flag. inherited is true when the control is authored on an ancestor model, meaning the fix lands on that model rather than the one being assessed; summary.inherited_candidate_controls counts them. Surface the owning model so the operator knows which high-leverage fixes belong to a parent model. A flat (non-composed) model reports every control as owned by the assessed model.

ParametersJSON Schema
NameRequiredDescriptionDefault
model_idYesID of the threat model.
server_versionYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the burden of behavioral disclosure. It states the operation is read-only, explains the returned structures, and highlights composed-model behavior: owner_model_id, owner_model_title, inherited flag, and summary.inherited_candidate_controls. This goes well beyond the schema and helps the agent understand side effects and ownership semantics.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is information-dense but every section earns its place: core output, use case, and composed-model nuance. The front-loaded first paragraph gives the essential purpose and returns, while later details address edge cases. It could be tightened slightly, but the structure is logical and not padded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool complexity and existing output schema, the description covers the main outputs, the use case, read-only behavior, and composed-model edge cases. The only notable gap is the missing semantics for server_version, which would matter for an agent trying to invoke this correctly in varied environments.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema only describes model_id as 'ID of the threat model' and leaves server_version undescribed. The tool description does not clarify server_version at all and adds no parameter-level detail beyond what the schema already provides. With 50% schema description coverage, this leaves one required parameter semantically opaque.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states that this tool returns a remediation-leverage plan ranking controls by leverage and a greedy minimal fix order. It is specific about the resource (a model) and the output shape (ranked, greedy_plan, summary), making it distinguish itself functionally from most siblings, though it does not explicitly name or contrast an alternative tool.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly says 'Use to prioritize implementation work' and explains that a single call gives the shortest path to coverage, contrasting with fixing objectives one at a time. It gives clear context for when to call this tool, but does not discuss when not to use it or name a specific alternative tool.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_review_queueGet Review QueueA

Returns the workspace's review queue: what needs a decision or a re-check, ranked. Read-only; no side effects.

Each row carries an item_type, one of escalation (a judgment an agent was refused and parked for a person), proposal (an open change of scope or design, or an assumption proposal: a precondition a strengthening run found only the environment can meet), unaccepted_assumption (an assumption something depends on that is not accepted — never attested, lapsed, or its text changed since it was attested — with the controls and objectives that wait on it), open_assumption, or stale_control (an implemented/verified control whose assertions have not been checked in 90+ days). Rows are ranked in that order. Escalations and proposals are decided with decide_proposal; an unaccepted assumption is accepted with submit_attestation; for each stale control, verify its assertions against the codebase. Accepting an assumption is a person's judgment unless the workspace delegates it. Start here for periodic maintenance.

ParametersJSON Schema
NameRequiredDescriptionDefault
server_versionYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does so well: it explicitly declares 'Read-only; no side effects,' explains the ranking order, and details what each item type represents (including the 90+ day threshold for stale controls). It omits auth/permission requirements, but coverage is otherwise rich.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core purpose and side-effect status, then elaborates the taxonomy in a structured way. The length is justified by the number of item types that must be distinguished, though the parenthetical definitions are dense and could be trimmed slightly.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be explained, yet the description adds essential semantic context by defining each item_type and how to act on it. Nothing needed to call and correctly interpret this tool is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0% and the single parameter (server_version) is never mentioned in the description. The parameter is standard plumbing rather than a semantic input, so the gap is minor, but the description does not compensate at all for the undocumented param.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('Returns the workspace's review queue') and immediately scopes it as 'what needs a decision or a re-check, ranked.' The item_type taxonomy (escalation, proposal, unaccepted_assumption, open_assumption, stale_control) further distinguishes it from sibling listing tools like list_proposals.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives clear routing guidance: 'Start here for periodic maintenance,' and maps each row type to its resolution tool (decide_proposal for escalations/proposals, submit_attestation for unaccepted assumptions). It doesn't explicitly contrast with alternatives like list_proposals or get_controls as when-to-use options, but the entry-point framing is strong.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_risk_viewGet Risk ViewA

Prioritized Risk View — one row per live Control Objective — at a chosen scope. Read-only; no side effects.

scope selects the aggregation boundary and how scope_id is interpreted:

  • "model" — a single threat model (scope_id = model id). One row per live CO with derived risk tier, asset impact, attacker likelihood, control coverage counts (coverage_ratio), and open-finding count (open_findings). Tombstoned COs are excluded; pair with get_threat_model if historical context is needed. Use to triage which COs need attention on one model — a single call ranks the work, no per-CO fan-out.

  • "tag" — every member model of a tag (scope_id = tag id). The same row shape with model_id and model_title added per row, so rows can be grouped by source model without an extra lookup, and delegation-aware (delegation_mitigated / delegating_controls): a CO mitigated via a verified cross-model delegation reads as covered, consistent with each model's own assessment. Use for a product, portfolio or audit-scope posture rollup.

ParametersJSON Schema
NameRequiredDescriptionDefault
scopeYesaggregation boundary — "model" or "tag".
scope_idYesid of the model or tag selected by ``scope``.
server_versionYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.9/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does so well: it declares 'Read-only; no side effects,' states that tombstoned COs are excluded, discloses the row shape (coverage_ratio, open_findings, risk tier, asset impact, attacker likelihood), and explains the delegation-aware semantics that make a cross-model-mitigated CO read as covered. These are non-obvious behavioral traits an agent could not infer from the schema.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded one-line summary followed by a clean two-bullet breakdown. Dense but almost every clause carries information; the embedded field lists in the 'model' bullet are slightly heavy, keeping it just short of maximally tight.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 3-parameter read tool with a low-complexity schema and an output schema that covers return structure, the description supplies everything an agent needs: purpose, scope semantics, exclusions, and delegation handling. Nothing material is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 67%, but the description compensates richly by defining how scope_id is interpreted under each scope value ('model id' vs 'tag id') and what the resulting rows contain per scope. This adds meaning well beyond the terse schema descriptions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Names a specific resource and return grain: 'Prioritized Risk View — one row per live Control Objective — at a chosen scope.' This distinguishes it from siblings like get_controls, get_findings_risks, and get_compliance_report without needing to open any schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly enumerates both scope modes and when each applies: 'model' for triaging one model's COs ('a single call ranks the work, no per-CO fan-out'), 'tag' for product/portfolio/audit-scope rollups. It also names the paired tool (get_threat_model) for historical context, giving clear when-not guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_scan_promptGet Scan PromptA

Get guidance prompts for scanning a codebase. Read-only; no side effects.

kind selects which scan brief to return:

  • "security" (default) — prompts telling the agent what evidence to look for per security control; only NOT_IMPLEMENTED controls are included (implemented ones need no scan). Use this to drive a gap-discovery pass, then record what is missing with submit_findings and what is present with submit_assertions. Pass control_id to scope the prompt to one control; empty (default) returns prompts for all not-yet-implemented controls.

  • "functional" — the agent brief for implementing functional-conformance tests. Generation specifies the functional tests, so for each test not yet verified this returns its implementation brief and the objectives it proves; it also reports objectives_without_tests (regenerate or add a test) and missing_objectives (applicable conditions with no objective yet). Drive test implementation from it, then call submit_assertions (functional_test_id) with TEST_EXISTS + TEST_ATTESTED assertions so CI verifies each test; read the resulting pass/fail state via get_functional_coverage. control_id does not apply to this kind and is ignored.

ParametersJSON Schema
NameRequiredDescriptionDefault
kindNo"security" (default) or "functional" — which scan brief.security
model_idYesID of the threat model.
control_idNoSecurity kind only — optional single control to scope the prompt to. Empty (default) returns prompts for all not-yet-implemented controls. Ignored when kind="functional".
server_versionYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.6/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does so reasonably: it declares "Read-only; no side effects" and discloses filtering behavior (only NOT_IMPLEMENTED controls included) and that control_id is ignored for functional. It omits auth/permission requirements and pagination/size characteristics, but for a read-only prompt getter these are minor.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with purpose followed by clearly delineated bullets per kind. It is on the longer side, but each clause carries mode-specific workflow information rather than restating the name, so the length is largely earned.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 4-parameter tool with an output schema, the description supplies the missing context an agent needs: what each kind returns, how control_id interacts with each kind, and which downstream tools consume the result. Nothing needed to invoke it correctly is absent.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 75% and the schema already documents kind and control_id, but the description adds meaningful semantics: default resolution of empty control_id ("returns prompts for all not-yet-implemented controls"), the mode-scoped applicability of control_id, and the effect of each kind on returned content. This goes beyond the schema text.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Starts with a specific verb+resource ("Get guidance prompts for scanning a codebase") and then precisely distinguishes the two modes it serves: security briefs vs functional briefs. An agent can tell exactly what the tool returns and how the two kinds differ without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicit when/when-to-use for each mode: security to "drive a gap-discovery pass", functional to "drive test implementation", each with the follow-up calls (submit_findings, submit_assertions, get_functional_coverage). It also states control_id does not apply to the functional kind, closing off a misuse path.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_setup_statusGet Setup StatusA

Get the workspace onboarding checklist with completed and pending steps. Read-only.

Call this before suggesting or performing setup actions so already-done steps aren't repeated; mark a step done with complete_setup_step. Takes no arguments beyond the version header.

ParametersJSON Schema
NameRequiredDescriptionDefault
server_versionYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden and clearly states the operation is read-only. It also clarifies that the only parameter is a version header, preventing an agent from expecting functional arguments. This is meaningful behavioral disclosure.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short sentences with zero waste. The purpose is front-loaded, the usage guidance follows immediately, and the parameter clarification is concise. Every sentence earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple read-only status tool with an output schema, the description gives enough context: what it returns, when to call it, and what arguments it needs. Nothing essential is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, but the description compensates by explaining that the tool 'takes no arguments beyond the version header,' which directly clarifies the sole required parameter, server_version. It doesn't specify format details, but it correctly prevents misuse.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('Get') and resource ('workspace onboarding checklist') and clearly distinguishes the tool from the large sibling set. It also names the complementary mutation tool, complete_setup_step, which helps an agent understand the tool's role.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives explicit usage guidance: call this before suggesting or performing setup actions to avoid repeating completed steps. It does not explicitly list exclusions or alternative query tools, but the context is clear and actionable.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_sufficiencyGet SufficiencyA

Whether the submitted assertions of one control, or of one functional test, together prove it. Read-only; name exactly one id.

For a control this explains verification_status: "partially_verified". status is sufficient | insufficient | pending, with freshness (fresh | stale | pending) beside it; insufficient carries details naming EACH uncovered clause and the evidence that would close it. A soundness_tier is the weakest clause's tier: a control is proven no more strongly than its thinnest clause. Reading does not queue a re-evaluation: the write that changed a control queues its own. For the whole model use get_verification_report.

Act by submitting the named assertions; a clause describing a mechanism the system does not use calls for refine_control, not evidence. get_control_work_order serves the per-clause list: where the order names a required class for a clause, required_evidence carries the clause id for covers and a suggested_submission skeleton whose <...> placeholders you replace. For a for-all clause prefer typed_boundary, else sink_default_deny. class_mismatch means the bound evidence is the wrong CLASS and more of it will not help. An attestation covers an existential clause, never a for-all one; that clause's only other exits are a risk acceptance or a not-applicable disposition.

For a functional test it is whether the test's evidence proves the objectives it is associated with, with the reasoning; computed after evidence is submitted, so it can read pending or absent until then.

ParametersJSON Schema
NameRequiredDescriptionDefault
model_idYes
control_idNo
server_versionYes
functional_test_idNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the burden and delivers key traits: read-only, no re-evaluation is queued by reading ('the write that changed a control queues its own'), and functional-test results are computed only after evidence is submitted, so they may read pending/absent. It does not cover auth/rate/permission context, but for a read-only query tool this is strong.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The purpose is front-loaded and clear, but the body is dense and mixes in remediation guidance (typed_boundary vs sink_default_deny, class_mismatch, attestation exits) that belongs to the action tools rather than to a read-only sufficiency lookup. Several sentences do not earn their place for selecting or invoking this tool.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a domain-complex tool with an output schema (so return-value detail is not required here), the description supplies the status/freshness/soundness_tier semantics and the pending-on-functional-test caveat an agent needs. The remaining gap is the unexplained model_id/server_version inputs and the ambiguity of whether 'one id' is enforced by the caller.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, so the description must compensate and it partially does: it explains that control_id and functional_test_id are mutually exclusive and that exactly one must be named, plus what each answers. However, model_id and server_version (both required) receive no semantic explanation anywhere, leaving half the parameters opaque.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific question the tool answers ('whether the submitted assertions of one control, or of one functional test, together prove it') plus its operation mode ('Read-only; name exactly one id'). It explicitly names the sibling that covers the other scope ('For the whole model use get_verification_report'), so an agent can distinguish it without opening a schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Names alternatives with the selecting condition: get_verification_report for the whole model, get_control_work_order for the per-clause list, submit_assertions/refine_control/show risk acceptance for the follow-up actions. It also constrains invocation ('name exactly one id') so the agent knows the required input shape.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_threat_modelGet Threat ModelA

Get a specific threat model by ID.

Returns the full threat model including trust boundaries, assets, attackers, control objectives, and assumptions.

Important for agents reading model state:

  • Assets and attackers may carry deleted: true (soft-deleted). Exclude these when showing "what's in the model now"; include them only when discussing history or offering restore. Restore an entity via restore_entity (entity_type="asset") / restore_entity (entity_type="attacker").

  • Control objectives may carry removed: true (tombstone — the (asset, attacker) pair was removed in a later version). Exclude these from coverage math and LLM prompts; they exist to keep CO IDs stable so controls referencing them can be detected as "orphaned" rather than silently rebinding.

ParametersJSON Schema
NameRequiredDescriptionDefault
versionNoOptional specific version number. Defaults to latest.
model_idYesID of the threat model.
include_cosNoInclude control objectives inline (default False: the answer carries no ``control_objectives`` key; read them with ``get_control_objectives``).
server_versionYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.9/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Without annotations the description carries the full behavioral burden, and it delivers valuable domain-specific semantics: soft-deleted entities (deleted:true) and tombstones (removed:true) and how to treat them. It stops short of stating permission requirements or output shape, but the model-state caveats are exactly the kind of non-obvious behavior an agent needs.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded purpose sentence followed by a payload list, then clearly headed agent notes. Slightly long but every line earns its place; minor redundancy between the header and body.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values needn't be explained, and the description focuses on the non-obvious entity-state semantics instead. The omission of usage routing to sibling tools is the main gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 75% and the schema already documents version, model_id, include_cos, and server_version. The description adds no parameter-level detail beyond what the schema provides, so baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource ('Get a specific threat model by ID') and enumerates the returned payload (trust boundaries, assets, attackers, control objectives, assumptions). This clearly distinguishes it from siblings like list_threat_models and query_threat_model.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is implied by 'get by ID' and the restore references, but the description never explicitly states when to use this versus query_threat_model or list_threat_models. No when-not guidance or alternatives are named.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_verdict_divergenceGet Verdict DivergenceA

Where the LLM's verdicts disagree with the model's authored state.

Two coverage divergence kinds, distinguished by the LLM's p_covers (probability the control covers the CO), shown as "model confidence":

  • missing_mapping: HIGH p_covers, but the CO is NOT mapped — the LLM is confident the control covers it, so it should be mapped. Accepting ADDS the mapping.

  • spurious_mapping: LOW p_covers, but the CO IS mapped — the LLM is confident the control does NOT cover it, so the mapping is likely wrong and inflates apparent coverage. Accepting REMOVES the mapping. Only confident rows surface; the uncertain middle band is dropped. So a ~100%-confidence row is a strong "add" and a ~0%-confidence row is a strong "remove" — both are actionable, in opposite directions.

Rows are sorted by confidence, so the strongest calls come first. Each section is paginated: its pagination.filtered_total reports the full count, so when it exceeds the rows returned, raise limit (up to 500) or page with offset to review every divergence — not only the first page.

Also returns group_sufficiency divergences (observation-only). Apply coverage rows with resolve_verdict_divergences(action="accept"); set aside rows the structural model got right with resolve_verdict_divergences(action="dismiss").

ParametersJSON Schema
NameRequiredDescriptionDefault
kindNoOptional filter — "missing_mapping", "spurious_mapping", or "group_sufficiency". Empty returns all kinds.
limitNoMax rows per section (clamped to 1-500, default 100). Set to 500 to pull an entire section in one call.
offsetNoSkip the first N rows of each section, for pagination.
model_idYesID of the threat model.
server_versionYes
include_dismissedNoWhen true, return ONLY previously-dismissed rows (the undo view) instead of the active list.

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden and does so well: only confident rows surface, the uncertain middle band is dropped, rows are sorted by confidence, and accepting a row either ADDS or REMOVES a mapping depending on kind. It also discloses pagination semantics (pagination.filtered_total) and the limit clamp. This is rich behavioral disclosure beyond any structured field.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The purpose is front-loaded and the bulleted breakdown of the two kinds is scannable. It runs long, and the sentence restating that ~100% confidence means 'add' and ~0% means 'remove' is partly redundant with the bullets above it, but overall the structure is efficient for the workflow described.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be re-explained, yet the description still supplies the operational context an agent needs: the two-kind model, direction of effect, confidence sorting, pagination, and the resolve/dismiss resolution path. Nothing required to call this correctly or act on its output is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is already 83%, so the baseline is 3, but the description adds genuine value beyond the schema: it explains the meaning of the two kind values and group_sufficiency, and clarifies limit (raise up to 500 to pull a whole section) and offset in terms of per-section pagination. The include_dismissed undo-view semantics are left to the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The opening line precisely defines the resource: rows where the LLM's verdicts disagree with the model's authored state. It names the two divergence kinds (missing_mapping, spurious_mapping) and explicitly distinguishes this read-only retrieval tool from the sibling that acts on it, resolve_verdict_divergences. An agent can tell what this returns and how it differs from get_sufficiency or model_coherence_report without opening a schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives clear routing guidance: apply coverage rows via resolve_verdict_divergences(action="accept"), set aside model-correct rows with action="dismiss", and treat group_sufficiency as observation-only. It also instructs when to raise limit or page with offset. It stops short of an explicit 'use this when X instead of Y' statement versus sibling diagnostic tools, but the next-step guidance is strong.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_verification_reportGet Verification ReportA

Get verification report with summary stats and sufficiency gaps.

Returns tier1/tier2 pass/fail/pending counts, per-control verification status, and sufficiency details.

Each per-control sufficiency block carries:

  • status: "sufficient" | "insufficient" | "pending" | "stale". "stale" means the stored verdict no longer reflects the current control description, active assertion set or the rules it was computed under. Reading does not queue a re-evaluation: the write that changed a control queues its own. Call this tool again later for the refreshed verdict.

  • details: human-readable LLM reasoning.

  • misaligned_assertion_ids: assertions whose stated subject is off-topic for the control's current description (common after a control has been refined or regenerated). Treat as a directive: rebind to the right control, supersede via delete_assertion, or rewrite. Do NOT treat them as evidence. A non-empty list forces the verdict to "insufficient".

  • stale: boolean shortcut for status == "stale", kept distinct so an INSUFFICIENT verdict that's also stale (the prior insufficient decision was computed under outdated inputs) can be flagged without overloading status.

A drift item means the accepted evidence changed (a test's definition, a witness's scope or allowlist) and its verdict was withdrawn until reviewed again.

By default returns summary only (no per-assertion details). Set summary_only=False to include full assertion details and drift items.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNoMax control entries to return (0=all).
offsetNoSkip first N control entries.
statusNoFilter by verification status: "verified", "partially_verified", "pending", "unverified".
model_idYesID of the threat model.
summary_onlyNoOmit per-assertion details and drift items (default True).
server_versionYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.6/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations the description carries the full load, and it delivers meaningful semantics: what 'stale' means, that reading does not queue a re-evaluation, that misaligned_assertion_ids is a directive (not evidence) and forces 'insufficient', and what a drift item implies. It still omits permission/auth requirements and rate-limit behavior, so not a 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Purpose is front-loaded in the first sentence, but the bulk of the text is spent describing return structure (sufficiency blocks, stale flag, drift items) even though an output schema exists and already governs those fields. The status/drift semantics partly justify the length, but the block is longer than it needs to be.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a read tool with an output schema, the description supplies the non-schema semantics an agent needs (stale semantics, re-evaluation timing, misaligned-assertion handling, default truncation). Only the sibling-selection guidance and any access prerequisites are missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 83%, so the baseline is 3. The description restates summary_only's effect in slightly richer terms (assertion details plus drift items) but adds nothing about limit/offset/status semantics beyond the schema's own text.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Get verification report') and enumerates the payload: tier1/tier2 pass/fail/pending counts, per-control verification status, and sufficiency details. That is enough to distinguish it from read-only siblings, though it never names an alternative like get_sufficiency or get_reachability_verdicts to sharpen the boundary.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The default-vs-full behavior is spelled out ('By default returns summary only... Set summary_only=False to include full assertion details and drift items'), which is real invocation guidance. But there is no when-to-use-this-vs-get_sufficiency reasoning and no exclusions, so usage is only implicit.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

import_compliance_frameworkImport Compliance FrameworkA

Import a custom compliance framework. Requires PRO tier.

Use this when your customer's program (regulatory, contractual, or internal) is not covered by Mipiti's built-in frameworks. After import, the framework is selectable on threat models exactly like a built-in.

Fields: name (required), version, description, requirements (required, non-empty), level_definitions.

Each requirement takes id and description (required), level (integer, default 1), the optional grouping chapter_id / chapter_name / section_id / section_name / title, scope (component, the default, or system: covered if ANY model satisfies it) and level_specific_text (per-level text).

level_definitions and level_specific_text are keyed by the level as a string integer ("1", "2"): the key is the ordinal the level <= target_level filter compares, so a non-integer key is refused (400). Labels ("Baseline", "SL3") go in each value's name. A level value is {"name", "description", "source"}, where source is authoritative (paraphrased from the published standard) or mipiti_convention (tiers you defined).

Example::

{
  "name": "ACME Tiered",
  "level_definitions": {
    "1": {"name": "Baseline", "description": "Minimum.",
          "source": "mipiti_convention"}
  },
  "requirements": [
    {"id": "ACME-PWD", "description": "Passwords meet policy",
     "level": 1, "level_specific_text": {"1": "Min 8 characters."}}
  ]
}
ParametersJSON Schema
NameRequiredDescriptionDefault
framework_jsonYesThe framework body as a JSON string.
server_versionYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does well: it discloses the PRO-tier entitlement requirement, the post-import state ('selectable exactly like a built-in'), and explicit error behavior (non-integer level keys are refused with 400). It omits what happens on duplicate framework names or whether import is idempotent, which is a meaningful gap for an import mutation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The definition is long but front-loads purpose, then entitlement, then usage condition, then field semantics, closing with a concrete example. Nearly every sentence conveys non-obvious payload structure, so the length is largely justified, though the field-by-field density could be tightened.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values needn't be explained, and the description covers the important non-schema behavior (entitlement, post-import availability, error codes) plus a full example of the payload. The main residual gap is documentation of server_version and duplicate-name handling, but overall it is sufficient to invoke the tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is low (50%) and the schema only says framework_json is 'the framework body as a JSON string', so the description must compensate. It does so thoroughly, documenting required vs optional fields, nested requirement semantics (scope 'any model' vs 'all'), level key ordinality, and the level value shape. The second parameter, server_version, remains entirely undocumented by both schema and description.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('import a custom compliance framework') and scopes it against the alternative path ('not covered by Mipiti's built-in frameworks'). It also clarifies the post-import behavior ('selectable on threat models exactly like a built-in'), which lets an agent distinguish it from siblings like list_compliance_frameworks or select_compliance_frameworks.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

'Use this when your customer's program (regulatory, contractual, or internal) is not covered by Mipiti's built-in frameworks' gives a clear triggering condition for the tool. It stops short of explicitly naming list_compliance_frameworks as the prerequisite check or naming an exclusion, but the context is unambiguous.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

import_controlsImport ControlsA

Import existing security controls into a threat model.

Accepts structured JSON or free-text. Controls are auto-mapped to COs and deduplicated against existing ones. The parse/map/dedup runs as a background job (polled for progress), then — because this mutates the model — you are asked to confirm before the controls are saved.

The saved controls are added to the model's current controls as one change (undoable with undo_model_change); no model version is created, and it is refused while a control build holds the model. Nothing runs for them unprompted: the result's awaiting_judgement lists them, and the mitigation groups they join credit nothing and read awaiting judgement until judge_imported_controls is called (estimate first, then confirm_estimate=True).

ParametersJSON Schema
NameRequiredDescriptionDefault
auto_mapNoAuto-map controls to COs using LLM (default: True).
model_idYesID of the threat model.
free_textNoFree-text controls (narrative/CSV/bullets).
source_labelNoOrigin label (e.g., "ISO 27001").
controls_jsonNoJSON array of {description, co_ids?, framework_refs?}.
server_versionYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.4/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description fully carries the burden and delivers: background parse/map/dedup with polling, a confirmation gate before mutation, undo semantics via undo_model_change, no model version created, refusal while a control build holds the model, and the awaiting_judgement lifecycle up through judge_imported_controls. This is unusually rich behavioral disclosure.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Purpose and the JSON/free-text input modes are front-loaded, and the workflow is layered in logical order. It is dense and slightly long, but nearly every sentence conveys a distinct behavioral fact rather than filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a high-complexity mutating import with a background job, confirmation, and judgement gating, the description covers the full lifecycle, including the undo path and the refusal condition. Since an output schema exists, it needn't explain return values, and nothing material is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 83%, so the schema already documents the six parameters well. The description's mention of 'structured JSON or free-text' loosely maps to controls_json/free_text but adds no format or constraint detail beyond the schema, so baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Import existing security controls into a threat model') and immediately distinguishes its scope by noting it maps to COs and deduplicated against existing ones, setting it apart from siblings like auto_map_controls and import_compliance_framework.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives clear context: accepts structured JSON or free-text, runs as a background job, requires confirmation before saving, and routes the agent to undo_model_change and judge_imported_controls as follow-ups. However, it never explicitly states when NOT to use it versus sibling importers (e.g., import_compliance_framework, convert_assumption_to_controls).

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

import_functional_testsImport Functional TestsA

Register tests that already exist in your codebase against a model's functional objectives, so tests you already have count toward functional conformance — not only Mipiti-specified tests. Mutating (bulk).

Scan the repo's test suite and pass the tests here. Optionally associate each with the objective ids it covers (from get_functional_objectives); the platform verifies each association is applicable before accepting it and returns any it rejected under rejected_mappings. A test with no (or a rejected) association is still imported, unmapped, so it can be associated later (see suggest_functional_test_mappings / associate_functional_test). For a single hand-authored test, use add_functional_test instead.

ParametersJSON Schema
NameRequiredDescriptionDefault
model_idYesID of the threat model.
tests_jsonYesA JSON array of test objects. Each object supports ``test_name``, ``file_path``, ``framework``, ``description``, ``status`` (not_implemented | implemented | verified — an operator claim; an independent CI run is what verifies it), and ``functional_objective_ids`` (list of objective ids the test covers). At least ``test_name`` or ``description`` is required per test; the rest are optional.
server_versionYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the behavioral burden and largely succeeds: it explicitly flags 'Mutating (bulk),' explains that associations are verified by the platform, and describes that rejected or missing associations still result in the test being imported unmapped. It does not detail side effects like idempotency or whether existing tests are overwritten, but the disclosed workflow behavior is substantial.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Every sentence earns its place: purpose, mutation warning, workflow, fallback behavior, and the alternative tool. The description is dense but well-structured, front-loading the core purpose and bulk-mutation nature before explaining details.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description is thorough for a bulk-import tool with an output schema: it covers the use case, parameters' semantics for the main payload, mutation behavior, rejection handling, and routing to related tools. The unresolved server_version parameter and the lack of clarity around 'Mipiti-specified tests' prevent a perfect score, but overall it is highly usable.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The description adds rich meaning to tests_json, explaining the optional fields, the status enum meanings, and association behavior beyond the schema. model_id is adequately described in the schema. However, server_version is a required parameter with no schema description and no mention in the tool description, leaving a significant semantic gap for a required input.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose: 'Register tests that already exist in your codebase against a model's functional objectives' so existing tests 'count toward functional conformance.' It names the specific resource (tests/functional objectives) and distinguishes itself from add_functional_test by noting bulk import of existing tests vs. a single hand-authored test.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives explicit guidance on when to use this tool: when importing existing tests from the codebase in bulk. It provides a clear exclusion: 'For a single hand-authored test, use add_functional_test instead.' It also explains follow-up workflows via suggest_functional_test_mappings / associate_functional_test for unmapped tests.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

import_threat_model_archiveImport Threat Model ArchiveA

Import a JSON audit archive (from export_report (scope="model", format="archive")) into a target workspace.

Mutating: creates a NEW threat model in the target workspace. Requires write access to that workspace. A fresh model_id is assigned on every import, so the same envelope can be imported any number of times without collisions; title collisions in the target workspace auto-suffix (imported YYYY-MM-DD). Non-destructive — never overwrites or touches an existing model.

Use to move or clone a model between workspaces or across instances; the envelope round-trips through export_report (scope="model", format="archive") first.

The archive carries the model's current state, and the import creates it as version 1 of a new model: its controls, live assertions, decisions in force and open findings. Earlier versions, activity and chat are not carried.

The restored model arrives UNVERIFIED. The tier verdicts on its assertions, the attested flag on a verification result, and the facts a verification run reported are the origin's record of what it claimed — kept with the model as that record, and not credited here: a verdict belongs to the run that produced it and the judge that decided it, and this workspace has neither. Verification is earned here by running it against code this workspace can reach, so plan for a restored model to read unverified until it has. The same holds for the judgements of its mitigation groups: the import queues none, and its objectives read awaiting judgement until someone asks for them.

ParametersJSON Schema
NameRequiredDescriptionDefault
envelopeYesThe full archive dict returned by ``export_report (scope="model", format="archive")``.
workspace_idYesTarget workspace to import into.
server_versionYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does so thoroughly: it declares the tool mutating, states the write-access requirement, explains that a fresh model_id is assigned so re-imports never collide, describes title auto-suffixing, and asserts non-destructiveness (never overwrites an existing model). It goes further by enumerating what the archive carries versus what it drops, and by warning that the restored model arrives unverified with origin verdicts kept only as a record.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The critical facts (mutating, creates new model, requires write access, non-destructive) are front-loaded before the longer provenance and verification caveats. The prose is dense and multi-paragraph, with some restatement of the envelope/export_report relationship, but given the tool's complexity nearly every sentence conveys distinct behavior.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need no explanation, and the description still covers everything an agent needs: prerequisites, identity/collision semantics, scope of what is restored, and the important caveat that verdicts and mitigation-group judgements do not carry over. Nothing material is missing for a 3-parameter tool with a nested envelope.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 67% (envelope and workspace_id documented; server_version bare), so the description must add some value and does: it clarifies that the envelope is the archive produced by export_report and spells out what that archive contains. It adds little for server_version, but overall it enriches the two meaningful parameters beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The first sentence states a specific verb (Import) and resource (a JSON audit archive) plus its destination (target workspace), and the second sentence specifies the exact effect: creating a NEW threat model. This is clearly distinguishable from siblings like import_controls, import_functional_tests, and import_compliance_framework, which handle different artifact types.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It explicitly names the use case ('move or clone a model between workspaces or across instances') and the required counterpart tool ('round-trips through export_report (scope="model", format="archive") first'), which tells the agent where the envelope comes from. It does not, however, state any when-not-to-use conditions or contrast with the other import_* siblings.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

judge_imported_controlsJudge Imported ControlsA

Have the imported controls awaiting their judgement judged. Mutating only with confirm_estimate=True; may consume credits then.

Controls saved by import_controls are not judged unprompted: the mitigation groups they join credit nothing until this runs. It judges every objective those controls join, priced and charged as judge_objectives prices and charges them.

  1. Call with confirm_estimate=False (the default). Nothing is queued and nothing is charged; the answer carries awaiting_judgement (the control ids), co_ids (the objectives they join), scope, ungrouped and estimate. Show the user the estimate.

  2. Call again with confirm_estimate=True once they agree. The judgements are queued (confirmed: true, queued) and the controls stop awaiting (awaiting_judgement comes back empty).

ungrouped lists objectives with no mitigation group: nothing can be judged there until their controls are grouped with set_mitigation_groups. When nothing awaits, the answer says so and does nothing.

Refusals come back as data, {confirmed: false, queued: 0, http_status, ...}: 409 while a control build is running, 402 when the balance this workspace bills to cannot cover the estimate, 503 when judging is unavailable on this deployment.

ParametersJSON Schema
NameRequiredDescriptionDefault
model_idYesID of the threat model.
server_versionYes
confirm_estimateNoFalse (default) returns the estimate and queues nothing; True queues the judgements.

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does so: it discloses that it is mutating only when confirm_estimate=True, that it may consume credits, that pricing matches judge_objectives, that it queues judgements, and it enumerates refusal statuses 409/402/503 returned as data rather than exceptions. This is unusually rich behavioral disclosure.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the mutating/credit caveat, then uses a numbered two-step list that is easy to follow. Dense but each paragraph (workflow, ungrouped, refusals) earns its place; only minor tightening is possible.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Even though an output schema exists, the description names the relevant return fields (awaiting_judgement, co_ids, scope, ungrouped, estimate, confirmed, queued) and explains refusal shapes, leaving no gap for correct invocation or interpretation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 67%; the description explains confirm_estimate's two modes more fully than the schema does, and ties model_id implicitly to the target threat model. It does not add meaning for server_version, but that param is boilerplate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Have the imported controls awaiting their judgement judged') and scopes it to controls saved by import_controls, distinguishing it from judge_objectives / judge_objective. An agent can identify the distinct role this plays in the control workflow.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives an explicit two-step protocol (estimate first with confirm_estimate=False, then confirm with True), names prerequisites (controls must be grouped, otherwise use set_mitigation_groups), and states the no-op case ('when nothing awaits, the answer says so and does nothing').

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

judge_objectiveJudge ObjectiveA

Have one control objective's mitigation group judged. Mutating — queues background work and may consume credits.

Use this for an objective whose risk_reason is awaiting_judgement: it has a built mitigation group and nothing has decided whether that group covers the objective — never evaluated, evaluated against inputs that have since changed, or the answer parked. The objective is not short of controls, so generating or implementing more will not move it; what is missing is the judgement.

This is not a repair. The judgement can come back insufficient, which moves the objective to coverage_gap / insufficient_by_design and names real work. That is the tool doing its job: it replaces "nobody has looked" with an answer, and the answer may be no.

Scoped to ONE objective, which is the difference that matters against recompute_verdicts: that tool force-enqueues every control's coverage verdict AND every live objective's group-sufficiency verdict, which on a large model runs to thousands of credits. This queues a single judgement. Any credits it consumes are metered at actuals as the work runs, like every other metered call — the account's usage is visible in its billing panel before and after.

Judging runs in the BACKGROUND; the call returns as soon as the work is queued. Re-read get_mitigation_groups (or assess_model / get_risk_view) shortly after to see the objective's new state. Calling again while a judgement is already queued is harmless and does not queue a second one.

A refusal comes back as data rather than an error, so it can be relayed:

  • {queued: false, http_status: 409, ...} — controls are still being generated for this model. Poll get_control_generation_status until terminal, then call again.

  • {queued: false, http_status: 503, ...} — judging is unavailable on this deployment.

  • {queued: false, http_status: 402, code, message} — the balance this workspace bills to cannot cover the judgement.

ParametersJSON Schema
NameRequiredDescriptionDefault
co_idYesID of the control objective to have judged (e.g. "CO5").
model_idYesID of the threat model.
server_versionYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden and does: mutating, queues background work, consumes credits metered at actuals, returns before work completes, re-calling is harmless/idempotent, and refusal semantics (409/503/402) returned as data rather than errors. This is unusually complete behavioral disclosure.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the core action and mutation warning, then structures the when-to-use, the not-a-repair caveat, and refusal codes. Longer than most, but each block (background behavior, refusal codes) earns its place; minor redundancy between the credit note and the metering note.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With an output schema present the description needn't detail return values, and it still covers background execution, how to observe the new state (get_mitigation_groups / assess_model / get_risk_view), and refusal payloads. Nothing an agent needs to invoke it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 67% and the description adds little per-parameter meaning beyond the implicit 'one objective' scoping of co_id. It does not clarify server_version (undocumented in schema) or parameter format. Adequate but not compensating for the coverage gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource (judge one control objective's mitigation group) and immediately names the sibling it differs from (recompute_verdicts), plus clarifies what it is NOT ('This is not a repair'). An agent can distinguish it from every sibling without opening a schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicit trigger: use for an objective whose risk_reason is awaiting_judgement, with the three sub-conditions spelled out. Also gives an explicit exclusion (the objective is not short of controls, so generating/implementing more won't move it) and names the alternative recompute_verdicts with its cost difference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

judge_objectivesJudge ObjectivesA

Have judged every objective whose mitigation group has no judgement for its current controls and none queued (the diagnosis's not_judged, or objectives reading awaiting_judgement). Mutating only with confirm_estimate=True; may consume credits then. Adding or implementing controls does not move such an objective; a judgement does. judging objectives are already queued: wait for them.

  1. Call with confirm_estimate=False (the default): nothing is queued or charged; the answer carries diagnosis, scope, ungrouped and estimate (credits, objectives, computed_at, rate_version). Show the user the estimate.

  2. Once they agree, call with confirm_estimate=True: each objective in scope is queued (confirmed: true, queued), metered at actuals as it runs, and status_detail is the fresh status.

ungrouped objectives have no mitigation group and are never judged: group their controls with set_mitigation_groups first.

A judgement is not a repair: it can come back insufficient or undecided, which counts the objective as uncovered or undecided, work for strengthen_controls. judge_objective does the same for one.

Refusals come back as data, {confirmed: false, queued: 0, http_status, code, message}: 409 control_generation_in_progress (poll get_control_generation_status), 402 insufficient_credits / quota_exceeded (with estimated_credits), 503 (judging unavailable). An unknown id in co_ids is a 400 error.

ParametersJSON Schema
NameRequiredDescriptionDefault
co_idsNoOptional comma-separated objective IDs; omit for every objective with no judgement and none queued.
model_idYesID of the threat model.
server_versionYes
confirm_estimateNoFalse (default) estimates; True queues.

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does so: it discloses that the tool is mutating only under confirm_estimate=True, that this can consume credits metered at actuals, and enumerates refusal payloads (409 control_generation_in_progress, 402 insufficient_credits/quota_exceeded, 503, 400 on unknown co_ids as data rather than exceptions).

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Well-structured with a numbered two-step list and grouped refusal codes, but the opening sentence is dense and awkwardly parsed, and the whole block is longer than strictly needed given the output schema exists. Every section is useful, so little is wasted.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a mutating, credit-consuming batch tool with an output schema, the description covers the decision procedure, the grouping prerequisite, the sibling alternative, outcome semantics (uncovered/undecided), and full refusal handling. Nothing an agent needs to invoke it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 75% and the schema already documents co_ids ('omit for every objective with no judgement') and confirm_estimate. The description reinforces these and adds the meaning of the estimate response subfields (credits, objectives, computed_at, rate_version), going somewhat beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (have judged), resource (objectives) and a precise scope (those with no judgement and none queued), and explicitly distinguishes itself from the single-objective sibling by noting 'judge_objective does the same for one.' An agent can pick this over judge_objective without opening either schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives an explicit two-step workflow (call with confirm_estimate=False to estimate, then True once the user agrees), plus when-not conditions: ungrouped objectives must be grouped via set_mitigation_groups first, judging objectives should be waited on, and insufficient/undecided outcomes route to strengthen_controls.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

lift_composition_entityLift Composition EntityA

Promote a shared-anchor entity from two sibling descendants to their lowest common ancestor. Mutates state across THREE models.

The operator has confirmed (via the composition lift-candidate view) that the entity local_id_a on descendant_a_id and the entity local_id_b on descendant_b_id are the same logical thing and should be modeled once on the LCA. The route's model_id is the operator's current context model — typically the LCA, but the server accepts any ancestor of both descendants.

Conflict resolution. The server re-detects field-level and attached-state conflicts against current live state before applying. If new conflicts have surfaced since the operator's last candidate fetch, the call returns 400 with the missing conflict keys; refresh the lift-candidate view and resubmit with resolutions covering every key. Each entry in field_resolutions / attached_state_resolutions is "keep_a" | "keep_b" | "keep_both" (union for list/set fields; falls back to B for scalars).

Over-application gate. The lift extends visibility to every descendant of the LCA, not just the two source descendants. The server runs an over-application gate that refuses lifts touching descendants outside an acknowledged set; pass acknowledged_third_party_subtrees to acknowledge specific subtrees, or skip_overapplication_gate=True to override entirely after explicit operator confirmation.

Each affected model (LCA + both descendants) bumps version and emits a model_refined activity event; a structured lift_applied event with the full lift_event payload lands on the LCA. The audit pack surfaces this under lift_history. Reverse it with undo_composition_event(event_type="lift"), which previews unless dry_run=False; the inverse operation is split_composition_entity.

ParametersJSON Schema
NameRequiredDescriptionDefault
kindYesEntity kind — one of ``"assets"``, ``"attackers"``, ``"components"``.
model_idYesOperator's context model — the model whose composition view surfaced the candidate. Treated as a route anchor only; doesn't have to be the LCA.
local_id_aYesLocal id of the entity on ``descendant_a_id``.
local_id_bYesLocal id of the entity on ``descendant_b_id``.
lca_model_idYesTarget ancestor model id (the LCA, or any ancestor higher up the chain).
server_versionYes
descendant_a_idYesFirst source descendant model id.
descendant_b_idYesSecond source descendant model id.
field_resolutionsNoOptional per-field resolution map (e.g. ``{"description": "keep_both", "tags": "keep_a"}``).
lca_descendant_idsNoOptional snapshot of the LCA's descendant set used by the over-application gate. Omit to let the server compute it via BFS.
skip_overapplication_gateNoWhen True, bypass the gate after explicit operator confirmation. Default False.
attached_state_resolutionsNoOptional per-state-key resolution map (e.g. ``{"state:assertions/AS3": "keep_b"}``).
acknowledged_third_party_subtreesNoOptional list of subtree roots the operator has acknowledged as in-scope for the lift.

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.9/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and delivers: it mutates state across THREE models, re-detects conflicts against live state and can return 400 with missing conflict keys, runs an over-application gate affecting descendants beyond the two sources, bumps versions, emits model_refined/lift_applied events, and is reversible via undo_composition_event (preview unless dry_run=False). These are precisely the behavioral traits an agent needs before invoking a destructive lift.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

It is long (roughly four dense paragraphs), but it is front-loaded with the purpose and each subsequent paragraph covers a distinct concern (conflict resolution, over-application gate, side effects/undo), so nearly every sentence earns its place. Minor tightening of the conflict-resolution prose would help.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 13-parameter destructive composition tool with no annotations, the description covers preconditions, failure modes, gate overrides, side effects, audit trail, and reversal. An output schema exists, so return values need not be explained, and nothing an agent needs to call this safely appears to be missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 92%, which sets a baseline of 3, but the description adds meaning the schema lacks: the allowed resolution values 'keep_a' | 'keep_b' | 'keep_both' (schema declares zero enums), the semantics of keep_both as a union for list/set fields with a fallback to B for scalars, the distinction between route model_id and the true LCA, and the purpose of acknowledged_third_party_subtrees versus skip_overapplication_gate. That is substantive value beyond structured fields.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The opening sentence states a specific verb+resource+scope: 'Promote a shared-anchor entity from two sibling descendants to their lowest common ancestor.' It also names the inverse sibling (split_composition_entity) and the reversal path (undo_composition_event), so an agent can distinguish it from related composition tools without opening a schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It states the precondition (operator has confirmed via the lift-candidate view that the two entities are the same logical thing), the retry condition when new conflicts surface (refresh the candidate view and resubmit), and the explicit alternatives (skip_overapplication_gate vs acknowledged_third_party_subtrees; split_composition_entity as inverse). This is close to a full when/when-not/alternatives statement.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_assertionsList AssertionsA

List active assertions for a control or assumption.

Provide exactly one of control_id or assumption_id.

Returns a flat list of assertions. Each assertion carries an origin field: "own" for assertions submitted directly against this model's control or assumption, "inherited" for assertions contributed through model composition (composed models whose assertions apply here). Inherited assertions are included in the listing.

Each assertion also carries three INDEPENDENT verdict fields. Read them together — a passing tier check is not the same as sufficient evidence:

  • tier1_status — mechanical check: the named file, symbol, or pattern is actually there. "pass" | "fail" | "pending".

  • tier2_status — semantic check: the cited code meaningfully implements the claim. "pass" | "fail" | "pending".

  • coherence_status — advisory consistency signal across the control's evidence set. "pending" here does NOT block the control from verifying, does NOT mean a verdict is missing, and is NOT a reason to trigger a recompute.

An assertion can pass BOTH tiers while its control stays unverified, because verification is decided per CONTROL, not per assertion: a control verifies only when its assertions collectively cover every clause of the control description. Read get_sufficiency for that verdict; never infer it from the tier fields here.

Each assertion also carries covers (the objective or clause ids it was declared to prove; empty when undeclared) and, where the platform surfaces it, tier1_attested and evidence_provenance (whether the run that verified it was signed and by what class of identity).

ParametersJSON Schema
NameRequiredDescriptionDefault
model_idYesID of the threat model.
control_idNoID of the control (omit if using assumption_id).
assumption_idNoID of the assumption (omit if using control_id).
server_versionYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden, and it delivers: it discloses that inherited assertions are included, that the three verdict fields are independent, that coherence_status 'pending' does NOT block verification, and that passing tiers does not mean the control verifies. It also explains the meaning of origin, covers, tier1_attested, and evidence_provenance. This is rich behavioral context beyond the schema.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is long but every section earns its place: the one-of constraint, the origin field, the three verdict fields with their caveats, and the pointer to get_sufficiency. It is front-loaded with the core purpose and the most critical usage rule. Slightly verbose in the verdict-field section, but the warnings are important enough to justify the length.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool has an output schema, so return values are already structured, but the description still explains the semantics of the non-obvious fields (origin, tier1_status, tier2_status, coherence_status, covers, tier1_attested, evidence_provenance). It also covers the key pitfall (per-control verification) and routes to the right sibling. For a list tool with this complexity, nothing essential is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 75%, and the description adds the critical constraint that exactly one of control_id or assumption_id must be provided, which the schema only hints at via 'omit if using...'. It also clarifies the semantics of the returned fields, which helps the agent interpret results. It doesn't document server_version or model_id, but those are self-explanatory and the schema covers them.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource: 'List active assertions for a control or assumption.' It immediately distinguishes the tool's scope (active assertions, flat list) and names the two valid target types. It also differentiates from siblings like get_sufficiency and get_verification_report by explicitly saying verification is decided per control and should be read from get_sufficiency, not inferred here.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives an explicit selection rule: 'Provide exactly one of control_id or assumption_id.' It also tells the agent when NOT to use this tool for a verdict: 'never infer it from the tier fields here' and directs to get_sufficiency. This is clear routing guidance relative to siblings.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_attestationsList AttestationsA

List an assumption's attestation history. Read-only; no side effects.

Returns the chronological record of attestation events recorded against the assumption (each with its actor, timestamp, and status/expiry as recorded), so you can trace why the assumption is currently attested, expired, or never attested. An assumption only mitigates its control objectives while it is active AND currently attested, so use this to diagnose coverage that depends on an attestation.

To record a new attestation use submit_attestation; for the assumption's current fields (status, description) use get_entity (entity_type="assumption").

ParametersJSON Schema
NameRequiredDescriptionDefault
model_idYesID of the threat model.
assumption_idYesID of the assumption whose attestation history to list.
server_versionYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the burden of behavioral disclosure. It explicitly states 'Read-only; no side effects' and describes the chronological return of events with actor, timestamp, and status/expiry. This is strong transparency beyond what the schema shows, though it does not mention potential limits or pagination.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is well-structured and front-loaded: the core action appears in the first sentence, followed by return-value context, a practical diagnostic purpose, and clear alternatives. Every sentence contributes information an agent needs to select and invoke the tool correctly.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool has an output schema and only three self-explanatory parameters, so the description need not expand return types. It explains what the returned history contains, why the tool matters for coverage diagnosis, and how it relates to neighboring tools. This is sufficient for correct selection and invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 67%, leaving server_version undocumented, but the parameter name is self-explanatory. The description clarifies the role of assumption_id by explaining that it identifies the assumption whose history is returned. It does not add detail about model_id or server_version, but the schema already handles most of the meaning.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource: 'List an assumption's attestation history.' It clearly distinguishes this from the many sibling tools by focusing on the attestation timeline of an assumption, and explicitly notes it is read-only with no side effects.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives concrete guidance on when to use the tool: to diagnose coverage that depends on an active attestation. It also names alternatives explicitly — submit_attestation for recording a new attestation and get_entity for current assumption fields — making the routing decision clear.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_co_dispositionsList Co DispositionsA

List the signed judgments recorded against this model's control objectives — risk acceptances, not-applicable dispositions, or both.

Read-only. Each entry carries the objective it names, the owner who signed it, the justification, the dates, and its status. Expired and revoked entries are included: a decision that lapsed is part of the audit trail, and hiding it would leave a reader unable to tell a judgment that was reviewed from one that was never made.

Read this before authoring a new judgment on an objective — an existing one may already cover it, or may have expired and need re-signing rather than duplicating.

ParametersJSON Schema
NameRequiredDescriptionDefault
kindNoOptional filter — "risk_accepted" or "not_applicable". Omit for both. Case and surrounding whitespace do not matter. A value that is neither is rejected by name rather than matched against nothing, so a typo cannot come back as an empty list you would read as "none recorded".
model_idYesID of the threat model.
server_versionYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.3/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden and does well: it declares the operation read-only, lists the fields each entry carries, and explains that expired and revoked entries are intentionally included for auditability. This goes beyond the schema and gives the agent important behavioral context.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is well structured and front-loaded: purpose first, return contents and inclusion behavior second, usage guidance last. Each sentence adds value, including the rationale for including expired and revoked entries, without excessive verbosity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given that an output schema exists and the description covers purpose, safety, return contents, and when to use it, the agent has enough context to invoke this tool correctly. Server_version is self-explanatory from the schema, and no pagination details are necessary for a list operation of this scope.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The description aligns with the schema's kind parameter by mentioning 'risk acceptances, not-applicable dispositions, or both', but the schema already documents kind in detail. It adds little for model_id and nothing for server_version, and with 67% schema coverage the description does not substantially compensate for the undocumented parameter.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: it lists signed judgments recorded against a model's control objectives, including risk acceptances and not-applicable dispositions. It is clear but does not explicitly differentiate itself from sibling tools like list_risk_acceptances, so the distinction is mostly implicit.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description tells the agent to read this before authoring a new judgment, warning that an existing or expired judgment may already cover the objective. It provides clear context for when the tool is useful, though it does not explicitly state when not to use it or name alternative tools.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_compliance_frameworksList Compliance FrameworksA

List the compliance frameworks available to map controls against.

Read-only; no side effects. Returns both built-in frameworks (e.g. OWASP ASVS) and any custom frameworks in the workspace. Use this to discover framework identifiers before select_compliance_frameworks (activate one for a model) or import_compliance_framework (add a custom one). Takes no arguments beyond the version guard.

ParametersJSON Schema
NameRequiredDescriptionDefault
server_versionYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full disclosure burden. It states 'Read-only; no side effects' and specifies the return scope: built-in frameworks (e.g. OWASP ASVS) plus any custom workspace frameworks. It could add slightly more nuance about the returned list's contents, but this is strong disclosure for a list tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Four concise sentences, each earning its place: purpose, behavior, usage guidance, and parameters. The core purpose is front-loaded and nothing is filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a single-parameter, read-only list tool with an output schema present, the description covers purpose, behavior, workflow placement, and the parameter's role. Return-value details are covered by the output schema, so nothing an agent needs to invoke it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% (server_version has no schema-level description), so the description must compensate. It clarifies that the parameter is a 'version guard' and not a functional argument, which helps, but it does not specify the expected version value or format, leaving the agent to guess what to pass.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource: 'List the compliance frameworks available to map controls against.' It distinguishes itself from siblings by naming select_compliance_frameworks (activate one) and import_compliance_framework (add a custom one) as the tools that follow on from this discovery call.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly states the workflow: 'Use this to discover framework identifiers before select_compliance_frameworks ... or import_compliance_framework.' It names both alternatives with their purposes, so an agent knows exactly when to call this tool instead of the related ones.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_control_revisionsList Control RevisionsA

List every change to a model version's set of controls. Read-only.

Each write to a version's published controls — a build's publish, an import, an edit, a deletion, an undo — is a set revision with its author. Returns {model_id, model_version, latest_version, discarded, revisions, undo_target}; each revision carries revision, job_id (the build that wrote it, if any), started_by, started_at, controls (the ids it touched), undo_of (the revision it undid, for an undo) and undone_by / undone_at. undo_target is the revision undo_model_change(target="controls") would undo (null when none, and for any version but the latest). discarded is true for a version a revert replaced.

ParametersJSON Schema
NameRequiredDescriptionDefault
versionNoModel version to read (0 = the latest live version).
model_idYesID of the threat model.
server_versionYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does well: it declares 'Read-only' outright and explains the semantics of non-obvious state such as 'discarded' (version replaced by a revert), 'undo_of'/'undone_by', and when 'undo_target' is null. It stops short of stating pagination, ordering, or permission requirements for reading another version.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the one-line purpose and read-only flag, then a dense but well-organized enumeration. The long field-by-field walkthrough overlaps with the existing output schema, so some content is not strictly earning its place, but the prose is tight and skimmable.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a read-only list tool with a 3-parameter schema and an output schema, the description covers what it does, its safety profile, and the meaning of ambiguous return flags. Since an output schema exists, the extended return-value enumeration is largely redundant rather than necessary, leaving little genuinely missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 67% (version and model_id documented, server_version not), so the schema does most of the work. The description's only parameter-adjacent statement is that undo_target is null 'for any version but the latest', which hints at the version parameter's behavior but adds no new input semantics beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('List every change to a model version's set of controls') with clear scope, and the read-only audit-history framing separates it from mutation siblings like apply_control_changeset or undo_model_change. It does not explicitly name an alternative listing tool (e.g. get_controls) to route against, but the resource is specific enough that an agent can distinguish it.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no explicit 'use this when / not when' guidance. The only usage signal is the cross-reference that 'undo_target' is the revision undo_model_change(target="controls") would undo, which implies this tool is used to find an undo target, but the agent must infer that.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_decisionsList DecisionsA

List the decision ledger of a model: every judgment recorded on it (finding dismissed or remediated, risk accepted, not-applicable declared, proposal accepted / rejected / reverted, assumption accepted, escalation resolved), newest first, with who made it and whether it was within the workspace's delegation policy. Read-only; no side effects.

Call this BEFORE raising a proposal or asking for a judgment, so you do not propose what a person rejected or ask again for what was already decided.

The ledger is append-only. There is no tool that edits it; to undo an accepted proposal, revert or re-decide, never edit the record. Rows with outcome == "refused" are judgments a program was refused; their escalation, if any, is in list_proposals. agent is null for a person's decision.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNoMaximum rows to return. 0 (default) uses the server default.
decisionNoOptional kind filter, one of ``finding_dismissed``, ``finding_remediated``, ``risk_accepted``, ``not_applicable_declared``, ``proposal_accepted``, ``proposal_rejected``, ``proposal_reverted``, ``assumption_accepted``, ``escalation_resolved``. Empty (default) returns every kind.
model_idYesID of the threat model.
agent_onlyNoOnly decisions made by a program (``agent`` not null). Default False.
server_versionYes
outside_policy_onlyNoOnly decisions made outside the delegation policy in force at the time (``within_policy`` false). Default False.

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.9/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Despite no annotations, the description discloses scope ('every judgment recorded'), ordering, read-only/no-side-effect status, and the append-only ledger invariant with a prohibition ('never edit the record'). It also explains the ``outcome == "refused"`` edge case and the null ``agent`` field, none of which the schema conveys.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core action, followed by the timing rule and then the append-only invariant. Every sentence carries operational value; no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Covers what the tool returns (fields, ordering, policy flag), the mutation boundary (append-only, use revert/re-decide), and the cross-tool handoff for refusals. Output schema exists, so return-shape explanation is correctly omitted, and no critical detail is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 83%, so the baseline is 3; the description goes beyond it by naming the decision kinds that populate the ``decision`` filter, explaining the ``outcome == "refused"`` semantics, and noting ``agent`` is null for a person's decision. It does not document ``limit`` or ``server_version`` explicitly, keeping it below a 5.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Names a specific verb (list) and resource (decision ledger of a model), then enumerates the exact decision kinds covered and the ordering ('newest first'). An agent can separate it from list_proposals and list_risk_acceptances without opening a schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives an explicit pre-call rule ('Call this BEFORE raising a proposal or asking for a judgment') plus the reason, and routes the refused-outcome case to list_proposals. Both the when-to-use and the when-to-go-elsewhere are stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_findingsList FindingsA

List negative findings recorded on a threat model. Read-only.

Returns finding rows with their lifecycle status; use to triage gaps or to find a finding_id for update_finding / remediate_finding. Each row carries an origin ("own" for findings recorded on this model, "inherited" for findings contributed through model composition, with inherited_from_* context); inherited findings are included in the listing.

ParametersJSON Schema
NameRequiredDescriptionDefault
statusNoOptional lifecycle filter, one of "discovered", "acknowledged", "remediated", "verified", "dismissed", "auto_resolved". Empty (default) returns all statuses. ``auto_resolved`` is closed by the platform, not by a person: the condition that produced the finding is no longer reproduced. It is deliberately distinct from ``remediated``/``verified`` (a person fixed and confirmed it) and from ``dismissed`` (a person judged it not worth fixing) — "the gap is gone" and "the gap does not matter" are opposite statements about residual risk, so they never share a status.
model_idYesID of the threat model.
control_idNoOptional filter to findings on one control. Empty (default) returns findings for all controls.
server_versionYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the behavioral burden and does so reasonably: it declares 'Read-only,' describes returned lifecycle status, and explains origin semantics including that inherited findings are included. It does not cover pagination, permissions, or rate limits, but it is substantially more transparent than a bare list operation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is three tightly structured sentences with no wasted wording. It front-loads the core operation and read-only nature, then adds usage and row-origin context in descending priority.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Because an output schema exists, the description need not explain return fields in full, and it still usefully clarifies the origin distinction for inherited findings. It is complete for a read-only list tool, though it omits edge-case behavior such as pagination or empty results.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 75%, so the schema largely documents the parameters, including the rich status filter semantics. The description adds no direct parameter-level guidance for status, model_id, or control_id, so a baseline 3 is appropriate rather than credit for compensating beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: 'List negative findings recorded on a threat model.' It also distinguishes this read-only listing from write-oriented siblings by naming update_finding and remediate_finding as consumers of the returned finding_id. An agent can identify what the tool does without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives clear usage contexts: 'use to triage gaps or to find a finding_id for update_finding / remediate_finding.' It does not explicitly state when not to use this tool or name a direct alternative such as get_findings_risks, so it falls short of full when/when-not guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_groupsList GroupsA

List the workspace's tags. Read-only; no side effects.

A tag is a named, overlapping grouping of threat models, for audit scopes, products, ad-hoc selections or portfolios. A model may carry many tags, and a tag never affects posture or credit. Returns {"tags": [...]}, each with id, name, description and model_ids.

Discover tag IDs here before the tag risk, compliance, dependency or export tools, or before adding/removing members. For a single model's tags use list_model_groups.

ParametersJSON Schema
NameRequiredDescriptionDefault
server_versionYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the burden, and it does disclose that the call is read-only with no side effects and that a tag never affects posture or credit. It also describes the return payload shape. It does not cover auth requirements or pagination, so it stops short of full transparency.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the action and read-only trait, then adds tag semantics, return shape and routing. Every sentence adds value with no padding.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Tag semantics, safety profile, return shape and sibling routing are all covered, and an output schema exists so return-value detail is not strictly required. The one gap is the undocumented required server_version parameter.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% for the single required parameter server_version, and the description never mentions it or its format/expected value. With an undocumented required parameter, the description does not compensate for the coverage gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('List the workspace's tags') and immediately distinguishes itself from the sibling list_model_groups by scoping. An agent can tell it apart from the tag-risk/export tools without opening a schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly routes the agent: discover tag IDs here before the risk, compliance, dependency or export tools, or before add/remove membership; and it names the alternative (list_model_groups) for a single model's tags. When-to-use and alternative are both stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_model_groupsList Model GroupsA

List the tags a given model belongs to. Read-only; no side effects.

A model may belong to many tags. Returns {model_id, tags: [...]}. Use list_groups for every tag in the workspace.

ParametersJSON Schema
NameRequiredDescriptionDefault
model_idYesthe model whose tags to list.
server_versionYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.9/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden, and it does declare 'Read-only; no side effects' plus the return shape. However it omits auth/permission requirements, pagination, and behavior when the model has no tags, so it is adequate rather than rich.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short sentences, front-loaded with the purpose, then the read-only note, then the return shape and alternative. Every sentence earns its place with no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Output shape is covered by the description and an output schema exists; the read-only nature and the list_groups alternative are stated. The main gap is the unaddressed server_version parameter, which leaves an agent guessing about a required argument.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is only 50%: model_id is documented in the schema while server_version is undocumented anywhere, and the description adds no parameter-level detail beyond implying a single model. With low coverage the description should compensate but does not.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('List the tags a given model belongs to') and explicitly contrasts with the sibling list_groups ('for every tag in the workspace'). An agent can distinguish this per-model listing from the workspace-wide listing without opening either schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Names an alternative (list_groups) and the condition that selects it (wanting every tag in the workspace vs. one model's tags). No explicit 'when not to use' or mention of the add/remove_model_from_group siblings, but the routing guidance is clear.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_proposalsList ProposalsA

List proposals and escalations on a model. Call this to poll the outcome of a proposal you raised, or of a judgment you were refused. Read-only; no side effects.

Statuses: proposed and applied_pending_review are open; accepted, rejected, reverted, superseded are closed. Kinds include add_component, remove_component, design_change, assumption, and decision_request: an escalation of a judgment this agent was refused. A 403 from update_finding, create_risk_acceptance, submit_attestation or decide_proposal carries an escalation_id; that escalation appears here as a decision_request. Poll it here until a person resolves it; do not retry the refused call.

ParametersJSON Schema
NameRequiredDescriptionDefault
statusNoOptional status filter (one of the values above). Empty (default) returns every proposal.
model_idYesID of the threat model.
server_versionYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.6/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden and does well: it declares 'Read-only; no side effects' and explains the open/closed status semantics and how a 403 with an escalation_id surfaces as a decision_request here. It stops short of describing pagination or result volume, but the behavioral core is unusually well covered for an unannotated tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with purpose, then usage, then reference material (statuses, kinds, escalation mechanics) in a fairly tight block. It is dense and slightly long, but nearly every sentence earns its place; the escalation paragraph is the only part that runs long.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return-value detail is unnecessary. The description covers the domain vocabulary (statuses, kinds), the escalation workflow, and the cross-tool trigger (403 from update_finding/create_risk_acceptance/submit_attestation/decide_proposal), leaving nothing essential for correct invocation missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 67% and no parameter enums exist in the schema, so the description compensates by enumerating the valid status values and the kinds, and by explaining that an empty status returns everything. model_id is covered by the schema and server_version is undocumented, so it falls short of a 5 but clearly adds meaning beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource (list proposals and escalations on a model) and clearly distinguishes itself from the create_proposal/decide_proposal/list_decisions siblings. An agent can tell this is a read-side poll of proposal state without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly says when to call it (to poll the outcome of a proposal you raised, or of a judgment you were refused), and gives a negative instruction ('do not retry the refused call') plus the escalation polling loop until a person resolves it. This is prescriptive routing, not implied.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_reconciliation_candidatesList Reconciliation CandidatesA

Entities a child model authored that look like ones it inherits. Read-only.

Use on a child model in a recursive tree to find duplicates before they distort coverage; decide each with decide_reconciliation_candidate.

  • disposition="active" (default): the open queue, paginated (page, page_size). {model_id, flag_enabled, total, tiers: {certain, heuristic}, page, page_size, candidates: [{kind, own_qid, inherited_qid, tier, reasons}]}. Tier certain is a deterministic match, safe to apply; heuristic is a fuzzy name/description match that needs review. Rejected pairs are left out.

  • disposition="rejected": the pairs recorded as NOT duplicates, oldest first and not paginated: {model_id, flag_enabled, rejections: [{id, model_id, kind, own_qid, inherited_qid, rejected_by, rejected_at}]}. An id is what an unreject names.

Where composition is not available both come back empty with flag_enabled: false.

ParametersJSON Schema
NameRequiredDescriptionDefault
pageNo
model_idYes
page_sizeNo
dispositionNoactive
server_versionYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does so well: it declares read-only, describes the full return shapes for both dispositions, explains the 'certain' vs 'heuristic' tier semantics and that certain is safe to apply, notes rejections are excluded, and discloses the composition-unavailable edge case returning empty with flag_enabled: false.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the purpose and read-only hint, then uses bullets for each disposition with inline response shapes. Dense but every clause adds selection or behavioral information; slightly heavy given how much overlaps the output schema.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 5-parameter read tool with an output schema, the description covers trigger conditions, both modes, tier meanings, the reject/unreject flow, and the degraded empty-result case. Nothing needed to call it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, so the description must compensate. It explains disposition values and the pagination contrast (active paginated via page/page_size, rejected not paginated), but model_id and server_version are never described and page/page_size semantics remain thin. Partial compensation only.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('list reconciliation candidates') and defines exactly what they are: entities a child model authored that duplicate ones it inherits. It distinguishes itself from the sibling decide_reconciliation_candidate by naming it as the follow-up action.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives an explicit when: 'on a child model in a recursive tree to find duplicates before they distort coverage.' It also routes the agent to decide_reconciliation_candidate for disposition, and explains which disposition value selects which result set.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_relianceList RelianceA

List a model's cross-model dependency edges, in both directions.

Read-only; no side effects. Returns {model_id, as_consumer: [...], as_provider: [...]}. Consumer edges are this model's declared delegations / reliances on other models' controls; provider edges are other models relying on this one (its blast radius if its controls change).

Use this to inspect existing dependencies before creating or deleting edges (manage_reliance / attach_foundation), or to understand what breaks if this model's controls change.

ParametersJSON Schema
NameRequiredDescriptionDefault
model_idYesID of the model to inspect.
server_versionYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the burden: it discloses 'read-only; no side effects' and explains the semantic meaning of both edge directions (consumer = declared delegations, provider = blast radius). It does not mention permissions/scoping or pagination, which are the remaining behavioral gaps for a full burden-of-proof case.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the operation and read-only guarantee, then return shape, then usage. The parenthetical about blast radius and the nested parentheses around sibling tool names add some density, but every sentence contributes meaning.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a read-only list tool with an output schema, the description covers purpose, usage, and interpretation of results (consumer vs provider edges). It is close to complete; only minor operational details (empty-model behavior, pagination/scoping) are absent.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 50%: model_id is documented in the schema, while server_version is undocumented in both places. The description implies model scope ('a model's') but adds no syntax, format, or semantics for either parameter, so the baseline 3 for schema-driven parameters applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('List a model's cross-model dependency edges') plus the scope ('in both directions'), and names the sibling mechanisms it relates to (manage_reliance / attach_foundation). An agent can distinguish it from generic dependency tools like get_system_dependencies or link_system_dependency.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly says when to use it: to inspect existing dependencies before creating or deleting edges, and to understand blast radius if controls change. It names the alternative tools (manage_reliance, attach_foundation) that perform the mutations, leaving no ambiguity about read vs write routing.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_risk_acceptancesList Risk AcceptancesA

List all risk acceptances on a specific threat model — risks that an operator explicitly accepted instead of mitigating.

Each entry carries the CO id, owner, justification, status (active / expired / revoked), and the review deadline. Use to inspect which gaps were intentionally accepted versus genuinely unaddressed when triaging at-risk COs.

Returns risk acceptances ONLY. An objective declared not applicable is a different claim — it is not an accepted risk, and counting it as one would read a "does not apply here" as "we are carrying this exposure". Use list_co_dispositions to see those, or both together.

ParametersJSON Schema
NameRequiredDescriptionDefault
model_idYesID of the threat model.
server_versionYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden. It discloses entry fields, statuses, and the key semantic boundary that not-applicable objectives are excluded. It does not explicitly note side-effect-free read behavior, but 'List' strongly implies it, and the output schema covers the return structure.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is front-loaded with the core function, then explains usage, and ends with an important cautionary note distinguishing acceptances from N/A dispositions. The length is justified by the conceptual subtlety, though it could be tightened slightly.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values are covered. The description explains scope, entry content, and its relationship to list_co_dispositions. Missing only server_version semantics and typical list-tool constraints like pagination, which are non-essential for correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is only 50%: model_id has a schema description while server_version does not. The description reinforces model_id with 'on a specific threat model' but never explains server_version, leaving a gap in parameter understanding.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: 'List all risk acceptances on a specific threat model.' It also explicitly distinguishes this from list_co_dispositions by clarifying that non-applicable objectives are not risk acceptances, making sibling differentiation clear.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It provides explicit usage context: 'Use to inspect which gaps were intentionally accepted versus genuinely unaddressed when triaging at-risk COs.' It also names the alternative tool for non-applicable objectives and suggests using both together, giving the agent clear selection guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_threat_modelsList Threat ModelsA

List saved threat models in the current workspace.

Read-only; no side effects. Returns {items: [{id, title, version, created_at, ...}], count}. Use this to discover model IDs to pass to other tools, or for a portfolio overview.

ParametersJSON Schema
NameRequiredDescriptionDefault
sourceNoFilter by the system that created each model. One of "web", "mcp", "jira", "api". Omit (default "") to list all models regardless of source.
server_versionYes
include_assessment_summaryNoIf True, include an `assessment_summary` object per model (counts of mitigated / at_risk / unassessed control objectives plus a human-readable `message`). Use for aggregate posture queries across the workspace in a single call (e.g. "which of my models are at risk?") instead of calling `assess_model` once per model. Adds roughly 100 bytes per model. Default False.

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations present, the description must carry the burden of behavioral disclosure. It explicitly states 'Read-only; no side effects' and describes the return shape as {items: [{id, title, version, created_at, ...}], count}. This is strong transparency for a list operation, though it does not mention pagination, authentication, or rate limits.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two sentences with no wasted words. It front-loads the action and resource, then states side-effect behavior, return shape, and intended use in a compact, scannable way.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description is largely complete for a read-only list tool: it covers purpose, side-effect profile, return shape, and primary use cases. The main missing context is an explicit note about the required server_version parameter and guidance distinguishing this from get_threat_model for detailed model retrieval.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 67%, covering source and include_assessment_summary, but the required server_version parameter has no description. The main description adds no extra parameter-level meaning beyond the schema, so the undocumented required parameter remains a gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states it lists saved threat models in the current workspace, with a specific verb and resource. It does not explicitly differentiate itself from related siblings like get_threat_model or query_threat_model, but the term 'list' plus the focus on model IDs and portfolio overview make the core purpose unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives concrete use cases: discovering model IDs to pass to other tools and portfolio overview. It does not explicitly say when to prefer alternatives such as get_threat_model or query_threat_model, but the use cases are helpful enough to orient an agent.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

manage_relianceManage RelianceA

Create, confirm or delete one cross-model reliance edge. Mutating.

action="create" declares that model_id relies on a provider control (the target is ALWAYS a control, so credit ends at a proven mechanism). mode is delegated (this model does not implement the objective; pass source_objective_id) or relied_upon (this model's own control depends on the provider's; pass source_control_id). The provider must be in the same workspace. The edge enters draft, runs LLM semantic validation, and carries no credit until confirmed. Returns the edge.

action="confirm" promotes the draft edge_id to active, the credit-soundness gate: refused unless validation returned valid; a partial result or a mode mismatch is never silently credited. Returns the edge.

action="delete" permanently removes edge_id, withdrawing any credit the consumer derived from it (its coverage can move); neither model's controls change. Returns {deleted: true, edge_id}.

list_reliance shows a model's edges and their ids.

ParametersJSON Schema
NameRequiredDescriptionDefault
modeNo
actionYes
edge_idNo
model_idNo
server_versionYes
provider_model_idNo
source_control_idNo
provider_control_idNo
source_objective_idNo
accept_partial_as_relied_uponNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does so richly: create enters a draft state, runs LLM semantic validation, and carries no credit until confirmed; confirm is refused unless validation returned valid and never silently credits a partial or mode mismatch; delete permanently removes the edge and withdraws derived credit without altering either model's controls. These are exactly the side effects and gating rules an agent needs.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Well structured with a front-loaded summary followed by per-action paragraphs, so an agent can jump to the relevant action. It is somewhat verbose and repeats 'Returns the edge' three times, which slightly dilutes density.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given an output schema exists, return values need not be spelled out, yet the description still notes the edge object and the {deleted, edge_id} shape. Combined with the per-action lifecycle and validation gating, nothing material is missing for correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0% across 10 parameters, so the description must compensate and largely does: it defines mode values (delegated vs relied_upon) and which source id each requires, the edge_id used by confirm/delete, and the same-workspace constraint on the provider. It omits accept_partial_as_relied_upon and server_version, leaving a small gap, so not a full 5.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Opens with a specific verb set and resource ('Create, confirm or delete one cross-model reliance edge') and immediately flags the mutating nature. It explicitly distinguishes itself from the sibling list_reliance, which it names as the read path. An agent can tell exactly what this tool does without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Each action is given its own condition: create declares a reliance with a mode-dependent source id, confirm is the credit-soundness gate that requires valid validation, delete withdraws derived credit. It also routes to list_reliance for retrieving edge ids. It lacks explicit 'when not to use' guidance, so it falls just short of a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

map_control_to_requirementMap Control To RequirementA

Manually map one security control to one compliance-framework requirement. Mutating: records a control-to-requirement mapping, which re-derives that requirement's coverage in the compliance report.

Use for a single, deliberate mapping you are asserting by hand. To let the LLM propose mappings across many requirements at once, use auto_map_controls; to close gaps end-to-end (map + exclude + fill), use auto_remediate_compliance.

ParametersJSON Schema
NameRequiredDescriptionDefault
notesNoOptional free-text note explaining the mapping rationale.
model_idYesID of the threat model.
confidenceNoProvenance label recorded on the mapping: "manual" (default, operator-asserted), "llm" (machine-suggested), or "verified" (human-confirmed).manual
control_idYesID of the control to map (e.g. "CTRL-01").
framework_idYesID of the compliance framework.
requirement_idYesID of the requirement to map to (e.g. "V2.1.1").
server_versionYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations exist, so the description carries the full behavioral burden. It opens with 'Mutating:' and explains the consequence: recording the mapping re-derives that requirement's coverage in the compliance report. It does not state idempotency or overwrite behavior, but the core side effect is clearly disclosed.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences: the first states the action and side effect, the second gives usage guidance and alternatives. No filler, and the most important information is front-loaded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With an output schema present and high parameter coverage, the description covers the key decision factors: what it does, that it mutates, the resulting report effect, and when to choose an alternative. Minor omissions like overwrite semantics and the unexplained required 'server_version' parameter keep it from a perfect score.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 86%, so the schema already documents most parameters with examples. The description adds only indirect context — 'asserting by hand' aligns with the confidence default 'manual' — but offers no new parameter semantics beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Description states a specific verb ('map'), resource ('one security control to one compliance-framework requirement'), and explicitly scopes to a single deliberate mapping. It distinguishes itself from siblings by naming auto_map_controls (many at once) and auto_remediate_compliance (end-to-end gap closure).

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly says 'Use for a single, deliberate mapping you are asserting by hand.' Names two alternatives with the exact conditions for choosing them: batch LLM proposals via auto_map_controls, and end-to-end gap closure via auto_remediate_compliance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

model_coherence_reportModel Coherence ReportA

How coherent a model's structure is: its component bindings, the repos its controls' assertions name, and whether every CO is structurally reachable. Read-only.

Pass co_id to keep only the findings about that CO (404 if it does not exist). Each finding carries type, severity, a message and the ids it concerns, so its fix can be called directly:

  • control_component_unknown, assertion_repo_orphan, control_unscoped_with_scoped_assertions: assign_to_components(target_type="control"). assertion_repo_mismatch: rebind the assertion or rescope the control.

  • asset_component_unknown: edit_asset with corrected component_ids.

  • component_unbound: for your own code, edit_component with the real repo_url. An external-zone component (a third-party service, a customer's IdP) stays unbound: the finding is a permanent marker of an external dependency, not a TODO, and client code touching it is no reason to bind it.

  • co_attacker_unpositioned: edit_attacker with trust_boundary_ids. co_asset_unbounded: assign_to_components(target_type="asset"). co_no_shared_boundary: reposition the attacker or rescope the asset; if the boundaries truly do not meet, that is the answer. co_missing_entity: restore_entity, or remove the CO.

An indeterminate reachability verdict means the structure it needs is missing: supply it. It never means the objective is inapplicable, which is a separate claim recorded with create_co_disposition. get_reachability_verdicts returns the raw verdicts.

ParametersJSON Schema
NameRequiredDescriptionDefault
co_idNo
model_idYes
server_versionYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.6/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the burden and does so well: it discloses the read-only nature, the 404-on-missing-CO behavior, the finding shape (type, severity, message, ids), and the nuanced semantics of an indeterminate reachability verdict. Remaining gaps (pagination, size limits, server_version semantics) are minor.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The purpose is front-loaded and the description is organized with bullets, but it is long and dense for a read-only report, devoting substantial space to finding-type-to-fix mappings that border on workflow documentation. The information is useful but not tightly scoped.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return-value detail is technically unneeded, yet the description still explains the finding structure and routes each finding to a fix. Combined with the reachability-verdict caveat, it is nearly complete, with the only real hole being the unexplained required parameters.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate, and it does so only for co_id (filter scope, 404 behavior, default = all). Both required parameters, model_id and server_version, are completely unexplained in the schema and the description, leaving the majority of parameters undocumented.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific purpose: reporting how coherent a model's structure is across component bindings, assertion-name repos, and CO reachability, and explicitly marks it read-only. It partially distinguishes itself from the sibling get_reachability_verdicts ('returns the raw verdicts'), but other report siblings (assess_model, get_verification_report) are not referenced.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives concrete usage context: pass co_id to scope findings, and each finding type is mapped to a specific remediation tool call (assign_to_components, edit_asset, edit_component, etc.), which is strong when-to-do-what guidance. It also clarifies the indeterminate-verdict case versus create_co_disposition. There is no explicit 'do not use this when...' exclusion, keeping it below a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

pause_control_generationPause Control GenerationA

Pause a model's background control generation. Mutating.

Use when the user asks to stop a build — for example one started by mistake — or before deleting a model whose controls are still being built. A running build stops at its next step (status pausing, then paused); a queued or waiting one is paused at once. Everything already done is kept in the build's staging copy; nothing is published, nothing new is started or billed, and nothing resumes it except resume_control_generation. A paused build still holds the model: to drop it instead, call discard_control_build once it shows paused. Pausing is idempotent.

Returns one of:

  • {paused: true, model_id, status, status_detail} — status is pausing (still stopping) or paused.

  • {paused: false, http_status: 409, code: "not_running", status} — there is no generation to pause; read status.

ParametersJSON Schema
NameRequiredDescriptionDefault
model_idYesID of the threat model whose control generation to pause.
server_versionYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden and does so: it discloses the state machine (pausing then paused), that queued builds pause at once, that completed work is only in the staging copy, nothing is published or billed, nothing auto-resumes, there is no undo other than resume, it holds the model while paused, and pausing is idempotent.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the action, the mutating flag, and the primary use case, then layers consequences and the return shape in clearly separated blocks. Dense but every clause (billing, idempotency, staging copy, alternatives) carries distinct information; only the length keeps it from a 5.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a mutation with no annotations, the description covers triggers, state transitions, side-effect boundaries, idempotency, recovery path, sibling disambiguation, and the two possible return shapes (including the 409 not_running case). Nothing needed to call it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 50%: model_id is documented in the schema, and the description's framing ('a model's background control generation') corroborates that target but adds no new semantics. server_version is undocumented in both places, and the description never compensates for it.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Pause a model's background control generation') plus the mutation nature in the first line. It is immediately distinguishable from resume_control_generation and discard_control_build, which are named later.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives explicit when-to-use triggers ('user asks to stop a build — for example one started by mistake — or before deleting a model whose controls are still being built') and names the two alternatives (resume_control_generation, discard_control_build) with the condition that selects each.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

query_threat_modelQuery Threat ModelA

Ask a natural-language question about an existing threat model.

Read-only; no side effects (no new version, no mutation). Uses AI to answer questions grounded in the model's assets, attackers, control objectives, assumptions, and current security posture, returning {model_id, answer} where answer is prose.

Use this for interpretation or summary questions ("what are the biggest gaps?", "which attackers target the token store?"). Do NOT use it to change the model — use refine_threat_model for that — and prefer get_threat_model / assess_model when you need structured data (entity lists, coverage counts) rather than a written answer.

ParametersJSON Schema
NameRequiredDescriptionDefault
model_idYesID of the threat model to query.
questionYesThe natural-language question to ask.
server_versionYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden, and it succeeds: it declares 'Read-only; no side effects (no new version, no mutation)', explains that AI is used, specifies that answers are grounded in model entities, and gives the return shape {model_id, answer}. This is unusually transparent.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is front-loaded with purpose, then behavior, then usage boundaries. Every sentence earns its place, and the three short paragraphs are easy to scan for an agent deciding whether to invoke this tool.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description is remarkably complete for a read-only query tool: it covers behavior, return shape, use cases, and sibling alternatives. The only notable omission is any explanation of the required server_version parameter, which prevents a perfect score.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The description adds useful semantics for the question parameter (interpretation/summary questions, with examples) and implies model_id must reference an existing model. However, the required server_version parameter is completely undocumented in both the schema and the description, leaving a clear gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource: 'Ask a natural-language question about an existing threat model.' It also distinguishes itself from likely siblings by explicitly saying it is not for mutation (refine_threat_model) and not for structured data (get_threat_model / assess_model).

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly states when to use the tool ('interpretation or summary questions'), when not to use it ('Do NOT use it to change the model'), and which alternatives to prefer for structured data. This leaves little to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

recompute_verdictsRecompute VerdictsA

Estimate, force, or retry a model's verdict evaluation.

  • mode="quote" (default): read-only. The cost of a recompute, enqueueing nothing: {estimated_credits, computed_at, rate_version, informational, total_enqueueable, already_evaluated, governor}. Subjects already carrying a verdict cost nothing, so it is an upper bound. Show the operator this number before recomputing: on a large model it runs to thousands of credits.

  • mode="recompute": mutating. Queues a fresh evaluation of every control's coverage verdict and every live objective's group-sufficiency verdict, bypassing quiet-period batching; usage is metered as it runs. Returns {model_id, model_version, enqueued_coverage, enqueued_group_sufficiency, total_enqueued, estimated_credits, quote, governor}.

  • mode="retry_parked": re-runs only the verdicts a transient failure (outage, exhausted credits, timeout) parked, of every kind, including per-control sufficiency and coherence; nothing else is touched. Returns {model_id, model_version, retried_slots, governor}.

Work runs in the background; when governor.exhausted it is queued and resumes at governor.resets_at, never dropped.

A recompute evaluates COVERAGE and GROUP SUFFICIENCY only. A control at partially_verified, or coherence_status: "pending" on an assertion, is not a reason to recompute: that verdict is computed on assertion write and read with get_sufficiency. Recompute when control-to-CO mappings look wrong (get_verdict_divergence). One objective awaiting judgement is judge_objective.

ParametersJSON Schema
NameRequiredDescriptionDefault
modeNoquote
model_idYes
server_versionYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations present, the description carries the full burden and does so: it labels mode='quote' read-only and mode='recompute' mutating, explains that usage is metered as it runs, that subjects already carrying a verdict cost nothing, that work runs in the background, and that a governor.exhausted run is queued and resumes at governor.resets_at rather than dropped. It also warns the operator to see the credit estimate first, which is exactly the kind of cost/side-effect disclosure an agent needs.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the verb and the three mode names, then bulleted per mode, with the caveats in a final block. Efficient for the amount of behavior being disclosed, though the per-mode return-object listings border on verbose and overlap the output schema.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a three-mode, mutating, cost-incurring tool with an output schema, the description covers mode selection, mutation/read-only semantics, cost, background execution, governor exhaustion behavior, and sibling alternatives. The only omission is model_id/server_version semantics, which is minor relative to the completeness elsewhere.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It fully documents the 'mode' enum (including the default and the distinct behavior/return shape of each value), but model_id and server_version — both required-adjacent inputs — get no explanation at all. Mode is the highest-value parameter and is well served, but the remaining gap keeps this at baseline-plus rather than high.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('Estimate, force, or retry') over a specific resource ('a model's verdict evaluation'), then breaks the three modes apart so an agent knows exactly which operation it is invoking. It explicitly differentiates itself from siblings like get_sufficiency, get_verdict_divergence, and judge_objective, so the agent can route without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives explicit when-to-use per mode (quote = read-only estimate, recompute = mutating full re-evaluation, retry_parked = only transient-failure-parked verdicts) and explicit when-NOT-to-use ('A control at partially_verified, or coherence_status: pending ... is not a reason to recompute') plus the correct alternatives for those cases. Nothing is left to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

reconcile_modelReconcile ModelA

Reconcile a threat model with the code it describes. Call this after reading the code and before (or instead of) editing the model by hand: report what changed and what you observed, and the platform decides the consequence of each observation. Mutating only where the platform applies an observation (see below).

Two inputs, both optional:

  • changed_paths: the file paths that changed since the model's recorded commit. For a code-derived model compute them with git diff --name-only <commit_sha>..HEAD (the commit_sha from the model's provenance). The platform maps them onto components and reports which components changed, which paths no component claims, and whether a refresh is recommended.

  • observations: what you saw in the code that the model does not say. Each observation lands in one of four buckets by kind:

    • mechanism_named - the control's mechanism exists under another name (subject_id = control id). Follow up with refine_control using the codebase_findings returned in refine_suggested.

    • component_present - the code has a component the model lacks; include a proposal ({name, repo_url?, path?, trust_boundary_ids?}).

    • component_absent - a modelled component has no code (subject_id = component id).

    • forbidden_behavior - the code does something the model rules out (subject_id = control id, or empty).

The platform decides the consequence. Proposals are never applied on the agent's word, with one exception: a component change on a code-derived model (provenance kind="code") is applied immediately and queued for a person's review as applied_pending_review. Every other proposal waits for decide_proposal. Forbidden behaviors become findings. Observations the platform could not use come back in ignored with the reason.

ParametersJSON Schema
NameRequiredDescriptionDefault
model_idYesID of the threat model.
repo_urlNoRepository the paths belong to (optional; helps map paths in multi-repo models).
observationsNoJSON string of an **array** of observation objects, each ``{kind, subject_id?, evidence?: {paths?: [], symbols?: [], note?: ""}, proposal?: {name, repo_url?, path?, trust_boundary_ids?}}`` with ``kind`` one of ``mechanism_named``, ``component_present``, ``component_absent``, ``forbidden_behavior``. Empty/None sends no observations.
changed_pathsNoComma- or newline-separated file paths that changed since the model's recorded commit. Empty/None skips path mapping (``changed_paths`` in the response is then null).
server_versionYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description fully carries the behavioral transparency burden. It discloses mutation semantics ('Mutating only where the platform applies an observation'), the immediate-apply exception with 'applied_pending_review', the fallback to decide_proposal, forbidden behaviors becoming findings, and ignored observations with reasons. This is exemplary disclosure of side effects.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is long but densely structured with a clear lead sentence, labeled optional inputs, and bullet-like kind explanations. Every section earns its place, and the most important guidance appears first. The formatting makes a complex set of behaviors scannable.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a complex reconciliation tool with no annotations, the description is remarkably thorough: inputs, observation buckets, consequences, exceptions, and ignored outcomes are all covered. The only notable omission is that server_version, a required parameter, is not explained anywhere; otherwise, an agent has enough context to invoke the tool and interpret its role in the workflow.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The description adds substantial meaning beyond the schema, especially for changed_paths (including the git diff command and provenance commit_sha) and observations (explaining each kind's meaning, required fields, and consequences). However, server_version is a required parameter with no schema description and is not mentioned in the description, leaving a small gap despite the 80% schema coverage.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource: 'Reconcile a threat model with the code it describes.' It clearly distinguishes this tool from hand-editing and from related sibling tools like refine_control and decide_proposal by explaining that the platform decides consequences rather than the agent directly editing.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicit usage guidance is front-loaded: 'Call this after reading the code and before (or instead of) editing the model by hand.' It also names follow-up actions ('Follow up with refine_control...'), states when proposals wait for decide_proposal, and explains the exception for code-derived models. This gives an agent clear decision rules.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

reevaluate_threat_model_factorsReevaluate Threat Model FactorsA

Re-run the LLM factor judgment on every asset and attacker in a threat model. Useful for re-baselining factors after a bug fix or feature-description change, without regenerating the whole model (which would destroy controls, assertions, components).

Each entity's factors and rationale are replaced with a fresh LLM-judged decomposition; the composed impact / likelihood is re-derived deterministically from the new factors. Each re-rating is recorded as a rating revision in the audit trail with change_reason (default: "LLM factor re-evaluation") so the starting-point regeneration is distinguishable from operator- supplied factor overrides via edit_asset / edit_attacker.

The platform's LLM factor judgment is a starting point. For deployment-specific factor adjustments (e.g., elevated regulatory_scope because your tenant is HIPAA-covered, or Commodity prevalence because your endpoint is public-internet exposed), use edit_asset / edit_attacker afterward with a change_reason documenting the operator override.

Per-entity soft-fail: an LLM failure on one entity is recorded in the response's failed_entities list (with id, kind, and reason); the remaining entities are still re-evaluated and their rating revisions persisted as they complete. The endpoint returns 503 only when every live entity failed — in which case nothing was persisted; retry when the evaluator is reachable.

Soft-deleted assets and attackers are skipped.

ParametersJSON Schema
NameRequiredDescriptionDefault
model_idYesID of the threat model to re-rate.
change_reasonNoOptional override of the audit-trail reason (default: "LLM factor re-evaluation"). Use this to thread a higher-level reason like "Re-eval after refinement bug fix shipped in vN.N.N" when running the tool as part of a broader workflow.
server_versionYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden and does so thoroughly. It discloses that factors and rationale are replaced, that impact/likelihood is re-derived, that rating revisions are audited, the default change_reason, the soft-fail per-entity behavior, the 503 condition, and that soft-deleted entities are skipped. This is exemplary transparency.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is long but well-structured: purpose first, then usage guidance, then behavioral details, then failure semantics. Each paragraph adds necessary information for a mutating re-rating operation. Minor redundancy and length prevent a 5, but the structure is strong.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description covers purpose, alternatives, side effects, audit behavior, failure modes, and exclusion of soft-deleted entities. The presence of an output schema reduces the need to explain return shapes. However, the required server_version parameter remains unexplained, which is a real gap for correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The description adds useful context for change_reason, explaining its default and audit-trail purpose, and reinforces model_id as the threat model to re-rate. However, server_version is a required parameter with no schema description and no explanation in the description, leaving its semantics unclear. The description partially compensates but does not fully cover the parameter set.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource: 'Re-run the LLM factor judgment on every asset and attacker in a threat model.' It clearly differentiates this from regenerating the whole model and from operator edits via edit_asset / edit_attacker, so an agent can identify when this tool is the right fit.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It explicitly states when to use the tool: re-baselining after a bug fix or feature-description change. It also names the alternatives for other cases: use edit_asset / edit_attacker for deployment-specific factor adjustments. This is concrete, actionable routing guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

refine_controlRefine ControlA

Refine a control's description with AI-gated CO sufficiency check.

Two modes:

  • Provide description: proposes a new description directly.

  • Provide codebase_findings: the platform proposes a description based on existing code that may already satisfy the control.

  • Both can be provided: the platform evaluates the proposed description with the codebase findings as context.

The AI evaluates whether the mitigation group still collectively satisfies all mapped control objectives. If rejected, returns {accepted: false, reason, per_co} with per-CO reasoning.

A refinement is rejected when the proposed description would reduce the protection the control currently states for an objective it is mapped to; per_co names each objective and explains why. This is a decision, not a transient error — re-wording the same narrowing will not pass it, and it applies however well-motivated the narrowing is. A control is a requirement that must be met to cover its objectives, so evidence that the system does not currently meet it means the control is UNMET, never that the control should ask for less.

After an accepted refinement the control's assertions are kept and judged again against the new description in the background: an assertion that still fits keeps counting as evidence, and one that no longer fits is flagged as not aligned with the control. Read get_sufficiency once that re-judgement lands, and replace the assertions it names. The refinement itself supersedes nothing; the response's superseded_assertions is always 0.

ParametersJSON Schema
NameRequiredDescriptionDefault
model_idYesID of the threat model.
control_idYesID of the control to refine (e.g., "CTRL-03").
descriptionNoProposed new control description (optional if codebase_findings provided).
justificationNoWhy this refinement is appropriate (10 to 2000 characters).
server_versionYes
codebase_findingsNoDescription of existing code that may already satisfy this control's objective (optional). When provided without description, the platform proposes a description.

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.3/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does so well: it discloses the rejection response shape ({accepted, reason, per_co}), the rejection rule (narrowing protection for a mapped objective), that rejection is a decision rather than a transient error that re-wording will not pass, and the background re-judgement of assertions after acceptance including that `superseded_assertions` is always 0.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Purpose is front-loaded in the first line, then modes, then rejection semantics, then the post-acceptance workflow — a logical progression. It is on the long side and the closing note about `superseded_assertions` being always 0 is somewhat tangential, but nearly every sentence carries operational content.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a high-complexity, AI-gated mutation tool, the description covers mode selection, the rejection contract, the non-retryable nature of rejection, and the required follow-up call to get_sufficiency. An output schema exists, so return-value enumeration is not needed, and the description still names the rejection payload fields.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is already 83%, so the schema defines most parameters. The description nonetheless adds real meaning beyond it by explaining the interaction between `description` and `codebase_findings` (each alone triggers a different mode; together they combine), which the schema only hints at with 'optional if codebase_findings provided'.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('Refine a control's description') plus the gating mechanism ('AI-gated CO sufficiency check'), and breaks out the two operating modes by parameter. It is clear what the tool does, though it never names which sibling to use instead (e.g., regenerate_controls, strengthen_controls, remap_control), so differentiation is inferential.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly tells the agent which parameters select which behavior: provide `description` for a direct proposal, `codebase_findings` for a platform proposal, or both for context-augmented evaluation. It also states the post-acceptance follow-up ('Read get_sufficiency once that re-judgement lands'). It lacks any when-not-to-use guidance relative to sibling tools.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

refine_threat_modelRefine Threat ModelA

Refine an existing threat model based on an instruction.

Updates the model's assets, attackers, trust boundaries, and control objectives based on the instruction. Creates a new version. Progress is reported automatically.

Refine CANNOT silently replace an entity's identity under a stable ID or silently drop an entity. Behavior:

  • Preserved entities where the LLM proposed an identity- bearing rewrite (name / description / security_properties on assets; capability / archetype / position on attackers) run through a semantic-preservation guard. Rewrites classified as replace or ambiguous (or unavailable if the gate LLM is down) have their identity fields REVERTED to the pre-refine values. Each rejection shows up as an entry in the semantic_rejections array in this tool's return value — surface these to the operator.

  • Entities the LLM drops from the refined output are re- appended to the model unchanged. The only sanctioned removal path is remove_entity (entity_type="asset") / remove_entity (entity_type="attacker") (soft-delete).

  • CO IDs are stable across refinements; pairs (asset, attacker) that disappear come back as tombstones with removed=True (not renumbered). Controls that only mapped to tombstoned COs become orphaned at read time.

ParametersJSON Schema
NameRequiredDescriptionDefault
model_idYesID of the threat model to refine.
instructionYesWhat to change, e.g. "Add CSRF attack vectors".
server_versionYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.3/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden and does so thoroughly. It discloses preservation guarantees, the semantic-preservation guard, reverting identity fields, semantic_rejections, re-appending dropped entities, stable CO IDs, tombstones, orphaned controls, and the sanctioned removal path through remove_entity.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is front-loaded with a concise summary and then structured into clear behavior bullets. The length is justified by the non-obvious preservation, revert, tombstone, and orphan semantics it must communicate; each sentence adds needed detail.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description covers purpose, side effects, edge cases, and the key return artifact (semantic_rejections), and an output schema is present to define return values. It provides enough behavioral context for an agent to invoke the tool correctly, especially given the complex preservation guarantees.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema already documents model_id and instruction, and the description adds some meaning by explaining what the instruction changes. However, server_version is a required parameter with no schema description and no mention in the tool description, so the parameter semantic coverage remains incomplete.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('Refine'), a specific resource ('an existing threat model'), and the exact elements updated (assets, attackers, trust boundaries, control objectives), plus the fact that it creates a new version. This clearly distinguishes it from generation or query siblings like generate_threat_model or get_threat_model.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The opening sentence implies the use case: refine an existing threat model based on an instruction. However, it does not explicitly say when to use this tool over siblings like generate_threat_model or edit_asset, and it does not provide exclusion conditions or alternative routing.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

regenerate_controlsRegenerate ControlsA

Propose a regeneration of the model's controls. Starts nothing.

A regeneration re-authors controls from the current COs; its publish creates the next model version. Controls whose descriptions survive unchanged KEEP their implementation status, evidence, notes, assertions, and Jira / compliance mappings. Controls whose descriptions change or disappear are soft-deleted (still queryable via get_controls(include_deleted=True)). When co_ids is given, only those COs' controls are regenerated — all other controls are left as-is.

This tool records the regeneration as the model's PROPOSED build and returns at once with status: "proposed" and proposal (mode, objective_ids, objective_count, estimated_credits, and the model_version and set_revision a start must name). A proposal merges with any already proposed for the model, the broader one winning. Show the user what it would build and cost; once they agree, call start_control_build with those values and confirm_estimate=True. A build someone started holds the model, so this is refused (409 generation_active) until it finishes, or is resumed and finishes, or is discarded.

To rebuild everything, omit co_ids. To fix only stale/orphaned CO mappings without re-authoring control text, prefer remap_control (mechanical, no LLM).

ParametersJSON Schema
NameRequiredDescriptionDefault
modeNo"batch" (default) or "per_co" (most thorough — one LLM call per CO).batch
co_idsNoOptional comma-separated CO IDs to regenerate (e.g. "CO1,CO5"). Omit to regenerate all controls.
model_idYesID of the threat model.
batch_sizeNoCOs per batch in batch mode (default 15). Smaller = more accurate and more granular progress, but more LLM calls.
server_versionYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden and does so: it discloses what survives (status, evidence, notes, assertions, mappings), what is destroyed (changed/removed controls soft-deleted, still queryable via get_controls(include_deleted=True)), merge semantics (broader proposal wins), and the immediate return shape (status "proposed" plus proposal fields). This is unusually complete behavioral disclosure for a mutation-adjacent tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the crucial distinction ("Starts nothing") and each subsequent sentence carries distinct information: retention rules, scoping, return payload, and the conflict path. It runs long across several paragraphs, but the density is high and little is redundant.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a proposal tool with an output schema, annotations absent, and 5 parameters, the description covers purpose, side effects, scoping, return contract, and the handoff to start_control_build. Nothing an agent needs to invoke it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 80%, so the schema already documents mode, co_ids, and batch_size. The description adds real meaning beyond it: co_ids scoping ("only those COs' controls are regenerated — all other controls are left as-is") and the omit-co_ids default ("To rebuild everything, omit co_ids"). batch_size and the batch/per_co tradeoff are left to the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ("Propose a regeneration of the model's controls") and immediately disambiguates from the actual build with "Starts nothing." It distinguishes itself from siblings like start_control_build and remap_control explicitly, so an agent can route without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Names the follow-up action and condition ("once they agree, call start_control_build with those values and confirm_estimate=True"), names an alternative for a narrower case ("prefer remap_control (mechanical, no LLM)"), and states the blocking condition (409 generation_active). Nothing is left to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

remap_controlRemap ControlA

Mechanical, non-AI-gated remap of a control's CO mappings.

Distinct from refine_control (AI-gated description edit) and set_mitigation_groups (AI-gated CO-centric group authoring). Use remap_control when the operator already knows the correct co_ids and just needs to persist the mapping change — e.g., restoring mappings after an asset/attacker edit left the control with stale or orphaned CO references. No LLM evaluation runs.

Rejects target co_ids that do not exist on the model or are tombstoned (the pair was removed in a later version) — map to live COs only.

ParametersJSON Schema
NameRequiredDescriptionDefault
co_idsYesComma-separated list of target CO IDs (e.g., "CO1,CO2,CO3"). Must include at least one CO.
model_idYesID of the threat model.
control_idYesID of the control to remap (e.g., "CTRL-03").
change_reasonYesWhy this remapping is appropriate (min 10 chars). Captured in the control's version history.
server_versionYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.6/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure. It does so well by noting that no LLM evaluation runs, that the operation is purely mechanical, and that non-existent or tombstoned target co_ids are rejected. It does not discuss permissions, irreversibility, or whether existing mappings are replaced, but the core behavioral constraints are clearly disclosed.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is compact and front-loaded: the first sentence states the exact purpose, the next sentences add sibling differentiation and usage context, and the final sentence documents validation behavior. Every sentence earns its place and there is no redundant filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's moderate complexity, the description covers purpose, usage, alternatives, and key rejection behavior. An output schema exists, so return values do not need to be explained. The only small gap is the unexplained server_version parameter and the lack of explicit statement about whether the remap replaces all existing CO mappings or merges with them, but the overall context is strong.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 80%, so the schema already documents most parameters. The description adds extra meaning for co_ids by clarifying that targets must be live COs and that tombstoned/nonexistent IDs are rejected, which goes beyond the schema's 'must include at least one CO'. The only notable gap is server_version, which remains undocumented in both schema and description.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource: 'Mechanical, non-AI-gated remap of a control's CO mappings.' It clearly names sibling tools it is distinct from (refine_control, set_mitigation_groups), eliminating ambiguity about what this tool does versus similar ones.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly states when to use this tool: when the operator already knows the correct co_ids and needs to persist a mapping change. It also provides a concrete example scenario (restoring stale/orphaned CO references) and distinguishes it from AI-gated alternatives. This leaves no doubt about selection.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

remediate_findingRemediate FindingA

Preview, and on confirmation apply, the platform's remediation of a finding.

Without apply it is read-only: it returns a structured diff of what the remediation would change, shaped by the finding's kind (for structural_duplicate_controls: which controls are kept, which dropped, and the CO mappings and framework refs the survivor takes). Show the operator that diff and get explicit confirmation.

With apply=True it commits the change, recording justification (a one-line operator rationale, required) on the audit trail. Never apply without having shown the preview: the platform records who acted but does not enforce the preview — the agent does.

404 when the finding does not exist; 422 when its kind has no automatic remediation (resolve those with the control tools); applying a finding already remediated or dismissed is refused (409).

ParametersJSON Schema
NameRequiredDescriptionDefault
applyNoFalse (default) previews; True applies.
finding_idYesID of the finding.
justificationNoWith ``apply=True``: the operator's rationale.
server_versionYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.9/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden and does so thoroughly: it discloses the read-only vs mutating nature, the audit-trail recording of justification, the fact that the platform records who acted but does not enforce the preview, and exact error conditions (404, 422, 409).

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is front-loaded with the core preview/apply distinction and then adds the confirmation, audit, and error details in tight, information-dense sentences. No sentence is wasted, and the structure mirrors the agent's decision path.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given a complex mutation with no annotations and an existing output schema, the description covers what an agent needs: preview vs commit, required confirmation, audit behavior, and failure modes. It also sketches the returned diff, so it is complete without duplicating the output schema.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 75%, so the schema already documents apply, finding_id, and justification. The description adds useful meaning for apply and justification by explaining the confirmation workflow and audit-trail requirement, but it does not mention server_version at all, leaving one required parameter without semantic support.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a precise verb+resource: preview and optionally apply the platform's remediation of a finding. It clearly distinguishes the read-only preview mode from the mutating apply mode, and the scope is unambiguous relative to siblings like update_finding or auto_remediate_compliance.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives explicit when-to-use guidance: without apply it is a preview, with apply=True it commits, and it says 'never apply without having shown the preview'. It also routes the agent away when the finding kind has no automatic remediation, telling it to resolve those with the control tools.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

remove_entityRemove EntityA

Soft-delete a single entity of any core type. Mutating: creates a new model version. Reversible with restore_entity using the same entity_type — the entity's ID is preserved (never reused) so a restore reinstates the same ID and all its links. To change an entity's fields instead of removing it, use the typed edit_* tool.

Dispatches on entity_type. Per-type consequence (all derived at read time; nothing is hard-destroyed):

  • asset — the asset's (asset × attacker) CO pairs are tombstoned, orphaning any controls mapped to them.

  • attacker — control objectives anchored to this attacker are tombstoned; controls left with no live anchor become orphaned.

  • component — controls scoped to this component have their component_id cleared (the controls themselves are kept) and the component's trust-boundary contribution to asset reachability is withdrawn.

  • trust_boundary — reachability widens: attacker vectors the boundary was filtering now pass freely and its sealed/isolation claim is dropped, so CO reachability verdicts past it can flip toward reachable/indeterminate.

  • assumption — marked deleted (kept for the audit trail); its CO links are cleared and its attestations retired; controls whose assumption_groups name it keep their groups, which credit nothing through it while it is deleted.

ParametersJSON Schema
NameRequiredDescriptionDefault
model_idYesID of the threat model.
entity_idYesID of the entity to soft-delete.
entity_typeYesWhich entity to soft-delete — one of ``asset``, ``attacker``, ``component``, ``trust_boundary``, ``assumption``.
server_versionYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations present, the description carries the full burden and delivers: it discloses soft-delete rather than hard destroy, that a new model version is created, that IDs are preserved and never reused, and enumerates per-entity-type consequences (tombstoning, orphaning, reachability widening, audit-trail retention) that are far beyond what any structured field provides.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the operation, reversibility, and alternative tools before the per-type bullet list. The length is justified by genuinely distinct per-type consequences, though the bullets are dense enough that some pruning is possible.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a multi-type soft-delete mutation with no annotations, the description covers reversibility, dispatch, and the differentiated side effects per type; an output schema exists so return values need not be explained. Nothing an agent needs to call this safely is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 75% and the description substantially enriches the key parameter: it explains that ``entity_type`` dispatches behavior and what each of the five values causes, and that ``entity_id`` is preserved across a restore. It adds little on ``model_id`` and ``server_version``, but those are already documented in the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource with scope: 'Soft-delete a single entity of any core type', followed by the mutating semantics and dispatch-on-entity_type. It is clearly distinguishable from sibling tools like edit_*, delete_control, delete_threat_model, and restore_entity.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly routes between alternatives: use ``restore_entity`` with the same ``entity_type`` to reverse, and use the typed ``edit_*`` tool to change fields rather than remove. The selecting conditions for each alternative are stated, not inferred.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

remove_model_from_groupRemove Model From GroupA

Remove a model from a tag. Mutating; the model itself is not deleted.

The frameworks the tag propagated to the model stay selected on it.

ParametersJSON Schema
NameRequiredDescriptionDefault
group_idYesthe tag.
model_idYesthe model to remove.
server_versionYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.9/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the behavior burden and delivers two non-obvious facts: the operation is mutating but non-destructive, and the frameworks propagated to the model remain selected. It still omits auth/permission needs and reversibility.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences, front-loaded with the action and its key side effect; nothing is wasted.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need no explanation, and the description covers the main mutation semantics. Given no annotations, it could add permission or reversibility details, but what is present is largely sufficient.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 67% and the description usefully equates 'tag' with the group id and confirms model_id is the item removed, adding vocabulary meaning beyond the schema's terse descriptions. server_version remains undocumented in both.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource (remove a model from a tag) and clarifies the mutation scope. It implicitly differentiates from add_model_to_group by naming the inverse operation, but never names a sibling explicitly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is implied by the name and the 'Remove a model from a tag' phrasing; the description gives no explicit when/when-not guidance or routing to alternatives like add_model_to_group or delete_group.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

resolve_verdict_divergencesResolve Verdict DivergencesA

Accept a batch of coverage divergences as mapping changes, or dismiss a batch of divergences. Mutating.

action="accept": each missing_mapping ADDS its CO to the control, each spurious_mapping REMOVES it, one version per affected control. Items are validated one by one: the answer separates applied from skipped (stale, would orphan, already so). To accept only confident rows, filter get_verdict_divergence's coverage rows by p_covers (near 1.0 for missing, near 0.0 for spurious) first. reason (min 10 chars) is recorded on each control's history.

action="dismiss": the structural model was right and the LLM was not; the model does not change. A dismissal is keyed to the row's current verdict input, so the row reappears once its control or objective changes. Works for coverage and group_sufficiency rows.

ParametersJSON Schema
NameRequiredDescriptionDefault
itemsYesJSON array of rows. accept: ``{"control_id", "co_id", "kind"}`` with kind ``missing_mapping`` or ``spurious_mapping``. dismiss: ``{"kind", "co_id", "control_id"?, "group_id"?}`` (control_id for coverage kinds, group_id for group_sufficiency).
actionYes"accept" or "dismiss".
reasonYesWhy (accept: min 10 chars; dismiss: non-empty).
model_idYesID of the threat model.
server_versionYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does so: it declares the tool is mutating, describes per-item validation producing applied vs skipped (stale, would orphan, already so), states that reason is recorded on each control's history, and explains that a dismissal is keyed to the row's current verdict input so it reappears once the control or objective changes. These are exactly the side-effect and persistence semantics an agent needs.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The core verb and mutating nature are front-loaded in the first sentence, and the action-specific blocks are clearly delimited. It is dense but every sentence carries operational information; only the formatting density slightly taxes readability.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 5-parameter mutating tool with an output schema (so return shape needn't be described), the description covers both actions, their validation outcomes, and the persistence behavior of dismissals. Nothing an agent needs to call it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 80% (baseline 3), but the description adds real meaning: it explains what each kind value does to the model ('missing_mapping ADDS its CO', 'spurious_mapping REMOVES it') and clarifies reason length thresholds per action. It doesn't document server_version, keeping it just short of a 5.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb pair and resource ('Accept a batch of coverage divergences as mapping changes, or dismiss a batch of divergences. Mutating.'), and explicitly names the sibling get_verdict_divergence as the source of the rows to act on. An agent can distinguish this mutating resolve tool from the read-side get_verdict_divergence without opening either schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives explicit when-to-use for each action: accept means 'the structural model was right and the LLM was not' changes; dismiss means the divergence should be suppressed. It even routes the agent to filter rows by p_covers in get_verdict_divergence before accepting, which is concrete pre-invocation guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

restore_entityRestore EntityA

Un-soft-delete a single entity of any core type, reversing a prior remove_entity. Mutating: creates a new model version. Only affects an entity that is currently soft-deleted.

Dispatches on entity_type. Per-type effect:

  • asset — revives the asset's tombstoned (asset × attacker) COs with their original IDs, un-orphaning any linked controls.

  • attacker — reinstates the attacker under its original ID, revives the COs tombstoned when it was removed, and un-orphans any controls that were anchored to it.

  • component — reinstates the component under its original ID, restoring its trust-boundary contribution to asset reachability.

  • trust_boundary — reinstates the boundary: the reachability it filtered re-narrows and its sealed/isolation claim is restored, so CO reachability verdicts past it can flip back toward unreachable.

  • assumption — returns the assumption to active status; controls whose assumption_groups referenced it keep their group structure intact. Its CO links are not restored (set them again with edit_assumption), and re-attestation is required before it counts anywhere it is linked or grouped.

Returns the entity-change result: {"model": <ThreatModel>, "controls_carried", "controls_orphaned", "orphaned_control_ids", ...}.

ParametersJSON Schema
NameRequiredDescriptionDefault
model_idYesID of the threat model.
entity_idYesID of the entity to restore.
entity_typeYesWhich entity to restore — one of ``asset``, ``attacker``, ``component``, ``trust_boundary``, ``assumption``.
server_versionYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does so thoroughly: it declares the mutation, the creation of a new model version, the soft-delete precondition, and per-entity-type side effects such as reviving tombstoned COs, un-orphaning controls, and requiring re-attestation for assumptions. It also notes what is not restored (assumption CO links) and how to fix it.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is long but front-loaded and well-structured, using a bulleted per-type breakdown that earns its place for a multi-type restore. The single sentence describing the return value is largely redundant given the output schema, a minor waste.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a complex mutating tool with no annotations and an existing output schema, the description covers the mutation semantics, preconditions, per-type effects, and limitations thoroughly. It omits only non-critical details like permission requirements or error behavior, leaving an agent fully equipped to call it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 75%, so the schema already documents model_id, entity_id, and entity_type labels. The description goes beyond the schema by detailing the behavioral consequence of each entity_type value, adding meaningful semantic depth for the dispatch parameter, though it adds nothing for the remaining parameters.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('Un-soft-delete'), resource ('a single entity of any core type'), and explicitly names the sibling operation it reverses (remove_entity). An agent can distinguish it from remove_entity and other mutation siblings without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides the key precondition — only affects an entity currently soft-deleted — and frames the tool as the inverse of remove_entity. It does not spell out explicit when-not-to-use cases or alternative routes beyond that inverse, so it is clear but not exhaustive.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

resume_control_generationResume Control GenerationA

Resume control generation that was paused, or retry one that stopped before finishing. Mutating.

Use when get_control_generation_status returns status: "paused" (someone paused it) or status: "blocked" (blocked.code dependency_unavailable or analysis_incomplete). A paused run resumes at once. For a blocked one the platform checks the services it depends on first, so a retry while one is still down costs nothing and changes nothing.

Returns one of:

  • {resumed: true, status: "queued", status_detail} — the run resumes where it stopped (only the unfinished work, billed to the original generation). Poll get_control_generation_status until complete.

  • {resumed: false, http_status: 409, code: "pause_in_progress"} — the run is still stopping after a pause; resume once it shows paused.

  • {resumed: false, http_status: 503, code: "dependency_unavailable", message, retry_after_seconds, ...} — still unavailable; relay the message and try again after retry_after_seconds.

  • {resumed: false, http_status: 409, code: "retry_too_soon", retry_after_seconds, ...} — a retry was just tried; wait.

  • {resumed: false, http_status: 409, code: "not_blocked", status} — nothing is paused; read status.

ParametersJSON Schema
NameRequiredDescriptionDefault
model_idYesID of the threat model whose paused control generation to resume.
server_versionYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden and does so: it declares the operation Mutating, explains that a resume continues where it stopped and is billed to the original generation, and documents the full set of outcome codes (pause_in_progress, dependency_unavailable, retry_too_soon, not_blocked) with associated HTTP statuses. This is far beyond what structured fields provide.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the one-line purpose and mutation flag before the usage condition and outcome list; the bulleted return shapes are scannable. It is somewhat long and duplicates return structure that the output schema likely already defines, but the error-code semantics earn most of the space.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a stateful mutating tool with no annotations and an output schema, the description supplies the missing context an agent needs: triggers, effects on billing/scope, and every terminal outcome code. The only real gap is that server_version is left entirely unexplained, which is minor for what appears to be a standard version field.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Only two parameters and schema coverage is 50%: model_id is documented in the schema, but server_version is undocumented in both schema and description. The description adds no parameter-level meaning beyond what the schema already states, so nothing compensates for the gap, and 3 is the baseline for a partially documented two-param tool.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Resume control generation that was paused, or retry one that stopped before finishing'), immediately flags it as Mutating, and implicitly distinguishes it from the sibling pause_control_generation by naming the resume/retry scope. An agent can tell what this does and which lifecycle state it targets without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives explicit preconditions tied to concrete values from a named sibling: use when get_control_generation_status returns status 'paused' or 'blocked' with specific blocked.code values. It also states when calling is harmless ('a retry while one is still down costs nothing and changes nothing'), which is real when/when-not guidance. Nothing is left to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

revalidate_entity_qualityRevalidate Entity QualityA

Re-run quality validation on a threat model's existing assets and attackers, as if they were freshly generated. A fast first-pass check judges every entity; only the ones it flags get a deeper review that confirms them, sharpens their wording, or flags them for you.

Use this to apply validation improvements to an already-generated model, or to clear stale quality warnings — without regenerating the whole model (which would destroy controls, assertions, and components). It is non-destructive: an entity that should be removed is left in place with a quality warning rather than deleted, so no control objective loses its asset or attacker anchor. It creates no new model version: the re-validation is queued and runs in the background, and the refreshed warnings appear on the next read of the model.

May consume credits for the entities that need the deeper review; a model already in good shape costs nothing. Returns at once with {"accepted": true, "queued": <entities queued>, "model": {...}}, where model is the model as it stands before the re-validation lands.

ParametersJSON Schema
NameRequiredDescriptionDefault
model_idYesID of the threat model whose assets and attackers to re-validate.
server_versionYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does so thoroughly: non-destructive (flagged entities are kept with a warning rather than deleted), no new model version, queued background execution, credit consumption conditional on entity count, and the exact synchronous return shape. Nothing about side effects is left to inference.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core action, then constraints, then cost, then return shape. Every sentence carries distinct information, though the description runs long and could trim slightly without loss.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Despite an output schema existing, the description still summarizes the return contract, and it covers the non-obvious traits (destructive-free behavior, async queue, credit cost) an agent needs before invoking. Complete for a background re-validation tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 50% (model_id documented, server_version undocumented). The description adds the semantic scope of model_id ('whose assets and attackers to re-validate') but says nothing about server_version. Baseline 3 is appropriate given the schema handles the main parameter.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('re-run quality validation on a threat model's existing assets and attackers') and describes the two-stage check (fast first pass, deeper review for flagged entities). An agent can distinguish it from regenerate_controls or refine_threat_model without opening schemas.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly names when to use it ('apply validation improvements to an already-generated model, or clear stale quality warnings') and the alternative it is not ('without regenerating the whole model, which would destroy controls, assertions, and components'). This is textbook when/when-not guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

select_compliance_frameworksSelect Compliance FrameworksA

Select (activate) compliance frameworks at a chosen scope. Requires PRO tier. Mutating.

Discover valid ids with list_compliance_frameworks (or add a custom one via import_compliance_framework); view the resulting gap analysis with get_compliance_report at the same scope. Re-calling replaces the scope's framework selection.

scope selects the target and how scope_id is read:

  • "model" — a single threat model (scope_id = model id). Activating a framework also kicks off background auto-remediation: it auto-maps existing controls to requirements, excludes non-applicable requirements by taxonomy, and suggests/applies new entities for the remaining gaps. The response includes auto_remediate_jobs, which run and complete on their own; re-trigger later with auto_remediate_compliance if the model changes.

  • "tag" — a tag (scope_id = tag id). Records the frameworks against the tag AND propagates them to every member model, and to every model added to the tag later, making the tag a compliance scope (e.g. an audit boundary) spanning several models.

ParametersJSON Schema
NameRequiredDescriptionDefault
scopeYestarget boundary — "model" or "tag".
scope_idYesid of the model or tag selected by ``scope``.
framework_idsYescomma-separated framework ids (e.g. "asvs-4.0,nist-csf").
server_versionYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden and delivers: it flags this as mutating, requires PRO tier, warns that re-calling replaces the scope's selection, and details that 'model' scope kicks off background auto-remediation jobs while 'tag' scope propagates to member and future models. This is rich behavioral context beyond what any schema field would convey.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Purpose, alternatives, and scope semantics are front-loaded and organized with a bulleted breakdown, so the structure aids scanning. It is somewhat long, but nearly every sentence adds operational information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a mutating, 4-param tool, the description covers prerequisites (PRO tier), scope semantics, side effects, and follow-up routes, and the existence of an output schema means return values needn't be detailed. Nothing an agent needs to call this correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 75%, and the description meaningfully extends it by explaining how scope_id is read under each scope value and the side effects of each ('model' vs 'tag'). It adds real semantic depth over the enum, though framework_ids format is left mostly to the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb+resource ('Select (activate) compliance frameworks') with explicit scope, and distinguishes itself from siblings list_compliance_frameworks, import_compliance_framework, auto_remediate_compliance, and get_compliance_report by name. An agent can tell this tool apart from its neighbors without opening any schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It names the alternatives and the conditions selecting them: use list_compliance_frameworks to discover ids, import_compliance_framework for custom ones, get_compliance_report to view results at the same scope, and auto_remediate_compliance to re-trigger. It also notes PRO tier requirement and that re-calling replaces the existing selection.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

set_control_assumption_groupsSet Control Assumption GroupsA

Declaratively set the assumption group structure for a control.

Replaces all assumption group assignments for this control. Each group is a set of assumption IDs that together externally handle the control; any one group being fully active+attested is sufficient.

  • Within a group: AND — all referenced assumptions must be active and attested for the group to count as complete

  • Across groups: OR — any one complete group marks the control as externally handled for mitigation purposes

To clear all assumption groups (revert to "not externally handled"), pass an empty JSON object: {}.

AI relevance gate (per group, no override): Each non-empty proposed group is evaluated independently. The behavior depends on how many groups pass:

  • All groups accepted → 200 success, structure persisted as submitted.

  • Some groups accepted (partial): the accepted groups ARE persisted (runtime OR-semantics activate immediately), the rejected groups are NOT saved, the call raises with HTTP 422 detailing both persisted_groups and rejected_groups (with per-group reasoning). Resubmit only the rejected groups with assumptions that cover the control, or sharpen those assumptions' descriptions.

  • All groups rejected: existing groups on this control are re-evaluated through the same gate. Relevant existing groups are preserved; irrelevant existing groups are dropped (assumptions themselves remain in the model — only this control's linkage is removed). The call raises with HTTP 422 detailing what was persisted, what was rejected, and what existing was dropped.

  • Empty submission ({}): clears all groups, no evaluation.

There is no force-override. To get a group accepted, choose assumptions whose descriptions actually cover the control or refine an assumption's description so coverage is explicit.

ParametersJSON Schema
NameRequiredDescriptionDefault
groupsYesJSON object mapping group numbers to assumption ID lists. Example: '{"1": ["AS1", "AS2"], "2": ["AS3"]}' Empty object `{}` clears all groups.
model_idYesID of the threat model.
control_idYesID of the control (e.g., "CTRL-03").
justificationNoWhy this group structure is appropriate (10 to 2000 characters when groups is non-empty; optional when clearing, at most 2000 characters if given).
server_versionYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure. It thoroughly explains the AI relevance gate, partial persistence, HTTP 422 behavior, preservation/dropping of existing groups, and the impossibility of force-override, which is far beyond a basic mutation description.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is long but the complexity demands detail. Information is front-loaded with the core replacement behavior, followed by group semantics and edge cases. There is some redundancy around 'no force-override,' but the overall structure remains focused and useful.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a complex mutation tool with no annotations, the description covers the main operation, success and failure modes, side effects on existing data, clearing behavior, and remediation guidance. Since an output schema exists, the lack of return-value detail is acceptable.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 80%, so the baseline is 3. The description adds important semantic meaning beyond the schema, especially the AND-within-group / OR-across-groups interpretation and the consequence of submitting rejected groups. This enriches the groups parameter meaning beyond the schema's simple JSON mapping.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource: 'Declaratively set the assumption group structure for a control.' It clearly states the full replacement behavior, which distinguishes it from get-style or mitigation/foundation sibling tools, even though no sibling is named explicitly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives clear operational guidance, including how to clear groups with an empty object, how partial acceptance behaves, and how to recover from rejected groups. It does not explicitly compare against alternatives like set_mitigation_groups, but the usage context is otherwise well defined.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

set_control_objective_calSet Control Objective CalA

Set the per-CO ISO/SAE 21434 Cybersecurity Assurance Level (CAL).

CAL is a 1-4 grade on each individual control objective that expresses how much assurance the control program owes for that specific objective. It lives on the control_objectives identity side-table — writes do NOT create a new threat-model version, and the value survives soft-delete + revival of the CO.

Pass cal=None (or omit it) to clear the value.

ParametersJSON Schema
NameRequiredDescriptionDefault
calNoISO/SAE 21434 CAL grade (1-4), or ``None`` to clear.
co_idYesControl-objective ID (e.g. ``CO3``).
model_idYesID of the threat model.
server_versionYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.4/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral disclosure burden. It does this exceptionally well by revealing that writes do NOT create a new threat-model version, that the value persists across soft-delete and revival, and that omitting or passing cal=None clears the value. These are non-obvious side effects an agent must know before invoking the tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is tightly structured and front-loaded with the core purpose, then adds only high-value behavioral and clearing details. The use of line breaks and code formatting for cal=None makes the critical instruction easy to parse. No sentence is wasted.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description covers the most important operational semantics: what CAL is, where it lives, how versioning behaves, and how to clear it. Since an output schema exists, return values need not be described. The main gap is server_version, whose purpose and allowed values are left entirely undefined, and there is no guidance on what happens if an out-of-range CAL is supplied.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 75%, and the description adds meaningful semantics beyond the schema for the cal parameter: it defines the 1–4 assurance grading, ties it to per-CO obligations, and clarifies the clearing behavior. co_id also gains an example ('CO3') through the schema. However, server_version remains entirely undocumented in both schema and description, preventing a higher score.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource: 'Set the per-CO ISO/SAE 21434 Cybersecurity Assurance Level (CAL).' It clarifies the resource is each individual control objective and adds a distinguishing behavioral fact — writes do not create a new threat-model version and live on the control_objectives identity side-table. This differentiates it from sibling control-update tools without requiring schema inspection.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies when to use the tool — when you need to set or clear a CAL value on a control objective — and explicitly explains the clearing behavior with cal=None. However, it does not name alternative tools or state conditions for choosing this over related siblings like update_control_status or refine_control. Usage context is present but exclusion/alternative guidance is absent.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

set_functional_satisfaction_groupsSet Functional Satisfaction GroupsA

Declaratively set (replace) a functional objective's satisfaction groups. Mutating.

Replaces the objective's group structure wholesale. Each group is a set of functional tests that together satisfy the objective (AND within a group); the objective counts as satisfied when any one complete group has all its tests verified (OR across groups). Tests you want to keep associated with the objective but outside any group go in ungrouped. Unlike set_control_assumption_groups, there is no AI relevance gate — the structure you submit is applied as-is. Read the current state first with get_functional_satisfaction_groups.

ParametersJSON Schema
NameRequiredDescriptionDefault
model_idYesID of the threat model.
ungroupedNoComma-separated functional-test ids to keep associated with the objective but unassigned to any group (optional).
groups_jsonYesJSON object mapping group label to a list of functional test ids, e.g. ``{"1": ["FT-1", "FT-2"], "2": ["FT-3"]}``. Pass ``{}`` to clear all groups.
server_versionYes
functional_objective_idYesThe objective whose groups to set.

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.9/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and delivers: 'Mutating' and 'Replaces the objective's group structure wholesale' clearly disclose the destructive behavior. It also explains the AND-within-group/OR-across-groups semantics and that the submitted structure is applied as-is, giving an agent a complete behavioral profile.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Four sentences with zero redundancy: mutation flag, wholesale replacement, group semantics, ungrouped explanation, sibling contrast, and a read-first tip. The verb and resource are front-loaded, and each sentence earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool's behavioral complexity is fully covered: semantics, destructive replace behavior, ungrouped handling, and the distinction from set_control_assumption_groups. An output schema exists, so return-value details are unnecessary. Nothing critical is missing for an agent to invoke this tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 80%, so the baseline is 3. The description adds value beyond the schema by explaining the purpose of the ungrouped parameter ('keep associated with the objective but outside any group') and clarifying the group semantics: a group is a set of tests that together satisfy the objective, and any complete group suffices. This goes beyond the schema's JSON shape and examples.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb-resource pair, 'Declaratively set (replace) a functional objective's satisfaction groups,' and signals 'Mutating.' It also explicitly distinguishes itself from set_control_assumption_groups by noting the absence of an AI relevance gate, which differentiates it from its closest sibling.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives explicit guidance to read the current state first with get_functional_satisfaction_groups, and contrasts this tool with set_control_assumption_groups ('no AI relevance gate — the structure you submit is applied as-is'). This establishes when to choose this tool over the sibling without ambiguity.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

set_mitigation_groupsSet Mitigation GroupsA

Declaratively set the mitigation-group structure for a control objective. Mutating; runs as a polled background job (an LLM sufficiency check evaluates whether the new structure satisfies the CO) and returns once complete.

Replaces ALL mitigation-group assignments for this CO. Call get_mitigation_groups first to see the current structure and the unmapped controls available for assignment.

Mitigation groups define alternative paths to satisfy a CO:

  • Within a group: AND — all controls must be implemented.

  • Across groups: OR — any one complete group mitigates the CO.

  • Defense-in-depth: tracked but not required for mitigation.

ParametersJSON Schema
NameRequiredDescriptionDefault
co_idYesID of the control objective (e.g., "CO5").
groupsYesJSON object mapping group numbers to control-ID lists. Example: '{"1": ["CTRL-01", "CTRL-02"], "2": ["CTRL-03"]}'.
model_idYesID of the threat model.
justificationNoWhy this group structure is appropriate (10 to 2000 characters). The gate weighs it alongside the objective and the controls' own descriptions, so state the reasoning, not evidence.
server_versionYes
defense_in_depthNoComma-separated control IDs tracked as defense-in-depth (not required for mitigation). Example: "CTRL-04,CTRL-05".

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden, and it delivers: it discloses that the operation is mutating, that it runs as a polled background job with an LLM sufficiency check, that it replaces ALL assignments, and that defense-in-depth controls are tracked but not required. This is substantial behavioral disclosure for a destructive-ish mutation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is front-loaded with the action and mutation risk, then adds the critical call-get-first guidance, then compactly explains the group logic in bullets. Every sentence and bullet earns its place; nothing is filler or redundant with the schema.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a mutation tool with no annotations, the description covers the essential context: replacement behavior, polling/completion semantics, group logic, and parameter intent. An output schema exists, so return-value explanation is not needed. The only noticeable gap is the completely undocumented server_version parameter, which keeps this from being fully complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is high (83%), so the baseline is 3, but the description adds real meaning beyond the schema: it explains the 'justification' parameter's role in the gate ('state the reasoning, not evidence') and clarifies the defense_in_depth semantics and groups JSON structure. The undocumented server_version parameter is not addressed at all, preventing a 5.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a precise verb and resource: 'Declaratively set the mitigation-group structure for a control objective.' It clearly distinguishes this from the read-side sibling by explicitly telling the agent to call get_mitigation_groups first to inspect current structure. The replacement semantics and group semantics make the tool's role unmistakable.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives concrete usage context: call get_mitigation_groups first, note that all existing assignments are replaced, and understand the AND/OR/defense-in-depth semantics. It stops short of explicitly naming alternatives or saying when not to use this tool versus other group-setting siblings like set_control_assumption_groups, so it misses the highest bar for exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

split_composition_entitySplit Composition EntityA

Push an ancestor-owned entity down to one or more descendants and soft-delete the ancestor's copy. Mutates state across the ancestor + every target descendant.

Inverse of lift_composition_entity. Use when an entity that currently lives on an ancestor is in fact descendant-specific and should be modeled separately per descendant — the operator chooses which descendants take a copy. A new local id is minted on each target; attached state on the ancestor's entity (assertions, jira mappings, risk acceptances, etc.) is duplicated to every target.

The route's model_id IS the ancestor (the entity being split lives on it). Each affected model (ancestor + every target descendant) bumps version and emits a model_refined activity event; a structured split_applied event with the full split_event payload lands on the ancestor. The audit pack surfaces this under split_history.

ParametersJSON Schema
NameRequiredDescriptionDefault
kindYesEntity kind — one of ``"assets"``, ``"attackers"``, ``"components"``.
model_idYesAncestor model id — the entity to split lives here.
server_versionYes
ancestor_local_idYesLocal id of the entity on the ancestor.
target_descendantsYesNon-empty list of descendant model ids that should each take a copy.

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.9/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden and does so thoroughly. It discloses mutation, soft-deletion, state duplication across targets, new local id minting, version bumps, activity events, structured split_applied events, and audit history visibility.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is front-loaded with the core action and side-effect summary, then expands into usage context and event behavior. Every sentence adds operational value, and the length is justified by the complexity of the mutation.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description covers purpose, side effects, parameter routing, event emissions, and audit behavior. Since an output schema exists, return-value documentation is not the description's job, and nothing essential is missing for an agent to invoke this correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 80%, so baseline is 3, but the description adds meaningful semantics: model_id is the ancestor, the entity being split lives on it, target_descendants are the models that each receive a copy, and ancestor_local_id identifies the entity to split. This goes beyond the raw parameter names.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource: it pushes an ancestor-owned entity down to descendants and soft-deletes the ancestor's copy. It also explicitly names itself as the inverse of lift_composition_entity, making its unique role clear among siblings.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly frames when to use it: when an entity on an ancestor is actually descendant-specific and should be modeled separately per descendant. It also names the inverse operation, giving the agent a clear decision boundary against the closest alternative.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

start_control_buildStart Control BuildA

Start the model's proposed control build. Mutating only with confirm_estimate=True; consumes credits then.

A model's controls are built only by a build someone starts. Generating or refining the model, editing an entity and regenerate_controls each PROPOSE one; get_control_generation_status shows it as proposal.

  1. Call with confirm_estimate=False (the default). Nothing starts and nothing is charged; the answer is {started: false, proposal, message}, the proposal carrying a fresh estimated_credits and the model_version and set_revision the model stands at. Show the user what it would build and cost.

  2. Call again with confirm_estimate=True and those two values once they have reviewed the model. The build starts ({started: true, job_id, model_version, status: "queued", status_detail}) and holds the model until it publishes, fails or is discarded: other writers of the model's controls are refused meanwhile, and reads show the last published controls. Poll get_control_generation_status until terminal; the publish is one step, after which get_controls shows the result.

Refusals come back as data:

  • {started: false, http_status: 404, code: "no_proposal"} — nothing is proposed for the model.

  • {started: false, http_status: 409, code: "review_stale", proposal, model_version, set_revision, estimated_credits} — the model or its controls changed since the values were read (or none were named). Review again and start with the values returned.

  • {started: false, http_status: 409, code: "generation_active", status} — a build already holds the model.

  • {started: false, http_status: 402, ...} — the balance this workspace bills to cannot cover the estimate.

ParametersJSON Schema
NameRequiredDescriptionDefault
model_idYesID of the threat model.
set_revisionNoWith ``confirm_estimate=True``: the proposal's ``set_revision``.
model_versionNoWith ``confirm_estimate=True``: the proposal's ``model_version``.
server_versionYes
confirm_estimateNoFalse (default) returns the proposal and its estimate and starts nothing; True starts the build.

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden and does so: credits are consumed only when confirm_estimate=True, the model is locked until publish/fail/discard, other writers are refused meanwhile, and reads show the last published controls. It also enumerates refusal codes (no_proposal, review_stale, generation_active, 402 balance) as structured returns, which an agent could not infer from the schema.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The critical decision (confirm_estimate gating cost and mutation) is front-loaded in the first sentence, and the numbered two-step flow is easy to scan. The refusal-code list is long but each entry adds actionable recovery information, so little is wasted, though the description is heavier than strictly necessary.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a two-phase, credit-consuming mutation that locks a model, the description supplies the estimate/confirm loop, the locking semantics, alternate proposers, and the polling path to terminal state. With an output schema present it need not restate return shapes, yet it still documents them concretely, leaving no gap an agent must guess at.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 80% and the schema already documents each parameter, so the baseline is 3; the description goes further by tying model_version and set_revision to the values returned in the proposal and explaining that confirm_estimate=False starts nothing and charges nothing. It does not explain server_version or model_id beyond the schema, keeping this short of a 5.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Opens with a precise verb+resource ('Start the model's proposed control build') and immediately scopes it against siblings by explaining that generate/refine/edit/regenerate_controls only PROPOSE and that this tool is what actually starts a build. An agent can distinguish it from regenerate_controls or resume_control_generation without opening any schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives an explicit ordered workflow: first call with confirm_estimate=False to get the estimate, then call with confirm_estimate=True plus model_version and set_revision after user review. It also names the alternatives that merely propose and the poll target (get_control_generation_status), so when-to-use and when-not-to-use are both covered.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

strengthen_controlsStrengthen ControlsA

Strengthen a model's controls: work on the objectives whose mitigation groups the background judge found do not cover them (the uncovered and undecided of get_control_generation_status's diagnosis). Mutating only with confirm_estimate=True; consumes credits then. not_judged objectives have nothing to strengthen from: judge them first with judge_objectives.

  1. Call with confirm_estimate=False (the default): nothing starts or is charged; the answer carries diagnosis, scope, estimate (credits, per_objective, basis) and the model_version / set_revision the model stands at. Show the user the estimate.

  2. Once they agree, call with confirm_estimate=True and those two values (their review of the model as it stood). A background run starts (started: true); poll get_control_generation_status. It holds the model like any build: pausable, resumable, discardable.

A gap only the environment can close (hosting, a third party) is never answered with a control: an accepted assumption stating it is bound into the group; otherwise an assumption proposal waits in the review queue (get_review_queue / decide_proposal) and the objective counts as awaiting_assumption. A rejected one is not proposed again.

Refusals come back as data, {started: false, http_status, code}: 409 review_stale (the model or its controls changed since the estimate: confirm again with the values returned), 409 generation_active (a build holds the model), 402 (the balance cannot cover the estimate).

ParametersJSON Schema
NameRequiredDescriptionDefault
co_idsNoOptional comma-separated objective IDs, of those diagnosed uncovered or undecided; omit for all.
model_idYesID of the threat model.
set_revisionNoWith ``confirm_estimate=True``: the estimate's value.
model_versionNoWith ``confirm_estimate=True``: the estimate's value.
server_versionYes
confirm_estimateNoFalse (default) estimates; True starts the run.

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.9/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does so thoroughly: it discloses that credits are only consumed with confirm_estimate=True, that the answer carries diagnosis/scope/estimate/model_version/set_revision, that runs are pausable/resumable/discardable, and it enumerates refusal outcomes as data with specific status codes (409 review_stale, 409 generation_active, 402).

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Long but justified by tool complexity, and strongly front-loaded: the purpose is in the first clause before the numbered workflow. The two-step procedure is enumerated, and the refusal section is compact and scannable. No filler sentences.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a multi-step, credit-consuming mutation with rich refusal semantics, the description covers the estimate/confirm cycle, preconditions, side effects, and error handling. Output schema exists, yet the description still usefully summarizes answer fields, leaving nothing an agent needs missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 83%, so the schema already documents most parameters. The description still adds meaning: co_ids is constrained to objectives diagnosed uncovered/undecided, and model_version/set_revision are tied to the estimate values that must be echoed back on confirmation. It does not add syntax beyond that, so not a 5.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('strengthen a model's controls') and precisely scopes the work to objectives whose mitigation groups the judge found uncovered/undecided. It distinguishes itself from sibling tools by naming get_control_generation_status, judge_objectives, and get_review_queue in their roles.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicit two-step workflow: call with confirm_estimate=False to estimate, then True to start. It names the preconditions (not_judged objectives need judge_objectives first) and the polling/alternative path (get_control_generation_status, review queue). When and when-not is fully covered.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

submit_assertionsSubmit AssertionsD

Typed claims about a control, assumption or functional test; CI checks later. get_assertion_types returns it all as data.

By class, strongest first, as name(required) [opt: optional]: [by_construction]

  • typed_boundary(scope, sinks, boundary_type, constructors, property) [opt: allowlist, wrappers] [sound_over_approximation]

  • sink_default_deny(scope, sinks, safe_forms, property) [opt: allowlist, wrappers] [existential_witness]

  • test_attested(test) [opt: env, mechanism] [under_approximating_scan]

  • pattern_matches(file, pattern) [opt: scope_start, scope_end, multiline, dotall, target]

  • pattern_absent(file, pattern) [opt: scope_start, scope_end, multiline, dotall, target]

  • no_plaintext_secret(file, patterns) [presence]

  • function_exists(file, name)

  • class_exists(file, name)

  • decorator_present(file, function, decorator)

  • function_calls(file, caller, callee)

  • import_present(file, module)

  • file_exists(file)

  • file_hash(file, algorithm, expected_hash, scope_file) [opt: scope_start, scope_end]

  • config_key_exists(file, key)

  • config_value_matches(file, key, pattern)

  • env_var_referenced(file, variable)

  • dependency_exists(manifest, package)

  • dependency_version(manifest, package, constraint)

  • parameter_validated(file, function, parameter)

  • error_handled(file, function)

  • middleware_registered(file, middleware)

  • http_header_set(file, header)

  • test_exists(pattern)

  • module_exists(file, name)

  • module_instantiated(file, parent, child)

  • port_exists(file, module, port) [opt: direction]

  • parameter_defined(file, parameter) [opt: module, pattern]

  • signal_exists(file, name) [opt: module, kind]

  • sva_assertion_present(file, name)

  • register_reset(file, signal) [opt: reset]

Each: type, params, description, repo ("/" or "no_repo"), covers beside them, never in params: the CO-NN or cls_ ids proved. A for-all clause takes only typed_boundary (sinks accept one boundary type) or, when they do not, sink_default_deny, bound with covers.

ParametersJSON Schema
NameRequiredDescriptionDefault
model_idYes
control_idNo
assumption_idNo
server_versionYes
assertions_jsonYes
functional_test_idNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

D1.5/5.0
Behavior1/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. It does not state that submission is a mutation, whether it requires a control/assumption/functional_test anchor, whether it is idempotent, what happens on invalid assertion_json, or what the output schema returns. The one behavioral note ('CI checks later') is vague and non-actionable.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The body is a long undifferentiated list of 30+ assertion types occupying most of the text, with the tool's actual purpose reduced to a single fragment. Structure is flat, poorly front-loaded, and buries the actionable information about required parameters and anchoring.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 6-param mutation tool with an output schema and no annotations, the description should explain anchoring (control_id/assumption_id/functional_test_id), the assertions_json payload format, and server_version. None of this is covered; the type taxonomy, however useful, does not substitute for tool-contract information.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% across 6 parameters, so the description must compensate, and it does not. It documents assertion-internal fields (scope, sinks, property, etc.) that are not input parameters, while model_id, server_version, assertions_json, control_id, assumption_id, and functional_test_id are never explained. The 'covers' rule is the only hint about anchoring, and it is cryptic.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a fragment ('Typed claims about a control, assumption or functional test; CI checks later') that gestures at purpose but never states a clear verb+resource like 'submit typed assertions for a model'. It spends most of its length on a taxonomy of assertion types rather than what the tool does. An agent must infer that this tool ingests/submits assertions.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is a single routing hint to get_assertion_types ('returns it all as data'), which is a sibling-lookup pointer rather than when-to-use-this-tool guidance. No statement of when to submit vs. list_assertions, delete_assertion, or submit_findings, and no prerequisites or exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

submit_attestationSubmit AttestationA

Record that a responsible party affirmed an assumption holds.

Only for external assumptions. Non-applicability assumptions require CI verification (submit assertions + run mipiti-verify) — manual attestation is rejected for them.

An assumption with a current attestation can mitigate linked COs. When the attestation expires, those COs become at-risk until re-attested or covered by controls.

Attesting accepts the assumption, which is a judgment about the world the platform cannot check: a program may do it only under a workspace delegation rule for assumption_accepted; otherwise the call is refused with HTTP 403 and an escalation_id and the attestation is parked for a person. An attestation holds for the text it was given for: editing the assumption's description retires it, and the assumption must be accepted again. An assumption need not be linked to an objective to be accepted — one bound into a control's group (see strengthen_controls) counts only while it is accepted.

An attestation is a responsible party's claim, never a proof over every site: it can cover an existential clause of a control (its tier reads claimed) and never a for-all one, where only a sound witness counts. An attestation the platform mints from CI results is no stronger than the weakest assertion behind it. The exits for a universal objective that cannot be proven are a risk acceptance or a not-applicable disposition.

ParametersJSON Schema
NameRequiredDescriptionDefault
model_idYesID of the threat model.
statementNoWhat was attested.
expires_atNoISO 8601 expiry date (e.g., "2026-06-30T00:00:00Z").
attested_byNoWho is attesting (name, role, organization).
evidence_urlNoOptional link to supporting documentation.
assumption_idYesID of the assumption (e.g., "AS1").
server_versionYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does so: it discloses the 403 refusal with escalation_id and that the attestation is parked for a person, that editing the assumption's description retires the attestation, and that expiry flips linked COs to at-risk. It further explains the semantics of coverage (existential 'claimed' tier vs for-all) so the agent knows what the write actually buys.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core action and the exclusion, and every paragraph carries behavioral information. It is however long and drifts into conceptual prose ('a judgment about the world the platform cannot check', 'no stronger than the weakest assertion behind it') that could be tightened without losing actionable content.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a mutation tool totalling a 7-parameter schema with an output schema present, the description covers preconditions, failure mode, side effects, lifetime, and interaction with linked COs. Nothing an agent needs to invoke it correctly is missing, and return values are covered by the output schema.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 86%, so the schema already documents model_id, assumption_id, expires_at, attested_by and evidence_url. The description adds contextual meaning around the attested statement ('holds for the text it was given for') but no new per-parameter syntax or format guidance, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('Record that a responsible party affirmed an assumption holds') and immediately scopes it apart from the sibling workflow: external assumptions only, not non-applicability ones. An agent can distinguish this from submit_assertions and submit_findings without opening any schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives explicit when-to-use ('Only for external assumptions'), when-not ('non-applicability assumptions... manual attestation is rejected'), and names the alternative path ('submit assertions + run mipiti-verify'). It also states the delegation-rule precondition that governs whether the call is allowed.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

submit_findingsSubmit FindingsA

Record negative findings (gaps discovered while scanning a codebase against a model's controls). Mutating: persists new finding records against the model.

Use after a gap-discovery scan (see get_scan_prompt) to log where expected control evidence was NOT found. Findings are the negative counterpart to assertions (positive proof via submit_assertions): a finding says "I looked here for this and it was missing." Once submitted, drive a finding through its lifecycle with update_finding and review them with list_findings.

ParametersJSON Schema
NameRequiredDescriptionDefault
model_idYesID of the threat model.
findings_jsonYesJSON string of an **array** of finding objects. Each object should carry: - ``control_id`` (str): the control the gap relates to. - ``title`` (str): short summary of the gap. - ``description`` (str): what is missing and why it matters. - ``severity`` (str): finding severity (e.g., "low"/"medium"/"high"/"critical"). - ``checked_locations`` (list): files/paths inspected. - ``checked_patterns`` (list): patterns/signals searched for. - ``expected_evidence`` (str): what implemented evidence would have looked like. Must parse as a JSON array; a single object or malformed JSON is rejected.
server_versionYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the burden of disclosing side effects. It clearly says 'Mutating: persists new finding records against the model' and notes the lifecycle with update_finding/list_findings. However, it does not mention idempotency, duplicate behavior, whether submissions can be overwritten, or any permission requirements—common gaps for a mutating tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences with no wasted words. The first sentence front-loads the core purpose and mutating nature; the second gives concrete usage context; the third adds the assertion counterpart and subsequent lifecycle tools. Every sentence earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with a detailed input schema and an existing output schema, the description is largely complete: it explains why and when to use it, what findings mean, and how to manage them afterward. The main residual gap is the under-documented server_version parameter and the lack of idempotency/duplicate caveats, but these are minor given the overall context.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema already documents model_id and findings_json in detail, including the JSON array format and rejection of malformed input. The description adds useful intent behind model_id and findings_json, but it does not clarify server_version, one of three required parameters that has no schema description.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('Record') and resource ('negative findings... gaps discovered while scanning a codebase against a model's controls'), and explicitly distinguishes findings from assertions via submit_assertions. An agent can clearly tell this tool apart from related siblings without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly says when to use it: 'Use after a gap-discovery scan (see get_scan_prompt)' and clarifies it is for logging where expected evidence was NOT found. It also contrasts with submit_assertions and points to update_finding/list_findings for lifecycle and review, giving clear alternatives and next steps.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

suggest_functional_test_mappingsSuggest Functional Test MappingsA

Suggest which functional objectives each imported test likely covers.

For unmapped tests (imported without an association, or added without objective ids), this proposes objective mappings so you can review and apply them with associate_functional_test. It only suggests — nothing is associated until you confirm.

ParametersJSON Schema
NameRequiredDescriptionDefault
model_idYesID of the threat model.
test_idsNoComma-separated functional-test ids to map. Empty means every currently-unmapped test.
server_versionYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full behavioral burden. It explicitly discloses the key side-effect boundary: 'It only suggests — nothing is associated until you confirm.' This directly tells the agent the tool is non-destructive and does not alter associations, which is critical for a proposal-style tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is compact and front-loaded. The first sentence states the core purpose, and the second provides exactly the necessary usage and safety context. No redundant phrasing or filler exists; every sentence earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description fully covers what the tool does, when to use it, what makes a test eligible, and the fact that no association happens until confirmation. An output schema exists, so the return shape is covered elsewhere. Nothing essential is missing for an agent to select and invoke this tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 67%, so model_id and test_ids already have meaning in the schema. The description reinforces the test_ids semantics ('unmapped tests', 'every currently-unmapped test') but adds little beyond that. The required server_version parameter remains undocumented in both the schema and the description, though coverage is high enough to keep the score at baseline.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource: 'Suggest which functional objectives each imported test likely covers.' It clearly scopes the tool to unmapped tests and explicitly distinguishes it from the sibling apply tool, associate_functional_test, so an agent can tell them apart.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description states exactly when to use this tool: for unmapped tests (imported without an association or added without objective ids). It also names the follow-up tool, associate_functional_test, and clarifies that this tool only proposes mappings—so the agent knows this is not the action for applying them.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

undo_composition_eventUndo Composition EventA

Preview, and on confirmation apply, the undo of a lift or split.

event_type is lift or split; event_id is the forward lift_applied / split_applied activity event's id, or the lift_id / split_id in its payload. model_id is the model the event was raised on (another model's event is 404).

By default (dry_run=True) it is read-only: {plan, refusal}, exactly one non-null. plan lists the inverse operations an undo would commit (lift: tombstone the LCA entity, restore the source copies, rewrite CO references; split: restore at the ancestor, tombstone the target copies). refusal lists why it cannot: state has moved since the event (assertions submitted on the entity, objectives referencing it, an edit). Show it to the operator.

With dry_run=False it is mutating, after explicit confirmation: the divergence check runs again (409 with detail.refusal.reasons when it refuses), the inverse is persisted across every affected model, and a lift_undone / split_undone event citing original_event_id is recorded. Returns {undone_event_id, original_event_id, applied_state_ops, models}; models is {lca_model, source_descendant_models} for a lift and {ancestor_model, descendant_models} for a split.

503 where composition is not available.

ParametersJSON Schema
NameRequiredDescriptionDefault
dry_runNo
event_idYes
model_idYes
event_typeYes
server_versionYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden and does so thoroughly: default read-only vs mutating mode, the exact {plan, refusal} return shape with one non-null, the 409 refusal payload path, the 503 unavailability, cross-model persistence, and the recorded undo event. This is unusually complete disclosure for a mutation tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core action, then structured into default dry_run behavior, confirmed mutation behavior, and error codes. It is dense but every sentence adds actionable detail; a short leading verb phrase for the mutating path would tighten it slightly.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a complex, multi-model destructive operation with 5 parameters at 0% schema coverage and no annotations, the description covers prerequisites, default safety, refusal conditions, error codes, and side effects. An agent can invoke it correctly on the first attempt.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, so the description must compensate, and it does: event_type (lift/split), event_id (forward lift_applied/split_applied id or the lift_id/split_id in the payload), model_id (404 for another model's event), and dry_run semantics are all explained. Only server_version, a required parameter, goes unexplained.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('undo of a lift or split') and clearly scopes it as a preview-then-apply operation. An agent can distinguish it from siblings like lift_composition_entity, split_composition_entity, and undo_model_change without opening any schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives strong conditional guidance: dry_run defaults to True (read-only preview), and dry_run=False mutates only 'after explicit confirmation', with the divergence check re-running. It does not, however, name the sibling alternative (e.g., why use this versus undo_model_change), so it stops short of explicit alternative routing.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

undo_model_changeUndo Model ChangeA

Undo the latest change to a model's controls, or revert its latest version. Mutating.

target="controls" restores exactly what the latest set revision of the latest version replaced (list_control_revisions' undo_target). Changes are undone latest first, one per call; the undo is itself recorded, and the next undo goes to the change before it. There is no redo. Verdicts the undo returns to are served again rather than re-judged. Answers {applied: true, model_id, model_version, undone, revision, controls}: undone is the revision undone, revision the one the undo recorded, controls the ids it restored.

target="version" creates a new version copying the latest earlier version not already discarded — the model, its controls and their objective metadata — and marks the replaced version discarded. Version numbers are never reused, and the discarded version stays readable in the history. Findings on controls the revert removes are resolved. Answers {applied: true, model_id, model_version, copied_from, discarded}: model_version is the new version.

A refusal comes back as {applied: false, http_status: 409, code, message}: generation_active (a build holds the model), and for controls nothing_to_undo or set_diverged (a control the change touched has changed since; undo the later change first), for a version no_earlier_version.

ParametersJSON Schema
NameRequiredDescriptionDefault
targetYes
model_idYes
server_versionYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does so well: it declares the tool mutating, states that the undo is itself recorded and irreversible (no redo), explains that verdicts are re-served rather than re-judged, that findings on removed controls are resolved, that version numbers are never reused, and that the discarded version stays readable. It even enumerates the 409 refusal codes with their meanings.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Well front-loaded: the core action and mutation warning come first, then per-target behavior and refusal handling. It is on the long side, and the two explicit response-shape listings duplicate what the existing output schema already conveys, which is slightly wasteful.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a destructive, two-mode mutation tool with an output schema, the description covers both branches, their side effects, failure modes, and recovery guidance. An agent has everything needed to invoke it correctly and interpret both success and refusal responses.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It fully specifies the meaning and side effects of the target enum values ('controls' vs 'version'), which is the only non-obvious parameter. model_id and server_version are left undefined, but they are self-explanatory identifiers, so the gap is minor.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a precise verb+resource ('Undo the latest change to a model's controls, or revert its latest version') and immediately splits the two modes by the target parameter. It also references the sibling tool that supplies the undo_target (list_control_revisions), so an agent can tell this apart from adjacent revision/history tools without opening a schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives clear operational context: one change undone per call, latest first, no redo, and the precondition that set_diverged must be resolved by undoing the later change first. It lacks explicit guidance on when to prefer this over neighbors like undo_composition_event or discard_control_build, so it falls short of a full when/when-not/alternatives treatment.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

update_control_statusUpdate Control StatusA

Update the implementation status of a security control. Mutating.

Sets the control's status to "implemented" or "not_implemented". Marking a control "implemented" REQUIRES at least one assertion on the control — check its assertion_count (via get_controls) first and submit assertions with submit_assertions if it is zero, or the call is rejected.

ParametersJSON Schema
NameRequiredDescriptionDefault
statusYesNew status — "implemented" or "not_implemented".
model_idYesID of the threat model the control belongs to.
control_idYesID of the control to update (e.g. "CTRL-01").
server_versionYes
implementation_notesNoOptional free-text notes recorded with the status change.

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden of behavioral disclosure. It openly states that the operation is mutating and documents a non-obvious rejection condition: marking a control implemented requires at least one assertion. It could add more about side effects or permissions, but the critical behavioral trait is covered.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three focused sentences with no filler. The core action and mutating nature are front-loaded, followed by a clearly formatted precondition block. Every sentence adds value, and the REQUIRES is emphasized without unnecessary prose.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With an output schema present, the description doesn't need to explain return values. Combined with the input schema, it gives an agent enough to call the tool correctly: required parameter names, permissible status values, and the critical assertion prerequisite. Minor gaps like error details or authentication are not essential for invoking the tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 80%, so the schema already documents most parameters clearly. The description adds context around the status transition and the assertion requirement, but doesn't elaborate on parameter syntax beyond what the schema provides. server_version still has no meaningful description, but the high schema coverage keeps this at baseline.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Description uses a specific verb and resource: it 'Sets the control's status to implemented or not_implemented'. It clearly distinguishes this from sibling tools like refine_control or remap_control by focusing solely on implementation status. The 'Mutating' label reinforces the action without ambiguity.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives explicit context for when to use this tool: when updating a control's implementation status. It also names prerequisite tools and conditions — check assertion_count via get_controls and use submit_assertions before marking implemented, or the call is rejected. It doesn't explicitly state when not to use this tool versus an alternative, but the guidance is clear enough.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

update_findingUpdate FindingA

Advance a finding through its lifecycle. Mutating: updates the finding's status and metadata.

Use to acknowledge, remediate, verify, or dismiss a finding previously recorded by submit_findings / list_findings. This records a MANUAL status transition — the machine-set auto_resolved state is not among the statuses it accepts; for gaps whose kind has an automatic fix, remediate_finding performs the actual cleanup instead.

ParametersJSON Schema
NameRequiredDescriptionDefault
notesNoOptional free-text notes recorded on the finding.
reasonNoOptional rationale; required when dismissing (status="dismissed").
statusYesNew lifecycle status, one of "discovered", "acknowledged", "remediated", "verified", "dismissed". ``auto_resolved`` is NOT settable here: it asserts that a condition is no longer reproduced, which is a claim only the platform can make from its own re-evaluation. Setting it by hand would forge that claim, so this tool refuses it — use ``dismissed`` (with a reason) to record that a gap does not matter, which is the judgment a person is entitled to make.
model_idYesID of the threat model.
finding_idYesID of the finding to update.
server_versionYes
remediation_assertion_idsNoOptional comma-separated assertion IDs that evidence the fix, linking the remediation to the assertions that prove it. Empty by default.

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden; it discloses that the tool is mutating, that it records a MANUAL transition, and that auto_resolved is deliberately refused because setting it by hand would forge a platform-only claim — a strong, non-obvious behavioral disclosure. It does not cover permissions/auth or reversibility, so it falls short of a 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences, front-loaded with the purpose and the mutating statement before the usage and exception details. Every sentence carries information, though the second and third run long and mix usage guidance with rationale.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With an output schema present, return values need not be explained, and the schema documents the parameters. The description adds the lifecycle framing, sibling routing, and the auto_resolved exclusion, leaving only minor gaps such as required permissions. Nearly complete for a mutation tool with no annotations.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 86%, so the schema already documents status, reason, notes, finding_id, and remediation_assertion_ids. The description largely restates the auto_resolved exclusion that the status field already explains, adding only the manual-vs-machine framing; at high coverage the baseline of 3 is correct.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('Advance a finding through its lifecycle. Mutating: updates the finding's status and metadata') and explicitly distinguishes itself from submit_findings/list_findings (creation/reading) and remediate_finding (automatic cleanup). An agent can tell exactly what this tool does and does not do without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives explicit when-to-use cases ('acknowledge, remediate, verify, or dismiss a finding previously recorded by submit_findings / list_findings') and a when-not / alternative ('for gaps whose kind has an automatic fix, remediate_finding performs the actual cleanup instead'). It also names the dismissal path and its condition (with a reason). Routing is unambiguous.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

update_organizationUpdate OrganizationA

Set per-organization level grades for IEC 62443-4-1 and NIST CSF.

Admin-only: the backend requires the caller to be an admin in the organization (or a superadmin). Non-admins will get a 403; do not invoke this tool unless you've verified admin role for the target org.

target_ml is the IEC 62443-4-1 Maturity Level the organization targets for its secure-development program (1-5). csf_tier is the NIST CSF Tier the organization targets for its cybersecurity risk-management posture (1-4).

Because None on the wire is indistinguishable from "field omitted", pass clear_target_ml=True or clear_csf_tier=True to explicitly reset a value to NULL. Omitting both the value and its clear_* flag leaves the existing server-side value untouched.

ParametersJSON Schema
NameRequiredDescriptionDefault
org_idYesOrganization ID.
csf_tierNoNIST CSF Tier (1-4), or ``None`` to leave unchanged.
target_mlNoIEC 62443-4-1 Maturity Level (1-5), or ``None`` to leave unchanged.
clear_csf_tierNoExplicitly reset ``csf_tier`` to NULL.
server_versionYes
clear_target_mlNoExplicitly reset ``target_ml`` to NULL.

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.9/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries full behavioral burden. It discloses the auth requirement, the 403 failure mode, the wire-level ambiguity of None, the clear_* flags for resetting values, and the behavior of leaving omitted values untouched. This goes well beyond what the schema alone conveys.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is organized into four purposeful paragraphs: purpose, auth warning, parameter definitions, and None/clear semantics. No filler; each sentence explains a behavior or precondition the agent needs before calling.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Very complete for a six-parameter tool: it covers auth, parameter ranges, reset semantics, and mutation behavior. It does not explain the required server_version parameter, which is also undocumented in the schema, so the definition has one minor blind spot.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The description adds ranges for target_ml (1-5) and csf_tier (1-4), explains the meaning of None versus omission, and clarifies the clear_* flags and their relationship to the value parameters. Schema coverage is high, but this description resolves semantics that the schema leaves ambiguous.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource: 'Set per-organization level grades for IEC 62443-4-1 and NIST CSF.' This clearly distinguishes it from the many compliance/control sibling tools and explains what 'update organization' means in this context.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It explicitly states the precondition: admin or superadmin required, and says not to invoke unless admin role is verified. It also implies the use case (updating org-level grades) and the when-not (non-admins will get a 403), so an agent can decide when this tool is appropriate.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

update_threat_modelUpdate Threat ModelA

Change a threat model's metadata: its name, its parent, where its description came from. Mutating; pass only what changes.

  • name (1-120 chars) renames it; no new version. Titles are unique within a workspace, case-insensitive (409 on a clash).

  • parent_id wires it under a parent on the recursive composition tree, so it inherits the parent's topology and objectives; clear_parent=True makes it a tree root. Cycles (409) and chains past the platform's maximum depth (400) are refused. No new version.

  • provenance_kind (code, ticket, document, manual, mixed) records where the description came from, with the other provenance_* values. code with provenance_commit_sha means the code is authoritative and the model follows it (reconcile_model measures it against the code); any other kind means the description is intent and the code is measured against it. Bumps the model version.

Changes apply in that order. Returns {model_id, name?, parent?, provenance?}, one entry per change applied. A failure raises and names the changes already applied.

ParametersJSON Schema
NameRequiredDescriptionDefault
nameNo
model_idYes
parent_idNo
clear_parentNo
provenance_refNo
server_versionYes
provenance_kindNo
provenance_repo_urlNo
provenance_commit_shaNo
provenance_source_refNo
provenance_source_urlNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.2/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations the description carries the full burden and does so: it discloses conflict codes (409 on name clash/cycle, 400 on max depth), which changes bump the model version versus which do not, the fixed application order of changes, and the partial-failure contract ('a failure raises and names the changes already applied'). That is unusually rich mutation semantics.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the purpose and the 'pass only what changes' rule, then bulleted per-parameter detail that earns its length. Dense but every clause carries information; slightly long for an agent skimming many tools.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a mutating, 11-param tool with no annotations, it covers ordering, error codes, version effects and failure reporting, and even restates the return shape despite an output schema existing. Missing pieces are the required server_version semantics and any auth/permission prerequisites.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0% across 11 params, yet the description explains name constraints, parent_id/clear_parent semantics, the provenance_kind enum values, and the meaning of provenance_commit_sha and the 'other provenance_* values'. The two required params, model_id and especially server_version (apparently a concurrency token), are left undocumented, which is the remaining gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource and narrows scope precisely to metadata (name, parent, provenance), with each field's effect spelled out. However it never names or contrasts a sibling (e.g. refine_threat_model or delete_threat_model), leaving the agent to infer which tool owns content-level changes versus metadata.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

'Mutating; pass only what changes' gives clear partial-update guidance, and per-field conditions (clear_parent=True to root, provenance_kind=code to make code authoritative) tell the agent when each argument is appropriate. Still no explicit when-to-use-this vs. the sibling refine/delete/undo tools.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 13 tool updatesv0.85.0
    • Changedadd_model_to_group3 fields changed
      • changedInput schema / properties / group_id / description
        Previous value: -"ID of the tag or system."New value: +"ID of the tag."
      • removedInput schema / properties / kind
        Removed value: -{
        -  "description": "``\"tag\"`` or ``\"system\"``.",
        -  "type": "string"
        -}
      • changedInput schema / required
        Previous value: -[
        -  "server_version",
        -  "kind",
        -  "group_id",
        -  "model_id"
        -]New value: +[
        +  "server_version",
        +  "group_id",
        +  "model_id"
        +]
    • Changedcreate_group4 fields changed
      • removedInput schema / properties / kind
        Removed value: -{
        -  "description": "``\"tag\"`` or ``\"system\"``.",
        -  "type": "string"
        -}
      • changedInput schema / properties / model_ids / description
        Previous value: -"optional initial member model ids — ``\"tag\"`` only."New value: +"optional initial member model ids."
      • changedInput schema / properties / name / description
        Previous value: -"the group name (unique within the workspace for its kind)."New value: +"the tag name (unique within the workspace)."
      • changedInput schema / required
        Previous value: -[
        -  "server_version",
        -  "kind",
        -  "name"
        -]New value: +[
        +  "server_version",
        +  "name"
        +]
    • Changeddelete_group3 fields changed
      • addedInput schema / properties / group_id
        Added value: +{
        +  "description": "ID of the tag to delete.",
        +  "type": "string"
        +}
      • removedInput schema / properties / tag_id
        Removed value: -{
        -  "description": "ID of the tag to delete.",
        -  "type": "string"
        -}
      • changedInput schema / required
        Previous value: -[
        -  "server_version",
        -  "tag_id"
        -]New value: +[
        +  "server_version",
        +  "group_id"
        +]
    • Changedget_compliance_report6 fields changed
      • changedInput schema / properties / limit / description
        Previous value: -"max requirement rows to return, 0 = no explicit limit (model/system scopes only)."New value: +"max requirement rows to return, 0 = no explicit limit (model scope only)."
      • changedInput schema / properties / offset / description
        Previous value: -"skip the first N requirement rows, pagination (model/system scopes only). Default 0."New value: +"skip the first N requirement rows, pagination (model scope only). Default 0."
      • changedInput schema / properties / scope / description
        Previous value: -"report boundary — \"model\", \"system\", or \"tag\"."New value: +"report boundary — \"model\" or \"tag\"."
      • changedInput schema / properties / scope / enum
        Previous value: -[
        -  "model",
        -  "system",
        -  "tag"
        -]New value: +[
        +  "model",
        +  "tag"
        +]
      • changedInput schema / properties / scope_id / description
        Previous value: -"id of the model, system, or tag selected by ``scope``."New value: +"id of the model or tag selected by ``scope``."
      • changedInput schema / properties / status / description
        Previous value: -"optional per-requirement status filter (model/system scopes only)."New value: +"optional per-requirement status filter (model scope only)."
    • Changedget_group3 fields changed
      • addedInput schema / properties / group_id
        Added value: +{
        +  "description": "ID of the tag.",
        +  "type": "string"
        +}
      • removedInput schema / properties / system_id
        Removed value: -{
        -  "description": "ID of the system to retrieve.",
        -  "type": "string"
        -}
      • changedInput schema / required
        Previous value: -[
        -  "server_version",
        -  "system_id"
        -]New value: +[
        +  "server_version",
        +  "group_id"
        +]
    • Addedget_group_dependencies
    • Changedget_risk_view3 fields changed
      • changedInput schema / properties / scope / description
        Previous value: -"aggregation boundary — \"model\", \"system\", or \"tag\"."New value: +"aggregation boundary — \"model\" or \"tag\"."
      • changedInput schema / properties / scope / enum
        Previous value: -[
        -  "model",
        -  "system",
        -  "tag"
        -]New value: +[
        +  "model",
        +  "tag"
        +]
      • changedInput schema / properties / scope_id / description
        Previous value: -"id of the model, system, or tag selected by ``scope``."New value: +"id of the model or tag selected by ``scope``."
    • Removedget_system_dependencies
    • Removedlink_system_dependency
    • Changedlist_groups2 fields changed
      • removedInput schema / properties / kind
        Removed value: -{
        -  "description": "``\"tag\"`` or ``\"system\"``.",
        -  "type": "string"
        -}
      • changedInput schema / required
        Previous value: -[
        -  "server_version",
        -  "kind"
        -]New value: +[
        +  "server_version"
        +]
    • Changedlist_model_groups1 field changed
      • changedInput schema / properties / model_id / description
        Previous value: -"the model whose groups (tags) to list."New value: +"the model whose tags to list."
    • Changedremove_model_from_group3 fields changed
      • addedInput schema / properties / group_id
        Added value: +{
        +  "description": "the tag.",
        +  "type": "string"
        +}
      • removedInput schema / properties / tag_id
        Removed value: -{
        -  "description": "the tag.",
        -  "type": "string"
        -}
      • changedInput schema / required
        Previous value: -[
        -  "server_version",
        -  "tag_id",
        -  "model_id"
        -]New value: +[
        +  "server_version",
        +  "group_id",
        +  "model_id"
        +]
    • Changedselect_compliance_frameworks3 fields changed
      • changedInput schema / properties / scope / description
        Previous value: -"target boundary — \"model\", \"system\", or \"tag\"."New value: +"target boundary — \"model\" or \"tag\"."
      • changedInput schema / properties / scope / enum
        Previous value: -[
        -  "model",
        -  "system",
        -  "tag"
        -]New value: +[
        +  "model",
        +  "tag"
        +]
      • changedInput schema / properties / scope_id / description
        Previous value: -"id of the model, system, or tag selected by ``scope``."New value: +"id of the model or tag selected by ``scope``."
  2. 58 tool updatesv0.84.0
    • Removedaccept_coverage_divergences
    • Removedadd_evidence
    • Removedapply_certain_reconciliation_match
    • Removedapply_finding_remediation
    • Changedattach_foundation6 fields changed
      • addedInput schema / properties / selections / anyOf
        Added value: +[
        +  {
        +    "items": {
        +      "additionalProperties": true,
        +      "type": "object"
        +    },
        +    "type": "array"
        +  },
        +  {
        +    "type": "null"
        +  }
        +]
      • addedInput schema / properties / selections / default
        Added value: +null
      • changedInput schema / properties / selections / description
        Previous value: -"list of {source_objective_id, provider_control_id} dicts."New value: +"the pairs to delegate; omit to get the candidates."
      • removedInput schema / properties / selections / items
        Removed value: -{
        -  "additionalProperties": true,
        -  "type": "object"
        -}
      • removedInput schema / properties / selections / type
        Removed value: -"array"
      • changedInput schema / required
        Previous value: -[
        -  "server_version",
        -  "model_id",
        -  "foundation_model_id",
        -  "selections"
        -]New value: +[
        +  "server_version",
        +  "model_id",
        +  "foundation_model_id"
        +]
    • Removedcheck_functional_gaps
    • Removedconfirm_reliance
    • Removedcreate_reliance
    • Addeddecide_reconciliation_candidate
    • Removeddelete_reliance
    • Addeddiscard_control_build
    • Removeddismiss_verdict_divergences
    • Addededit_evidence
    • Changedgenerate_threat_model1 field changed
      • changedInput schema / properties / provenance_kind / description
        Previous value: -"Where the description came from, one of\n``code``, ``ticket``, ``document``, ``manual``, ``mixed``.\nEmpty (default) records nothing. For an existing repository\npass ``provenance_kind=\"code\"`` with ``provenance_repo_url``\nand ``provenance_commit_sha`` (the HEAD you gathered from):\nthe code is then authoritative and the model follows it.\nAny other kind means the description is intent and the code\nis measured against it. The same record can be set later\nwith ``set_model_provenance``."New value: +"Where the description came from, one of\n``code``, ``ticket``, ``document``, ``manual``, ``mixed``.\nEmpty (default) records nothing. For an existing repository\npass ``provenance_kind=\"code\"`` with ``provenance_repo_url``\nand ``provenance_commit_sha`` (the HEAD you gathered from):\nthe code is then authoritative and the model follows it.\nAny other kind means the description is intent and the code\nis measured against it. The same record can be set later\nwith ``update_threat_model``."
    • Addedget_capabilities
    • Removedget_capability
    • Addedget_composition
    • Removedget_composition_overview
    • Changedget_control_generation_status1 field changed
      • removedInput schema / properties / model_id / description
        Removed value: -"ID of the threat model whose control-generation status to poll."
    • Changedget_control_work_order2 fields changed
      • removedInput schema / properties / control_id / description
        Removed value: -"ID of the control to implement (e.g. \"CTRL-03\")."
      • removedInput schema / properties / model_id / description
        Removed value: -"ID of the threat model."
    • Changedget_controls11 fields changed
      • removedInput schema / properties / co_id / description
        Removed value: -"List-mode filter — control objective ID."
      • removedInput schema / properties / component_id / description
        Removed value: -"List-mode filter — component ID (e.g., \"CMP1\")."
      • removedInput schema / properties / control_id / description
        Removed value: -"If set, detail mode — return this one control's full\nrecord directly (e.g. ``CTL-12``). If omitted, list mode."
      • removedInput schema / properties / include_deleted / description
        Removed value: -"List mode — include soft-deleted controls\n(default False)."
      • removedInput schema / properties / include_orphaned / description
        Removed value: -"List mode — include controls mapped only to\ntombstoned COs (default False)."
      • removedInput schema / properties / limit / description
        Removed value: -"List mode — max controls to return (0 = all)."
      • removedInput schema / properties / model_id / description
        Removed value: -"ID of the threat model."
      • removedInput schema / properties / offset / description
        Removed value: -"List mode — skip the first N controls (pagination)."
      • removedInput schema / properties / status / description
        Removed value: -"List-mode filter — \"implemented\", \"not_implemented\", or\n\"verified\"."
      • removedInput schema / properties / summary_only / description
        Removed value: -"List mode — if True, returns only id, description,\nstatus, assertion_count, and assumed_by per control (much\nsmaller response)."
      • removedInput schema / properties / version / description
        Removed value: -"Detail mode only — model version to read the control\nfrom. 0 (default) uses the latest. Ignored in list mode."
    • Removedget_effective_coverage
    • Changedget_functional_coverage2 fields changed
      • addedInput schema / properties / gaps_only
        Added value: +{
        +  "default": false,
        +  "description": "Return only the actionable gaps (default False).",
        +  "type": "boolean"
        +}
      • changedInput schema / properties / model_id / description
        Previous value: -"ID of the threat model whose functional coverage to report."New value: +"ID of the threat model."
    • Removedget_functional_test_sufficiency
    • Changedget_reachability_verdicts6 fields changed
      • removedInput schema / properties / co_id / description
        Removed value: -"FLAT mode only. Optional CO id — when set, returns a single\nverdict; 404 if the CO doesn't exist or is tombstoned. Ignored\nwhen ``composed=True``."
      • removedInput schema / properties / composed / description
        Removed value: -"When False (default), derive over this model's own\ntopology (flat). When True, derive over the composed effective\ntree (own ⊕ inherited)."
      • removedInput schema / properties / kind_filter / description
        Removed value: -"COMPOSED mode only. Restrict verdicts to one kind —\none of ``\"reachable\" | \"unreachable\" | \"indeterminate\"``. Named\n``kind_filter`` (not ``kind``) to disambiguate from the verdict\nobject's own ``kind`` field. When omitted, all verdict kinds are\nreturned. Ignored when ``composed=False``."
      • removedInput schema / properties / model_id / description
        Removed value: -"ID of the threat model."
      • removedInput schema / properties / page / description
        Removed value: -"COMPOSED mode only. 1-indexed page number (default ``1``).\nIgnored when ``composed=False``."
      • removedInput schema / properties / page_size / description
        Removed value: -"COMPOSED mode only. Verdicts per page (default ``100``).\nIgnored when ``composed=False``."
    • Changedget_sufficiency5 fields changed
      • addedInput schema / properties / control_id / default
        Added value: +""
      • removedInput schema / properties / control_id / description
        Removed value: -"ID of the control (e.g., \"CTRL-01\")."
      • addedInput schema / properties / functional_test_id
        Added value: +{
        +  "default": "",
        +  "type": "string"
        +}
      • removedInput schema / properties / model_id / description
        Removed value: -"ID of the threat model."
      • changedInput schema / required
        Previous value: -[
        -  "server_version",
        -  "model_id",
        -  "control_id"
        -]New value: +[
        +  "server_version",
        +  "model_id"
        +]
    • Changedget_threat_model1 field changed
      • changedInput schema / properties / include_cos / description
        Previous value: -"Include control objectives inline."New value: +"Include control objectives inline (default False:\nthe answer carries no ``control_objectives`` key; read them\nwith ``get_control_objectives``)."
    • Changedimport_compliance_framework1 field changed
      • changedInput schema / properties / framework_json / description
        Previous value: -"A JSON string containing the framework body.\n(String not dict so the JSON shape stays explicit on the wire.)"New value: +"The framework body as a JSON string."
    • Addedjudge_imported_controls
    • Changedjudge_objectives2 fields changed
      • changedInput schema / properties / co_ids / description
        Previous value: -"Optional comma-separated objective IDs to restrict the call\nto. Omit for every objective with no judgement for its current\ncontrols and none queued."New value: +"Optional comma-separated objective IDs; omit for every\nobjective with no judgement and none queued."
      • changedInput schema / properties / confirm_estimate / description
        Previous value: -"False (default) returns the estimate and queues\nnothing; True queues the judgements."New value: +"False (default) estimates; True queues."
    • Removedlist_capabilities
    • Addedlist_control_revisions
    • Removedlist_effective_attack_paths
    • Removedlist_effective_control_objectives
    • Removedlist_effective_entities
    • Changedlist_reconciliation_candidates4 fields changed
      • removedInput schema / properties / disposition / description
        Removed value: -"Which side of the queue to read — ``\"active\"``\n(default, open candidates) or ``\"rejected\"`` (persisted\nnot-a-duplicate decisions)."
      • removedInput schema / properties / model_id / description
        Removed value: -"ID of the descendant threat model."
      • removedInput schema / properties / page / description
        Removed value: -"ACTIVE disposition only. 1-indexed page number. Default 1.\nIgnored when ``disposition=\"rejected\"``."
      • removedInput schema / properties / page_size / description
        Removed value: -"ACTIVE disposition only. Items per page. Default 50.\nIgnored when ``disposition=\"rejected\"``."
    • Addedmanage_reliance
    • Changedmodel_coherence_report2 fields changed
      • removedInput schema / properties / co_id / description
        Removed value: -"Optional CO id to scope the report to a single CO."
      • removedInput schema / properties / model_id / description
        Removed value: -"ID of the threat model."
    • Removedpreview_finding_remediation
    • Removedpreview_undo_composition
    • Removedpropose_attach_foundation
    • Changedrecompute_verdicts3 fields changed
      • removedInput schema / properties / dry_run
        Removed value: -{
        -  "default": false,
        -  "description": "When True, return only the pre-flight estimate and enqueue\nnothing. When False (default), enqueue the recompute.",
        -  "type": "boolean"
        -}
      • addedInput schema / properties / mode
        Added value: +{
        +  "default": "quote",
        +  "enum": [
        +    "quote",
        +    "recompute",
        +    "retry_parked"
        +  ],
        +  "type": "string"
        +}
      • removedInput schema / properties / model_id / description
        Removed value: -"ID of the threat model to re-evaluate (or estimate for)."
    • Removedreject_reconciliation_candidate
    • Addedremediate_finding
    • Removedremove_evidence
    • Removedrename_threat_model
    • Addedresolve_verdict_divergences
    • Removedretry_verdicts
    • Removedset_model_provenance
    • Removedset_threat_model_parent
    • Addedstart_control_build
    • Changedstrengthen_controls4 fields changed
      • changedInput schema / properties / co_ids / description
        Previous value: -"Optional comma-separated objective IDs to restrict the run\nto, of those the diagnosis lists as uncovered or undecided. Omit\nfor all of them."New value: +"Optional comma-separated objective IDs, of those diagnosed\nuncovered or undecided; omit for all."
      • changedInput schema / properties / confirm_estimate / description
        Previous value: -"False (default) returns the estimate and starts\nnothing; True starts the run."New value: +"False (default) estimates; True starts the run."
      • addedInput schema / properties / model_version
        Added value: +{
        +  "anyOf": [
        +    {
        +      "type": "integer"
        +    },
        +    {
        +      "type": "null"
        +    }
        +  ],
        +  "default": null,
        +  "description": "With ``confirm_estimate=True``: the estimate's value."
        +}
      • addedInput schema / properties / set_revision
        Added value: +{
        +  "anyOf": [
        +    {
        +      "type": "integer"
        +    },
        +    {
        +      "type": "null"
        +    }
        +  ],
        +  "default": null,
        +  "description": "With ``confirm_estimate=True``: the estimate's value."
        +}
    • Changedsubmit_assertions1 field changed
      • addedInput schema / properties / functional_test_id
        Added value: +{
        +  "anyOf": [
        +    {
        +      "type": "string"
        +    },
        +    {
        +      "type": "null"
        +    }
        +  ],
        +  "default": null
        +}
    • Removedsubmit_functional_test_assertions
    • Changedundo_composition_event4 fields changed
      • addedInput schema / properties / dry_run
        Added value: +{
        +  "default": true,
        +  "type": "boolean"
        +}
      • removedInput schema / properties / event_id / description
        Removed value: -"Either the surrogate id of the forward\n``lift_applied`` / ``split_applied`` activity event, or the\nstructured ``lift_id`` / ``split_id`` carried in the event\npayload."
      • removedInput schema / properties / event_type / description
        Removed value: -"Which forward composition event to undo. One of:\n  - ``\"lift\"``: undo a ``lift_applied`` event. On success,\n    persists the inverse across the LCA + every affected\n    source descendant and emits a ``lift_undone`` event. The\n    returned ``models`` block carries ``lca_model`` and\n    ``source_descendant_models``.\n  - ``\"split\"``: undo a ``split_applied`` event. On success,\n    restores the ancestor's entity, tombstones the duplicated\n    copies on every target descendant, persists across all\n    affected models, and emits a ``split_undone`` event. The\n    returned ``models`` block carries ``ancestor_model`` and\n    ``descendant_models``."
      • removedInput schema / properties / model_id / description
        Removed value: -"The model whose composition view originated the event.\nMust match the cited event's ``threat_model_id`` — the server\nrejects cross-model citations with 404."
    • Addedundo_model_change
    • Removedunreject_reconciliation_candidate
    • Addedupdate_threat_model
  3. 7 tool updatesv0.83.1
    • Changedcreate_proposal1 field changed
      • changedInput schema / properties / kind / description
        Previous value: -"One of ``add_component`` (payload ``{name, repo_url?, path?,\ntrust_boundary_ids?}``), ``remove_component`` (payload\n``{component_id}``), ``design_change`` (payload ``{target_kind:\n\"attacker\"|\"asset\", target_id, design_move}``; take ``design_move``\nfrom ``get_design_leverage``)."New value: +"One of ``add_component`` (payload ``{name, repo_url?, path?,\ntrust_boundary_ids?}``), ``remove_component`` (payload\n``{component_id}``), ``design_change`` (payload ``{target_kind:\n\"attacker\"|\"asset\", target_id, design_move}``; take ``design_move``\nfrom ``get_design_leverage``), ``assumption`` (payload ``{co_id,\ngroup_id, precondition, assumption_id?, gap?}``: a precondition\nabout the environment that only something outside this system\ncan meet, one declarative sentence; ``assumption_id`` names an\nexisting assumption that states it). A precondition a person\nalready rejected for that objective is refused."
    • Changeddecide_proposal1 field changed
      • addedInput schema / properties / expires_at
        Added value: +{
        +  "default": "",
        +  "description": "ISO 8601 date an accepted assumption lapses (e.g.\n\"2027-03-29T00:00:00Z\"). Applies to accepting an ``assumption``\nproposal; omitted, the acceptance lasts a year.",
        +  "type": "string"
        +}
    • Addedjudge_objective
    • Addedjudge_objectives
    • Changedlist_decisions1 field changed
      • changedInput schema / properties / decision / description
        Previous value: -"Optional kind filter, one of ``finding_dismissed``,\n``finding_remediated``, ``risk_accepted``,\n``not_applicable_declared``, ``proposal_accepted``,\n``proposal_rejected``, ``proposal_reverted``,\n``escalation_resolved``. Empty (default) returns every kind."New value: +"Optional kind filter, one of ``finding_dismissed``,\n``finding_remediated``, ``risk_accepted``,\n``not_applicable_declared``, ``proposal_accepted``,\n``proposal_rejected``, ``proposal_reverted``,\n``assumption_accepted``, ``escalation_resolved``. Empty\n(default) returns every kind."
    • Addedpause_control_generation
    • Addedstrengthen_controls
  4. 1 tool updatev0.78.1
    • Addedresume_control_generation
  5. 3 tool updatesv0.77.1
    • Changedrefine_control1 field changed
      • changedInput schema / properties / justification / description
        Previous value: -"Why this refinement is appropriate (min 10 chars)."New value: +"Why this refinement is appropriate (10 to 2000\ncharacters)."
    • Changedset_control_assumption_groups1 field changed
      • changedInput schema / properties / justification / description
        Previous value: -"Why this group structure is appropriate (min 10 chars\nwhen groups is non-empty; optional when clearing)."New value: +"Why this group structure is appropriate (10 to 2000\ncharacters when groups is non-empty; optional when clearing, at\nmost 2000 characters if given)."
    • Changedset_mitigation_groups1 field changed
      • changedInput schema / properties / justification / description
        Previous value: -"Why this group structure is appropriate (min 10 chars)."New value: +"Why this group structure is appropriate (10 to 2000\ncharacters). The gate weighs it alongside the objective and the\ncontrols' own descriptions, so state the reasoning, not evidence."
  6. 4 tool updatesv0.77.0
    • Changedadd_attacker2 fields changed
      • addedInput schema / properties / change_reason
        Added value: +{
        +  "anyOf": [
        +    {
        +      "type": "string"
        +    },
        +    {
        +      "type": "null"
        +    }
        +  ],
        +  "default": null,
        +  "description": "Required when ``surface_extent`` is supplied —\ndocuments the declaration for the audit trail."
        +}
      • addedInput schema / properties / surface_extent
        Added value: +{
        +  "anyOf": [
        +    {
        +      "type": "string"
        +    },
        +    {
        +      "type": "null"
        +    }
        +  ],
        +  "default": null,
        +  "description": "``\"whole\"`` when, from its position, the attacker's\noperations range over ANY entry of the interface it reaches\n(any endpoint, request, row, file, message or frame). Recorded\nas attested by this call and requires ``change_reason``. Omit\nto leave it undeclared, which is the ordinary case."
        +}
    • Changedadd_trust_boundary1 field changed
      • changedInput schema / properties / sealed / description
        Previous value: -"Optional. Set True to declare the boundary has NO lateral\ningress — the only way into its zone is crossing the perimeter\n(an air-gap / network-segmented enclave). A sealed boundary that\nblocks the attacker's vector lets reachability decisively rule the\nasset unreachable instead of indeterminate. Default False (assume a\nlateral pivot is possible). Set it only when the isolation is real\nand attestable."New value: +"Optional. Set True to declare the boundary has NO lateral\ningress — the only way into its zone is crossing the perimeter\n(an air-gap / network-segmented enclave). On its own this is a\nsuggestion: only an ATTESTED seal lets reachability decisively\nrule an asset unreachable instead of indeterminate, and the\nattestation is recorded with ``edit_trust_boundary``\n(``seal_source=\"attested\"`` with a ``change_reason``). Default\nFalse (assume a lateral pivot is possible). Set it only when the\nisolation is real and attestable."
    • Changededit_attacker3 fields changed
      • addedInput schema / properties / attest_surface_extent
        Added value: +{
        +  "anyOf": [
        +    {
        +      "type": "boolean"
        +    },
        +    {
        +      "type": "null"
        +    }
        +  ],
        +  "default": null,
        +  "description": "Record the extent already on the attacker as\nattested, without changing its value. Pass ``true`` to attest;\nrequires ``change_reason``."
        +}
      • changedInput schema / properties / change_reason / description
        Previous value: -"Required when any factor field is supplied —\ndocuments the operator override of LLM-generated factors."New value: +"Required when any factor field, ``surface_extent``\nor ``attest_surface_extent`` is supplied — documents the\noperator override for the audit trail."
      • addedInput schema / properties / surface_extent
        Added value: +{
        +  "anyOf": [
        +    {
        +      "type": "string"
        +    },
        +    {
        +      "type": "null"
        +    }
        +  ],
        +  "default": null,
        +  "description": "``\"whole\"`` (operations range over ANY entry of the\ninterface reached) or ``\"point\"`` (one named entry). Supplying\nit attests it; requires ``change_reason``."
        +}
    • Changedsubmit_functional_test_assertions1 field changed
      • changedInput schema / properties / assertions_json / description
        Previous value: -"JSON array of assertion objects, each {\"type\": \"test_attested\" | \"test_exists\" | ..., \"params\": {...}, \"description\": \"...\", \"repo\": \"<owner>/<repo>\"}. Every assertion must carry an explicit repo, or the \"no_repo\" sentinel when the check is not tied to a repository."New value: +"JSON array of assertion objects, each {\"type\": \"test_attested\" | \"test_exists\" | ..., \"params\": {...}, \"description\": \"...\", \"repo\": \"<owner>/<repo>\"}. Every assertion must carry an explicit repo, or the \"no_repo\" sentinel when the check is not tied to a repository. These assertions count toward functional conformance; a ``covers`` declaration is refused here, because a binding to a control clause is declared on ``submit_assertions``."
  7. 11 tool updatesv0.75.0
    • Addedcreate_proposal
    • Addeddecide_proposal
    • Changedgenerate_threat_model6 fields changed
      • addedInput schema / properties / provenance_commit_sha
        Added value: +{
        +  "default": "",
        +  "description": "Commit SHA the description was gathered\nat (``code`` kind).",
        +  "type": "string"
        +}
      • addedInput schema / properties / provenance_kind
        Added value: +{
        +  "default": "",
        +  "description": "Where the description came from, one of\n``code``, ``ticket``, ``document``, ``manual``, ``mixed``.\nEmpty (default) records nothing. For an existing repository\npass ``provenance_kind=\"code\"`` with ``provenance_repo_url``\nand ``provenance_commit_sha`` (the HEAD you gathered from):\nthe code is then authoritative and the model follows it.\nAny other kind means the description is intent and the code\nis measured against it. The same record can be set later\nwith ``set_model_provenance``.",
        +  "type": "string"
        +}
      • addedInput schema / properties / provenance_ref
        Added value: +{
        +  "default": "",
        +  "description": "Branch or tag name at that commit (optional).",
        +  "type": "string"
        +}
      • addedInput schema / properties / provenance_repo_url
        Added value: +{
        +  "default": "",
        +  "description": "Repository URL the description was\ngathered from (``code`` kind).",
        +  "type": "string"
        +}
      • addedInput schema / properties / provenance_source_ref
        Added value: +{
        +  "default": "",
        +  "description": "Identifier of the ticket or document the\ndescription came from (``ticket`` / ``document`` kinds).",
        +  "type": "string"
        +}
      • addedInput schema / properties / provenance_source_url
        Added value: +{
        +  "default": "",
        +  "description": "URL of that ticket or document.",
        +  "type": "string"
        +}
    • Addedget_control_work_order
    • Addedget_design_leverage
    • Addedlist_decisions
    • Addedlist_proposals
    • Removedlist_workspaces
    • Addedreconcile_model
    • Addedset_model_provenance
    • Changedsubmit_functional_test_assertions1 field changed
      • changedInput schema / properties / assertions_json / description
        Previous value: -"JSON array of assertion objects, each {\"type\": \"test_passes\" | \"test_exists\" | ..., \"params\": {...}, \"description\": \"...\", \"repo\": \"<owner>/<repo>\"}. Every assertion must carry an explicit repo, or the \"no_repo\" sentinel when the check is not tied to a repository."New value: +"JSON array of assertion objects, each {\"type\": \"test_attested\" | \"test_exists\" | ..., \"params\": {...}, \"description\": \"...\", \"repo\": \"<owner>/<repo>\"}. Every assertion must carry an explicit repo, or the \"no_repo\" sentinel when the check is not tied to a repository."
  8. 2 tool updatesv0.71.1
    • Addedget_assertion_types
    • Changedget_controls1 field changed
      • changedInput schema / properties / control_id / description
        Previous value: -"If set, detail mode �� return this one control's full\nrecord directly (e.g. ``CTL-12``). If omitted, list mode."New value: +"If set, detail mode — return this one control's full\nrecord directly (e.g. ``CTL-12``). If omitted, list mode."
  9. 5 tool updatesv0.71.0
    • Addedcreate_co_disposition
    • Changedget_controls1 field changed
      • changedInput schema / properties / control_id / description
        Previous value: -"If set, detail mode — return this one control's full\nrecord directly (e.g. ``CTL-12``). If omitted, list mode."New value: +"If set, detail mode �� return this one control's full\nrecord directly (e.g. ``CTL-12``). If omitted, list mode."
    • Addedlist_co_dispositions
    • Changedlist_findings1 field changed
      • changedInput schema / properties / status / description
        Previous value: -"Optional lifecycle filter, one of \"discovered\", \"acknowledged\", \"remediated\", \"verified\", \"dismissed\". Empty (default) returns all statuses."New value: +"Optional lifecycle filter, one of \"discovered\", \"acknowledged\", \"remediated\", \"verified\", \"dismissed\", \"auto_resolved\". Empty (default) returns all statuses.\n``auto_resolved`` is closed by the platform, not by a person: the\ncondition that produced the finding is no longer reproduced. It is\ndeliberately distinct from ``remediated``/``verified`` (a person\nfixed and confirmed it) and from ``dismissed`` (a person judged it\nnot worth fixing) — \"the gap is gone\" and \"the gap does not matter\"\nare opposite statements about residual risk, so they never share a\nstatus."
    • Changedupdate_finding1 field changed
      • changedInput schema / properties / status / description
        Previous value: -"New lifecycle status, one of \"discovered\", \"acknowledged\", \"remediated\", \"verified\", \"dismissed\"."New value: +"New lifecycle status, one of \"discovered\", \"acknowledged\", \"remediated\", \"verified\", \"dismissed\". ``auto_resolved`` is NOT settable here: it asserts that a condition is no longer reproduced, which is a claim only the platform can make from its own re-evaluation. Setting it by hand would forge that claim, so this tool refuses it — use ``dismissed`` (with a reason) to record that a gap does not matter, which is the judgment a person is entitled to make."
  10. 85 tool updatesv0.68.2
    • Changedadd_functional_test1 field changed
      • changedInput schema / properties / functional_objective_ids / description
        Previous value: -"Comma-separated objective ids the test satisfies (at least one required; get them from list_functional_objectives)."New value: +"Comma-separated objective ids the test satisfies (at least one required; get them from get_functional_objectives)."
    • Addedadd_model_to_group
    • Removedadd_model_to_system
    • Removedadd_model_to_tag
    • Removedassign_asset_to_components
    • Removedassign_control_to_components
    • Addedassign_to_components
    • Removedassume_control
    • Removedauto_remediate
    • Addedauto_remediate_compliance
    • Addedcreate_group
    • Removedcreate_system
    • Removedcreate_tag
    • Addeddelete_group
    • Removeddelete_tag
    • Addedexport_report
    • Removedexport_tag_report
    • Removedexport_threat_model
    • Removedexport_threat_model_archive
    • Removedget_asset
    • Removedget_assumption
    • Removedget_attacker
    • Changedget_compliance_report9 fields changed
      • changedInput schema / properties / framework_id / description
        Previous value: -"ID of the compliance framework (as listed by list_compliance_frameworks)."New value: +"framework to report on (already selected at this scope; see ``list_compliance_frameworks``)."
      • changedInput schema / properties / level / description
        Previous value: -"Optional level filter for level-aware frameworks — returns only requirements at or below this level (e.g. 1 for L1 only). Omit for all levels."New value: +"optional level filter; omit for all levels."
      • changedInput schema / properties / limit / description
        Previous value: -"Maximum requirement rows to return. Default 0 = no explicit limit."New value: +"max requirement rows to return, 0 = no explicit limit (model/system scopes only)."
      • removedInput schema / properties / model_id
        Removed value: -{
        -  "description": "ID of the threat model.",
        -  "type": "string"
        -}
      • changedInput schema / properties / offset / description
        Previous value: -"Number of requirement rows to skip, for pagination. Default 0."New value: +"skip the first N requirement rows, pagination (model/system scopes only). Default 0."
      • addedInput schema / properties / scope
        Added value: +{
        +  "description": "report boundary — \"model\", \"system\", or \"tag\".",
        +  "enum": [
        +    "model",
        +    "system",
        +    "tag"
        +  ],
        +  "type": "string"
        +}
      • addedInput schema / properties / scope_id
        Added value: +{
        +  "description": "id of the model, system, or tag selected by ``scope``.",
        +  "type": "string"
        +}
      • changedInput schema / properties / status / description
        Previous value: -"Optional status filter: \"covered\", \"partial\", \"uncovered\", \"unmapped\", or \"excluded\". Empty = all statuses."New value: +"optional per-requirement status filter (model/system scopes only)."
      • changedInput schema / required
        Previous value: -[
        -  "server_version",
        -  "model_id",
        -  "framework_id"
        -]New value: +[
        +  "server_version",
        +  "scope",
        +  "scope_id",
        +  "framework_id"
        +]
    • Removedget_component
    • Removedget_control
    • Removedget_control_objective
    • Changedget_control_objectives3 fields changed
      • addedInput schema / properties / co_id
        Added value: +{
        +  "anyOf": [
        +    {
        +      "type": "string"
        +    },
        +    {
        +      "type": "null"
        +    }
        +  ],
        +  "default": null,
        +  "description": "If set, single mode — return this one control objective\n(e.g. ``CO3``) with its verdict. If omitted, matrix mode."
        +}
      • changedInput schema / properties / limit / description
        Previous value: -"Max to return (0 = summary only, no per-CO records)."New value: +"Matrix mode — max to return (0 = summary only, no per-CO\nrecords)."
      • changedInput schema / properties / offset / description
        Previous value: -"Skip the first N control objectives."New value: +"Matrix mode — skip the first N control objectives."
    • Changedget_controls10 fields changed
      • changedInput schema / properties / co_id / description
        Previous value: -"Filter by control objective ID."New value: +"List-mode filter — control objective ID."
      • changedInput schema / properties / component_id / description
        Previous value: -"Filter by component ID (e.g., \"CMP1\")."New value: +"List-mode filter — component ID (e.g., \"CMP1\")."
      • changedInput schema / properties / control_id / description
        Previous value: -"Optional specific control id for detail mode."New value: +"If set, detail mode — return this one control's full\nrecord directly (e.g. ``CTL-12``). If omitted, list mode."
      • changedInput schema / properties / include_deleted / description
        Previous value: -"Include soft-deleted controls (default False)."New value: +"List mode — include soft-deleted controls\n(default False)."
      • changedInput schema / properties / include_orphaned / description
        Previous value: -"Include controls mapped only to tombstoned COs\n(default False)."New value: +"List mode — include controls mapped only to\ntombstoned COs (default False)."
      • changedInput schema / properties / limit / description
        Previous value: -"Max controls to return (0 = all)."New value: +"List mode — max controls to return (0 = all)."
      • changedInput schema / properties / offset / description
        Previous value: -"Skip the first N controls (pagination)."New value: +"List mode — skip the first N controls (pagination)."
      • changedInput schema / properties / status / description
        Previous value: -"Filter by \"implemented\", \"not_implemented\", or \"verified\"."New value: +"List-mode filter — \"implemented\", \"not_implemented\", or\n\"verified\"."
      • changedInput schema / properties / summary_only / description
        Previous value: -"If True, returns only id, description, status,\nassertion_count, and assumed_by per control (much smaller\nresponse)."New value: +"List mode — if True, returns only id, description,\nstatus, assertion_count, and assumed_by per control (much\nsmaller response)."
      • addedInput schema / properties / version
        Added value: +{
        +  "default": 0,
        +  "description": "Detail mode only — model version to read the control\nfrom. 0 (default) uses the latest. Ignored in list mode.",
        +  "type": "integer"
        +}
    • Addedget_entity
    • Removedget_functional_objective
    • Addedget_functional_objectives
    • Removedget_functional_scan_prompt
    • Addedget_group
    • Removedget_model_risk_view
    • Removedget_reach_verdicts
    • Changedget_reachability_verdicts5 fields changed
      • changedInput schema / properties / co_id / description
        Previous value: -"Optional CO id. When set, returns a single verdict;\n404 if the CO doesn't exist or is tombstoned."New value: +"FLAT mode only. Optional CO id — when set, returns a single\nverdict; 404 if the CO doesn't exist or is tombstoned. Ignored\nwhen ``composed=True``."
      • addedInput schema / properties / composed
        Added value: +{
        +  "default": false,
        +  "description": "When False (default), derive over this model's own\ntopology (flat). When True, derive over the composed effective\ntree (own ⊕ inherited).",
        +  "type": "boolean"
        +}
      • addedInput schema / properties / kind_filter
        Added value: +{
        +  "anyOf": [
        +    {
        +      "type": "string"
        +    },
        +    {
        +      "type": "null"
        +    }
        +  ],
        +  "default": null,
        +  "description": "COMPOSED mode only. Restrict verdicts to one kind —\none of ``\"reachable\" | \"unreachable\" | \"indeterminate\"``. Named\n``kind_filter`` (not ``kind``) to disambiguate from the verdict\nobject's own ``kind`` field. When omitted, all verdict kinds are\nreturned. Ignored when ``composed=False``."
        +}
      • addedInput schema / properties / page
        Added value: +{
        +  "default": 1,
        +  "description": "COMPOSED mode only. 1-indexed page number (default ``1``).\nIgnored when ``composed=False``.",
        +  "type": "integer"
        +}
      • addedInput schema / properties / page_size
        Added value: +{
        +  "default": 100,
        +  "description": "COMPOSED mode only. Verdicts per page (default ``100``).\nIgnored when ``composed=False``.",
        +  "type": "integer"
        +}
    • Removedget_recompute_quote
    • Addedget_risk_view
    • Changedget_scan_prompt2 fields changed
      • changedInput schema / properties / control_id / description
        Previous value: -"Optional single control to scope the prompt to. Empty (default) returns prompts for all not-yet-implemented controls."New value: +"Security kind only — optional single control to scope\nthe prompt to. Empty (default) returns prompts for all\nnot-yet-implemented controls. Ignored when kind=\"functional\"."
      • addedInput schema / properties / kind
        Added value: +{
        +  "default": "security",
        +  "description": "\"security\" (default) or \"functional\" — which scan brief.",
        +  "type": "string"
        +}
    • Removedget_system
    • Removedget_system_compliance_report
    • Removedget_system_risk_view
    • Removedget_tag_compliance_report
    • Removedget_tag_risk_view
    • Removedget_trust_boundary
    • Changedimport_threat_model_archive1 field changed
      • changedInput schema / properties / envelope / description
        Previous value: -"The full archive dict returned by\n``export_threat_model_archive``."New value: +"The full archive dict returned by\n``export_report (scope=\"model\", format=\"archive\")``."
    • Removedlink_dependency
    • Addedlink_system_dependency
    • Removedlist_functional_objectives
    • Addedlist_groups
    • Addedlist_model_groups
    • Removedlist_model_tags
    • Changedlist_reconciliation_candidates4 fields changed
      • addedInput schema / properties / disposition
        Added value: +{
        +  "default": "active",
        +  "description": "Which side of the queue to read — ``\"active\"``\n(default, open candidates) or ``\"rejected\"`` (persisted\nnot-a-duplicate decisions).",
        +  "type": "string"
        +}
      • changedInput schema / properties / model_id / description
        Previous value: -"ID of the threat model."New value: +"ID of the descendant threat model."
      • changedInput schema / properties / page / description
        Previous value: -"1-indexed page number. Default 1."New value: +"ACTIVE disposition only. 1-indexed page number. Default 1.\nIgnored when ``disposition=\"rejected\"``."
      • changedInput schema / properties / page_size / description
        Previous value: -"Items per page. Default 50."New value: +"ACTIVE disposition only. Items per page. Default 50.\nIgnored when ``disposition=\"rejected\"``."
    • Removedlist_reconciliation_rejections
    • Removedlist_systems
    • Removedlist_tags
    • Addedpreview_undo_composition
    • Removedpreview_undo_lift_composition
    • Removedpreview_undo_split_composition
    • Changedrecompute_verdicts2 fields changed
      • addedInput schema / properties / dry_run
        Added value: +{
        +  "default": false,
        +  "description": "When True, return only the pre-flight estimate and enqueue\nnothing. When False (default), enqueue the recompute.",
        +  "type": "boolean"
        +}
      • changedInput schema / properties / model_id / description
        Previous value: -"ID of the threat model to re-evaluate."New value: +"ID of the threat model to re-evaluate (or estimate for)."
    • Removedremove_asset
    • Removedremove_assumption
    • Removedremove_attacker
    • Removedremove_component
    • Addedremove_entity
    • Addedremove_model_from_group
    • Removedremove_model_from_tag
    • Removedremove_trust_boundary
    • Removedrestore_asset
    • Removedrestore_assumption
    • Removedrestore_attacker
    • Addedrestore_entity
    • Addedrevalidate_entity_quality
    • Removedrevalidate_threat_model_entities
    • Changedselect_compliance_frameworks5 fields changed
      • changedInput schema / properties / framework_ids / description
        Previous value: -"Comma-separated framework IDs (e.g. \"asvs-4.0,nist-csf\")."New value: +"comma-separated framework ids (e.g. \"asvs-4.0,nist-csf\")."
      • removedInput schema / properties / model_id
        Removed value: -{
        -  "description": "ID of the threat model.",
        -  "type": "string"
        -}
      • addedInput schema / properties / scope
        Added value: +{
        +  "description": "target boundary — \"model\", \"system\", or \"tag\".",
        +  "enum": [
        +    "model",
        +    "system",
        +    "tag"
        +  ],
        +  "type": "string"
        +}
      • addedInput schema / properties / scope_id
        Added value: +{
        +  "description": "id of the model, system, or tag selected by ``scope``.",
        +  "type": "string"
        +}
      • changedInput schema / required
        Previous value: -[
        -  "server_version",
        -  "model_id",
        -  "framework_ids"
        -]New value: +[
        +  "server_version",
        +  "scope",
        +  "scope_id",
        +  "framework_ids"
        +]
    • Removedselect_system_compliance_frameworks
    • Removedselect_tag_compliance_frameworks
    • Removedset_co_cal
    • Addedset_control_objective_cal
    • Addedsubmit_functional_test_assertions
    • Removedsubmit_functional_tests
    • Removedunassume_control
    • Addedundo_composition_event
    • Removedundo_lift_composition_event
    • Removedundo_split_composition_event
  11. 1 tool updatev0.67.0
    • Addedcreate_risk_acceptance
  12. 62 tool updatesv0.66.0
    • Addedaccept_coverage_divergences
    • Changedadd_evidence3 fields changed
      • changedInput schema / properties / label / description
        Previous value: -"Description of evidence (required)."New value: +"Human-readable description of the evidence (required)."
      • changedInput schema / properties / type / description
        Previous value: -"Evidence type: \"code\", \"test\", \"config\", \"document\", \"link\"."New value: +"Evidence type — one of \"code\", \"test\", \"config\",\n\"document\", \"link\" (default \"code\")."
      • changedInput schema / properties / url / description
        Previous value: -"Optional file path or URL."New value: +"Optional file path or URL pointing at the artifact."
    • Changedadd_functional_test2 fields changed
      • changedInput schema / properties / functional_objective_ids / description
        Previous value: -"Comma-separated objective ids the test satisfies."New value: +"Comma-separated objective ids the test satisfies (at least one required; get them from list_functional_objectives)."
      • changedInput schema / properties / status / description
        Previous value: -"not_implemented | implemented | verified (an operator claim;\nan independent CI run is what actually verifies it)."New value: +"not_implemented | implemented | verified — an operator claim only; an independent CI run is what actually verifies the test. Defaults to not_implemented."
    • Changedapply_certain_reconciliation_match1 field changed
      • changedInput schema / properties / model_id / description
        Previous value: -"ID of the descendant threat model the duplicate is\non."New value: +"ID of the descendant threat model the duplicate is on."
    • Addedapply_control_changeset
    • Changedassess_model4 fields changed
      • changedInput schema / properties / limit / description
        Previous value: -"Max to return (0=all)."New value: +"Max control objectives to return (0 = all)."
      • changedInput schema / properties / offset / description
        Previous value: -"Skip first N."New value: +"Skip the first N control objectives."
      • changedInput schema / properties / status / description
        Previous value: -"Filter: \"mitigated\", \"at_risk\", \"unassessed\"."New value: +"Optional filter — \"mitigated\", \"at_risk\", or \"unassessed\"."
      • changedInput schema / properties / summary_only / description
        Previous value: -"If True, returns only summary counts (no per-CO details)."New value: +"If True, return only summary counts (no per-CO details)."
    • Addedattach_foundation
    • Changedauto_map_controls1 field changed
      • changedInput schema / properties / control_id / description
        Previous value: -"Optional specific control to map."New value: +"Optional single control ID to map; omit to map all of the model's controls."
    • Changedcheck_functional_gaps1 field changed
      • addedInput schema / properties / model_id / description
        Added value: +"ID of the threat model to analyse for functional gaps."
    • Changedcomplete_setup_step1 field changed
      • changedInput schema / properties / step_id / description
        Previous value: -"One of: mcp_configured, mipiti_verify_installed,\nci_secret_added, ci_pipeline_added."New value: +"The step to mark complete, one of \"mcp_configured\", \"mipiti_verify_installed\", \"ci_secret_added\", \"ci_pipeline_added\"."
    • Addedconfirm_reliance
    • Changedcreate_system1 field changed
      • changedInput schema / properties / name / description
        Previous value: -"System name (e.g., \"Mobile Banking Platform\")."New value: +"System name (e.g. \"Mobile Banking Platform\")."
    • Changeddeclare_foundation2 fields changed
      • changedInput schema / properties / provides / description
        Previous value: -"List of advertised-control dicts (control_id required)."New value: +"List of advertised-control dicts. ``control_id`` is required per entry; ``capability_label`` and ``description`` describe what the control provides to consumers."
      • changedInput schema / properties / visibility / description
        Previous value: -"\"workspace\" or \"explicit\"."New value: +"Who may delegate to this foundation. \"workspace\" (default) makes it discoverable to every model in the workspace; \"explicit\" limits it to models explicitly attached."
    • Changeddelete_assertion2 fields changed
      • changedInput schema / properties / assumption_id / description
        Previous value: -"ID of the assumption (omit if using control_id)."New value: +"ID of the assumption the assertion belongs to (omit if it belongs to a control)."
      • changedInput schema / properties / control_id / description
        Previous value: -"ID of the control (omit if using assumption_id)."New value: +"ID of the control the assertion belongs to (omit if it belongs to an assumption)."
    • Changeddelete_control1 field changed
      • changedInput schema / properties / reason / description
        Previous value: -"Justification for deletion."New value: +"Optional justification recorded in the audit trail\n(recommended)."
    • Addeddelete_reliance
    • Addeddismiss_verdict_divergences
    • Changededit_assumption10 fields changed
      • changedInput schema / properties / assumption_id / description
        Previous value: -"ID of the assumption (e.g., \"AS1\")."New value: +"ID of the assumption to edit (e.g., \"AS1\")."
      • changedInput schema / properties / clear_exclusion / description
        Previous value: -"When True, clears the predicate. Mutually\nexclusive with the exclusion_* params (those win if both\nare sent)."New value: +"When True, removes the predicate entirely (the\nassumption becomes prose-only). Mutually exclusive with the\nexclusion_* params — if both are sent, the exclusion_* params win."
      • changedInput schema / properties / description / description
        Previous value: -"New description."New value: +"New description (omit to leave unchanged)."
      • addedInput schema / properties / exclusion_asset_component_id / description
        Added value: +"\"*\" or a concrete component ID."
      • addedInput schema / properties / exclusion_asset_id / description
        Added value: +"\"*\" or a concrete asset ID."
      • addedInput schema / properties / exclusion_attacker_id / description
        Added value: +"Predicate match — \"*\" wildcard or a concrete\nattacker ID."
      • addedInput schema / properties / exclusion_attacker_vector / description
        Added value: +"One of \"Network\" | \"Adjacent\" | \"Local\" |\n\"Physical\" | \"*\"."
      • addedInput schema / properties / exclusion_co_ids / description
        Added value: +"Comma-separated CO IDs the predicate matches\nexplicitly; when non-empty, overrides the match fields. Supplying\nany exclusion_* param rewrites the whole predicate (unspecified\nfields default to \"*\")."
      • addedInput schema / properties / exclusion_property_match / description
        Added value: +"\"C\" | \"I\" | \"A\" | \"U\" | \"*\"."
      • changedInput schema / properties / linked_co_ids / description
        Previous value: -"New comma-separated CO IDs (replaces existing linkage)."New value: +"New comma-separated CO IDs; replaces the existing\nlinkage (omit to leave unchanged)."
    • Changedget_capability2 fields changed
      • addedInput schema / properties / capability_id / description
        Added value: +"ID of the capability to fetch."
      • addedInput schema / properties / model_id / description
        Added value: +"ID of the threat model the capability belongs to."
    • Changedget_compliance_report5 fields changed
      • changedInput schema / properties / framework_id / description
        Previous value: -"ID of the compliance framework."New value: +"ID of the compliance framework (as listed by list_compliance_frameworks)."
      • changedInput schema / properties / level / description
        Previous value: -"Optional level filter (e.g., 1 for L1 only)."New value: +"Optional level filter for level-aware frameworks — returns only requirements at or below this level (e.g. 1 for L1 only). Omit for all levels."
      • changedInput schema / properties / limit / description
        Previous value: -"Max to return."New value: +"Maximum requirement rows to return. Default 0 = no explicit limit."
      • changedInput schema / properties / offset / description
        Previous value: -"Skip first N."New value: +"Number of requirement rows to skip, for pagination. Default 0."
      • changedInput schema / properties / status / description
        Previous value: -"Filter: \"covered\", \"partial\", \"uncovered\", \"unmapped\", \"excluded\"."New value: +"Optional status filter: \"covered\", \"partial\", \"uncovered\", \"unmapped\", or \"excluded\". Empty = all statuses."
    • Changedget_control_generation_status1 field changed
      • addedInput schema / properties / model_id / description
        Added value: +"ID of the threat model whose control-generation status to poll."
    • Changedget_control_objectives2 fields changed
      • changedInput schema / properties / limit / description
        Previous value: -"Max to return (0=summary only)."New value: +"Max to return (0 = summary only, no per-CO records)."
      • changedInput schema / properties / offset / description
        Previous value: -"Skip first N."New value: +"Skip the first N control objectives."
    • Changedget_controls7 fields changed
      • changedInput schema / properties / control_id / description
        Previous value: -"Optional specific control for detail mode."New value: +"Optional specific control id for detail mode."
      • changedInput schema / properties / include_deleted / description
        Previous value: -"Include soft-deleted controls."New value: +"Include soft-deleted controls (default False)."
      • changedInput schema / properties / include_orphaned / description
        Previous value: -"Include controls mapped only to tombstoned\nCOs (default False)."New value: +"Include controls mapped only to tombstoned COs\n(default False)."
      • changedInput schema / properties / limit / description
        Previous value: -"Max to return (0=all)."New value: +"Max controls to return (0 = all)."
      • changedInput schema / properties / offset / description
        Previous value: -"Skip first N (for pagination)."New value: +"Skip the first N controls (pagination)."
      • changedInput schema / properties / status / description
        Previous value: -"Filter by \"implemented\", \"not_implemented\", \"verified\"."New value: +"Filter by \"implemented\", \"not_implemented\", or \"verified\"."
      • changedInput schema / properties / summary_only / description
        Previous value: -"If True, returns only id, description, status,\nassertion_count, and assumed_by per control (much smaller response)."New value: +"If True, returns only id, description, status,\nassertion_count, and assumed_by per control (much smaller\nresponse)."
    • Changedget_effective_coverage1 field changed
      • changedInput schema / properties / origin / description
        Previous value: -"filter coverage rows by contributing-control origin —\none of ``\"own\" | \"cross\" | \"inherited\"``. When omitted,\nrows with any origin mix are returned."New value: +"filter coverage rows by contributing-control origin — one\nof ``\"own\" | \"cross\" | \"inherited\"``. When omitted, rows with\nany origin mix are returned."
    • Changedget_functional_coverage1 field changed
      • addedInput schema / properties / model_id / description
        Added value: +"ID of the threat model whose functional coverage to report."
    • Changedget_functional_objective2 fields changed
      • addedInput schema / properties / functional_objective_id / description
        Added value: +"ID of the functional objective to fetch."
      • addedInput schema / properties / model_id / description
        Added value: +"ID of the threat model the objective belongs to."
    • Changedget_functional_scan_prompt1 field changed
      • addedInput schema / properties / model_id / description
        Added value: +"ID of the threat model to build the functional brief for."
    • Changedget_scan_prompt1 field changed
      • changedInput schema / properties / control_id / description
        Previous value: -"Optional specific control ID."New value: +"Optional single control to scope the prompt to. Empty (default) returns prompts for all not-yet-implemented controls."
    • Changedget_system_compliance_report6 fields changed
      • changedInput schema / properties / framework_id / description
        Previous value: -"ID of the compliance framework."New value: +"ID of a framework already selected for this system."
      • changedInput schema / properties / level / description
        Previous value: -"Optional level filter."New value: +"Optional framework level/tier filter (e.g., baseline level number). Omit for all levels."
      • changedInput schema / properties / limit / description
        Previous value: -"Max to return."New value: +"Max requirement rows to return; 0 (default) returns all."
      • changedInput schema / properties / offset / description
        Previous value: -"Skip first N."New value: +"Skip the first N requirement rows (pagination). Default 0."
      • changedInput schema / properties / status / description
        Previous value: -"Filter: \"covered\", \"partial\", \"uncovered\", \"unmapped\", \"excluded\"."New value: +"Optional per-requirement filter, one of \"covered\", \"partial\", \"uncovered\", \"unmapped\", \"excluded\". Empty (default) returns all."
      • changedInput schema / properties / system_id / description
        Previous value: -"ID of the system."New value: +"ID of the system to report on."
    • Changedget_tag_compliance_report1 field changed
      • changedInput schema / properties / level / description
        Previous value: -"optional framework level filter (0 = all)."New value: +"optional framework level filter; 0 (default) reports all levels."
    • Addedget_verdict_divergence
    • Changedimport_functional_tests1 field changed
      • changedInput schema / properties / tests_json / description
        Previous value: -"A JSON array of test objects. Each object supports\n``test_name``, ``file_path``, ``framework``, ``description``,\n``status`` (not_implemented | implemented | verified — an operator\nclaim; an independent CI run is what verifies it), and\n``functional_objective_ids`` (list of objective ids the test covers).\nAt least ``test_name`` or ``description`` is required per test; the\nrest are optional."New value: +"A JSON array of test objects. Each object supports ``test_name``, ``file_path``, ``framework``, ``description``, ``status`` (not_implemented | implemented | verified — an operator claim; an independent CI run is what verifies it), and ``functional_objective_ids`` (list of objective ids the test covers). At least ``test_name`` or ``description`` is required per test; the rest are optional."
    • Changedimport_threat_model_archive1 field changed
      • changedInput schema / properties / envelope / description
        Previous value: -"The full archive dict returned by\n`export_threat_model_archive`."New value: +"The full archive dict returned by\n``export_threat_model_archive``."
    • Changedlift_composition_entity4 fields changed
      • changedInput schema / properties / acknowledged_third_party_subtrees / description
        Previous value: -"Optional list of subtree\nroots the operator has acknowledged as in-scope for the\nlift."New value: +"Optional list of subtree roots\nthe operator has acknowledged as in-scope for the lift."
      • changedInput schema / properties / attached_state_resolutions / description
        Previous value: -"Optional per-state-key resolution\nmap (e.g. ``{\"state:assertions/AS3\": \"keep_b\"}``)."New value: +"Optional per-state-key resolution map\n(e.g. ``{\"state:assertions/AS3\": \"keep_b\"}``)."
      • changedInput schema / properties / lca_descendant_ids / description
        Previous value: -"Optional snapshot of the LCA's descendant\nset used by the over-application gate. Omit to let the\nserver compute it via BFS."New value: +"Optional snapshot of the LCA's descendant set\nused by the over-application gate. Omit to let the server\ncompute it via BFS."
      • changedInput schema / properties / lca_model_id / description
        Previous value: -"Target ancestor model id (the LCA, or any\nancestor higher up the chain)."New value: +"Target ancestor model id (the LCA, or any ancestor\nhigher up the chain)."
    • Changedlist_attestations1 field changed
      • changedInput schema / properties / assumption_id / description
        Previous value: -"ID of the assumption."New value: +"ID of the assumption whose attestation history to list."
    • Changedlist_capabilities1 field changed
      • addedInput schema / properties / model_id / description
        Added value: +"ID of the threat model whose capabilities to list."
    • Changedlist_findings2 fields changed
      • changedInput schema / properties / control_id / description
        Previous value: -"Optional filter by control ID."New value: +"Optional filter to findings on one control. Empty (default) returns findings for all controls."
      • changedInput schema / properties / status / description
        Previous value: -"Optional filter: \"discovered\", \"acknowledged\", \"remediated\",\n\"verified\", \"dismissed\"."New value: +"Optional lifecycle filter, one of \"discovered\", \"acknowledged\", \"remediated\", \"verified\", \"dismissed\". Empty (default) returns all statuses."
    • Changedlist_functional_objectives1 field changed
      • addedInput schema / properties / model_id / description
        Added value: +"ID of the threat model whose functional objectives to list."
    • Addedlist_reliance
    • Addedlist_tags
    • Changedlist_threat_models2 fields changed
      • changedInput schema / properties / include_assessment_summary / description
        Previous value: -"If True, include an `assessment_summary`\nobject with each model (counts of mitigated / at_risk /\nunassessed COs plus a human-readable `message`). Useful for\naggregate posture queries across the workspace in a single\ncall — e.g. \"which of my models are at risk?\" — instead of\ncalling `assess_model` once per model (N+1 at the agent layer).\nAdds ~100 bytes per model to the response."New value: +"If True, include an `assessment_summary` object per model (counts of mitigated / at_risk / unassessed control objectives plus a human-readable `message`). Use for aggregate posture queries across the workspace in a single call (e.g. \"which of my models are at risk?\") instead of calling `assess_model` once per model. Adds roughly 100 bytes per model. Default False."
      • changedInput schema / properties / source / description
        Previous value: -"Filter by source system. One of \"web\", \"mcp\", \"jira\", \"api\".\nOmit to list all models regardless of source."New value: +"Filter by the system that created each model. One of \"web\", \"mcp\", \"jira\", \"api\". Omit (default \"\") to list all models regardless of source."
    • Changedmap_control_to_requirement4 fields changed
      • changedInput schema / properties / confidence / description
        Previous value: -"Mapping confidence: \"llm\", \"manual\", \"verified\"."New value: +"Provenance label recorded on the mapping: \"manual\" (default, operator-asserted), \"llm\" (machine-suggested), or \"verified\" (human-confirmed)."
      • changedInput schema / properties / control_id / description
        Previous value: -"ID of the control (e.g., \"CTRL-01\")."New value: +"ID of the control to map (e.g. \"CTRL-01\")."
      • changedInput schema / properties / notes / description
        Previous value: -"Optional notes about mapping."New value: +"Optional free-text note explaining the mapping rationale."
      • changedInput schema / properties / requirement_id / description
        Previous value: -"ID of the requirement (e.g., \"V2.1.1\")."New value: +"ID of the requirement to map to (e.g. \"V2.1.1\")."
    • Changedpreview_undo_lift_composition2 fields changed
      • changedInput schema / properties / lift_id / description
        Previous value: -"Either the surrogate id of the ``lift_applied``\nactivity event, or the structured ``lift_id`` carried in\nthe event's payload — both lookups are supported."New value: +"Either the surrogate id of the ``lift_applied`` activity\nevent, or the structured ``lift_id`` carried in the event's\npayload — both lookups are supported."
      • changedInput schema / properties / model_id / description
        Previous value: -"The model whose composition view originated the\nlift. Must match the ``threat_model_id`` carried by the\ncited activity event; the server rejects with 404 when a\ncaller tries to undo a sibling model's lift through a\ndifferent model's URL."New value: +"The model whose composition view originated the lift.\nMust match the ``threat_model_id`` carried by the cited\nactivity event; the server rejects with 404 when a caller\ntries to undo a sibling model's lift through a different\nmodel's URL."
    • Changedquery_threat_model1 field changed
      • changedInput schema / properties / question / description
        Previous value: -"The question to ask."New value: +"The natural-language question to ask."
    • Addedrefine_threat_model
    • Changedregenerate_controls3 fields changed
      • changedInput schema / properties / batch_size / description
        Previous value: -"COs per batch in batch mode (default: 15). Smaller\n= more accurate + granular progress, more LLM calls."New value: +"COs per batch in batch mode (default 15). Smaller =\nmore accurate and more granular progress, but more LLM calls."
      • changedInput schema / properties / co_ids / description
        Previous value: -"Optional comma-separated CO IDs to regenerate (e.g.\n\"CO1,CO5\"). When omitted, regenerates all controls."New value: +"Optional comma-separated CO IDs to regenerate (e.g.\n\"CO1,CO5\"). Omit to regenerate all controls."
      • changedInput schema / properties / mode / description
        Previous value: -"\"batch\" (default) or \"per_co\" (most thorough, one LLM\ncall per CO)."New value: +"\"batch\" (default) or \"per_co\" (most thorough — one LLM call\nper CO)."
    • Changedremove_attacker1 field changed
      • changedInput schema / properties / attacker_id / description
        Previous value: -"ID of the attacker to soft-delete."New value: +"ID of the attacker to soft-delete (e.g. \"T1\")."
    • Changedremove_component1 field changed
      • changedInput schema / properties / component_id / description
        Previous value: -"ID of the component to remove."New value: +"ID of the component to remove (e.g. \"CMP1\")."
    • Changedremove_evidence1 field changed
      • changedInput schema / properties / evidence_index / description
        Previous value: -"Zero-based index to remove."New value: +"Zero-based position of the item to remove within\nthe control's ``evidence`` array (default 0 = first item)."
    • Addedremove_model_from_tag
    • Changedremove_trust_boundary1 field changed
      • changedInput schema / properties / tb_id / description
        Previous value: -"ID of the trust boundary to remove."New value: +"ID of the trust boundary to remove (e.g., \"TB1\")."
    • Addedretry_verdicts
    • Changedrevalidate_threat_model_entities1 field changed
      • addedInput schema / properties / model_id / description
        Added value: +"ID of the threat model whose assets and attackers to\nre-validate."
    • Addedselect_tag_compliance_frameworks
    • Changedset_functional_satisfaction_groups2 fields changed
      • changedInput schema / properties / groups_json / description
        Previous value: -"A JSON object mapping group label to a list of functional\ntest ids, e.g. ``{\"1\": [\"FT-1\", \"FT-2\"], \"2\": [\"FT-3\"]}``."New value: +"JSON object mapping group label to a list of functional test ids, e.g. ``{\"1\": [\"FT-1\", \"FT-2\"], \"2\": [\"FT-3\"]}``. Pass ``{}`` to clear all groups."
      • changedInput schema / properties / ungrouped / description
        Previous value: -"Comma-separated functional-test ids to keep unassigned to any\ngroup (optional)."New value: +"Comma-separated functional-test ids to keep associated with the objective but unassigned to any group (optional)."
    • Changedset_mitigation_groups2 fields changed
      • changedInput schema / properties / defense_in_depth / description
        Previous value: -"Comma-separated control IDs for defense-in-depth.\nExample: \"CTRL-04,CTRL-05\""New value: +"Comma-separated control IDs tracked as\ndefense-in-depth (not required for mitigation). Example:\n\"CTRL-04,CTRL-05\"."
      • changedInput schema / properties / groups / description
        Previous value: -"JSON object mapping group numbers to control ID lists.\nExample: '{\"1\": [\"CTRL-01\", \"CTRL-02\"], \"2\": [\"CTRL-03\"]}'"New value: +"JSON object mapping group numbers to control-ID lists.\nExample: '{\"1\": [\"CTRL-01\", \"CTRL-02\"], \"2\": [\"CTRL-03\"]}'."
    • Changedsubmit_findings1 field changed
      • changedInput schema / properties / findings_json / description
        Previous value: -"JSON array of finding objects with control_id, title,\ndescription, severity, checked_locations, checked_patterns,\nexpected_evidence."New value: +"JSON string of an **array** of finding objects. Each object should carry:\n- ``control_id`` (str): the control the gap relates to.\n- ``title`` (str): short summary of the gap.\n- ``description`` (str): what is missing and why it matters.\n- ``severity`` (str): finding severity (e.g., \"low\"/\"medium\"/\"high\"/\"critical\").\n- ``checked_locations`` (list): files/paths inspected.\n- ``checked_patterns`` (list): patterns/signals searched for.\n- ``expected_evidence`` (str): what implemented evidence would have looked like.\nMust parse as a JSON array; a single object or malformed JSON is rejected."
    • Changedsubmit_functional_tests1 field changed
      • changedInput schema / properties / assertions_json / description
        Previous value: -"JSON array of assertions, each\n{\"type\": \"test_passes\"|..., \"params\": {...}, \"description\": \"...\",\n \"repo\": \"<owner>/<repo>\"}. Each assertion must carry an explicit\nrepo (or the \"no_repo\" sentinel)."New value: +"JSON array of assertion objects, each {\"type\": \"test_passes\" | \"test_exists\" | ..., \"params\": {...}, \"description\": \"...\", \"repo\": \"<owner>/<repo>\"}. Every assertion must carry an explicit repo, or the \"no_repo\" sentinel when the check is not tied to a repository."
    • Changedunassume_control1 field changed
      • changedInput schema / properties / control_id / description
        Previous value: -"ID of the control."New value: +"ID of the control (e.g., \"CTRL-03\")."
    • Changedunreject_reconciliation_candidate1 field changed
      • changedInput schema / properties / model_id / description
        Previous value: -"ID of the descendant threat model the rejection is\non."New value: +"ID of the descendant threat model the rejection is on."
    • Addedupdate_control_status
    • Changedupdate_finding5 fields changed
      • changedInput schema / properties / finding_id / description
        Previous value: -"ID of the finding."New value: +"ID of the finding to update."
      • changedInput schema / properties / notes / description
        Previous value: -"Optional notes."New value: +"Optional free-text notes recorded on the finding."
      • changedInput schema / properties / reason / description
        Previous value: -"Optional reason (required for dismissal)."New value: +"Optional rationale; required when dismissing (status=\"dismissed\")."
      • changedInput schema / properties / remediation_assertion_ids / description
        Previous value: -"Comma-separated assertion IDs linking fix."New value: +"Optional comma-separated assertion IDs that evidence the fix, linking the remediation to the assertions that prove it. Empty by default."
      • changedInput schema / properties / status / description
        Previous value: -"New status."New value: +"New lifecycle status, one of \"discovered\", \"acknowledged\", \"remediated\", \"verified\", \"dismissed\"."
  13. 36 tool updatesv0.62.2
    • Addedadd_attacker
    • Addedapply_certain_reconciliation_match
    • Removedconfirm_reliance
    • Addedcreate_reliance
    • Addeddeclare_foundation
    • Addedexport_threat_model
    • Addedgenerate_threat_model
    • Addedget_assumption
    • Addedget_attacker
    • Addedget_component
    • Addedget_composition_overview
    • Addedget_control_generation_status
    • Addedget_control_objective
    • Addedget_controls
    • Addedget_effective_coverage
    • Addedget_tag_compliance_report
    • Addedget_tag_risk_view
    • Addedget_trust_boundary
    • Addedimport_threat_model_archive
    • Addedlist_model_tags
    • Addedlist_reconciliation_candidates
    • Addedlist_reconciliation_rejections
    • Removedlist_reliance
    • Removedlist_tags
    • Addedlist_workspaces
    • Addedmodel_coherence_report
    • Addedpreview_undo_lift_composition
    • Addedquery_threat_model
    • Addedremap_control
    • Addedremove_asset
    • Removedremove_model_from_tag
    • Addedrename_threat_model
    • Removedselect_tag_compliance_frameworks
    • Addedsplit_composition_entity
    • Addedundo_split_composition_event
    • Removedupdate_control_status
  14. 114 tool updatesv0.62.2
    • Addedadd_asset
    • Addedadd_assumption
    • Addedadd_component
    • Addedadd_evidence
    • Addedadd_functional_test
    • Addedadd_model_to_system
    • Addedadd_model_to_tag
    • Addedapply_finding_remediation
    • Addedassess_model
    • Addedassociate_functional_test
    • Addedassume_control
    • Removedattach_foundation
    • Addedauto_map_controls
    • Addedauto_remediate
    • Addedcheck_control_gaps
    • Addedcheck_functional_gaps
    • Addedclassify_model_cwe
    • Addedcomplete_setup_step
    • Addedconvert_assumption_to_controls
    • Addedcreate_system
    • Addedcreate_tag
    • Removeddeclare_foundation
    • Addeddelete_assertion
    • Addeddelete_control
    • Removeddelete_reliance
    • Addeddelete_tag
    • Addededit_asset
    • Addededit_assumption
    • Addededit_attacker
    • Addededit_component
    • Addededit_trust_boundary
    • Addedexport_threat_model_archive
    • Addedgenerate_functional_objectives
    • Removedgenerate_threat_model
    • Addedget_asset
    • Addedget_capability
    • Addedget_compliance_report
    • Removedget_composition_overview
    • Addedget_control
    • Addedget_control_objectives
    • Addedget_cwe_catalog
    • Addedget_findings_risks
    • Addedget_functional_coverage
    • Addedget_functional_objective
    • Addedget_functional_satisfaction_groups
    • Addedget_functional_scan_prompt
    • Addedget_functional_test_sufficiency
    • Addedget_mitigation_groups
    • Addedget_model_cwe_tags
    • Addedget_model_risk_view
    • Addedget_reach_verdicts
    • Addedget_recompute_quote
    • Addedget_remediation_leverage
    • Addedget_review_queue
    • Addedget_scan_prompt
    • Addedget_setup_status
    • Addedget_sufficiency
    • Addedget_system_compliance_report
    • Addedget_system_dependencies
    • Addedget_system_risk_view
    • Removedget_tag_compliance_report
    • Removedget_tag_risk_view
    • Addedget_verification_report
    • Addedimport_compliance_framework
    • Addedimport_controls
    • Addedimport_functional_tests
    • Addedlink_dependency
    • Addedlist_assertions
    • Addedlist_attestations
    • Addedlist_capabilities
    • Addedlist_compliance_frameworks
    • Addedlist_effective_attack_paths
    • Addedlist_effective_control_objectives
    • Addedlist_findings
    • Addedlist_functional_objectives
    • Removedlist_model_tags
    • Addedlist_reliance
    • Addedlist_risk_acceptances
    • Addedlist_systems
    • Addedlist_tags
    • Removedmodel_coherence_report
    • Addedpreview_finding_remediation
    • Addedpreview_undo_split_composition
    • Removedquery_threat_model
    • Addedrecompute_verdicts
    • Addedreevaluate_threat_model_factors
    • Removedrefine_threat_model
    • Addedreject_reconciliation_candidate
    • Removedremap_control
    • Addedremove_assumption
    • Addedremove_evidence
    • Addedremove_trust_boundary
    • Removedrename_threat_model
    • Addedrestore_asset
    • Addedrestore_assumption
    • Addedrestore_attacker
    • Addedrevalidate_threat_model_entities
    • Addedselect_compliance_frameworks
    • Addedselect_system_compliance_frameworks
    • Addedset_co_cal
    • Addedset_control_assumption_groups
    • Addedset_functional_satisfaction_groups
    • Addedset_mitigation_groups
    • Removedsplit_composition_entity
    • Addedsubmit_assertions
    • Addedsubmit_attestation
    • Addedsubmit_findings
    • Addedsubmit_functional_tests
    • Addedsuggest_functional_test_mappings
    • Addedunassume_control
    • Addedundo_lift_composition_event
    • Addedunreject_reconciliation_candidate
    • Addedupdate_finding
    • Addedupdate_organization
  15. 59 tool updatesv0.62.1
    • Removedadd_asset
    • Removedadd_component
    • Removedadd_model_to_system
    • Removedadd_model_to_tag
    • Addedadd_trust_boundary
    • Removedapply_certain_reconciliation_match
    • Removedassess_model
    • Removedcheck_control_gaps
    • Removedcreate_reliance
    • Removedcreate_system
    • Removedcreate_tag
    • Removeddelete_control
    • Removeddelete_tag
    • Removededit_asset
    • Removededit_component
    • Removedexport_threat_model
    • Removedexport_threat_model_archive
    • Removedget_asset
    • Removedget_assumption
    • Removedget_attacker
    • Removedget_compliance_report
    • Removedget_component
    • Removedget_control
    • Addedget_control_assumption_groups
    • Removedget_control_generation_status
    • Removedget_control_objective
    • Removedget_control_objectives
    • Removedget_controls
    • Removedget_effective_coverage
    • Removedget_mitigation_groups
    • Removedget_reach_verdicts
    • Removedget_recompute_quote
    • Removedget_system_dependencies
    • Removedget_trust_boundary
    • Removedimport_compliance_framework
    • Removedimport_controls
    • Removedimport_threat_model_archive
    • Removedlink_dependency
    • Removedlist_compliance_frameworks
    • Removedlist_effective_attack_paths
    • Removedlist_effective_control_objectives
    • Removedlist_reconciliation_candidates
    • Removedlist_reconciliation_rejections
    • Removedlist_reliance
    • Removedlist_tags
    • Removedpreview_undo_lift_composition
    • Removedpreview_undo_split_composition
    • Removedrecompute_verdicts
    • Removedreevaluate_threat_model_factors
    • Removedreject_reconciliation_candidate
    • Removedremove_evidence
    • Removedrestore_attacker
    • Removedrevalidate_threat_model_entities
    • Removedselect_compliance_frameworks
    • Removedset_co_cal
    • Removedset_mitigation_groups
    • Removedundo_lift_composition_event
    • Removedundo_split_composition_event
    • Removedunreject_reconciliation_candidate
  16. 65 tool updatesv0.62.0
    • Removedadd_assumption
    • Removedadd_attacker
    • Removedadd_evidence
    • Removedadd_functional_test
    • Removedadd_trust_boundary
    • Removedapply_finding_remediation
    • Removedassociate_functional_test
    • Removedassume_control
    • Removedauto_map_controls
    • Removedauto_remediate
    • Removedcheck_functional_gaps
    • Removedclassify_model_cwe
    • Removedcomplete_setup_step
    • Removedconvert_assumption_to_controls
    • Removeddelete_assertion
    • Removededit_assumption
    • Removededit_attacker
    • Removededit_trust_boundary
    • Removedgenerate_functional_objectives
    • Removedget_capability
    • Removedget_control_assumption_groups
    • Addedget_control_generation_status
    • Removedget_cwe_catalog
    • Removedget_findings_risks
    • Removedget_functional_coverage
    • Removedget_functional_objective
    • Removedget_functional_satisfaction_groups
    • Removedget_functional_scan_prompt
    • Removedget_functional_test_sufficiency
    • Removedget_model_cwe_tags
    • Removedget_model_risk_view
    • Removedget_remediation_leverage
    • Removedget_review_queue
    • Removedget_scan_prompt
    • Removedget_setup_status
    • Removedget_sufficiency
    • Removedget_system_compliance_report
    • Removedget_system_risk_view
    • Removedget_verification_report
    • Removedimport_functional_tests
    • Removedlist_assertions
    • Removedlist_attestations
    • Removedlist_capabilities
    • Removedlist_findings
    • Removedlist_functional_objectives
    • Removedlist_risk_acceptances
    • Removedlist_systems
    • Removedlist_workspaces
    • Removedpreview_finding_remediation
    • Removedremove_asset
    • Removedremove_assumption
    • Removedremove_trust_boundary
    • Removedrestore_asset
    • Removedrestore_assumption
    • Removedselect_system_compliance_frameworks
    • Removedset_control_assumption_groups
    • Removedset_functional_satisfaction_groups
    • Removedsubmit_assertions
    • Removedsubmit_attestation
    • Removedsubmit_findings
    • Removedsubmit_functional_tests
    • Removedsuggest_functional_test_mappings
    • Removedunassume_control
    • Removedupdate_finding
    • Removedupdate_organization
  17. 7 tool updatesv0.60.1
    • Addedclassify_model_cwe
    • Changededit_attacker2 fields changed
      • addedInput schema / properties / attest_position
        Added value: +{
        +  "anyOf": [
        +    {
        +      "type": "boolean"
        +    },
        +    {
        +      "type": "null"
        +    }
        +  ],
        +  "default": null,
        +  "description": "Operator-attest the attacker's current position without\nchanging it — e.g. to confirm a fully external attacker's empty\ncrossed set so an objective blocked on an unpositioned attacker can\nbe resolved. Pass ``true`` to attest."
        +}
      • changedInput schema / properties / trust_boundary_ids / description
        Previous value: -"Comma-separated trust boundary IDs (replaces existing)."New value: +"Comma-separated trust boundary IDs — the boundaries\nthis attacker has crossed (its position). Replaces the existing set.\nChanging it operator-attests the position, which lets reachability\ntrust it for a decisive verdict."
    • Changededit_trust_boundary3 fields changed
      • changedInput schema / properties / change_reason / description
        Previous value: -"Required when ``passes`` or ``sealed`` actually changes.\nCaptured in the audit trail; documents why the boundary's vector\nfilter was tightened/widened or its isolation claim changed."New value: +"Required when ``passes``, ``sealed``, or the seal\nattestation actually changes. Captured in the audit trail; documents\nwhy the boundary's vector filter, isolation claim, or attestation\nchanged."
      • addedInput schema / properties / seal_source
        Added value: +{
        +  "anyOf": [
        +    {
        +      "type": "string"
        +    },
        +    {
        +      "type": "null"
        +    }
        +  ],
        +  "default": null,
        +  "description": "\"attested\" | \"unattested\". Only an operator-attested seal\nlets reachability decisively rule an objective unreachable past the\nboundary; an unattested (default/model-suggested) seal is treated as\npivotable. Use \"attested\" to attest a boundary already marked sealed\nwithout re-toggling it; \"unattested\" retracts. An attested seal\nimplies ``sealed``. Requires ``change_reason``."
        +}
      • changedInput schema / properties / sealed / description
        Previous value: -"New isolation flag. True declares NO lateral ingress (the only\nway in is crossing the perimeter — an air-gap / segmented enclave),\nwhich lets reachability decisively rule the boundary unreachable;\nFalse assumes a lateral pivot is possible. Reach-relevant — changing\nit can flip CO verdicts. Omit to leave unchanged."New value: +"New isolation flag. True declares NO lateral ingress (the only\nway in is crossing the perimeter — an air-gap / segmented enclave),\nwhich lets reachability decisively rule the boundary unreachable;\nFalse assumes a lateral pivot is possible. Reach-relevant — changing\nit can flip CO verdicts. Setting it records an operator attestation\nof the seal. Omit to leave unchanged."
    • Addedget_cwe_catalog
    • Addedget_model_cwe_tags
    • Addedget_recompute_quote
    • Addedrecompute_verdicts

TDQS

B3.2/5.0

Scored across 128 tools

Disambiguation3/5

Descriptions are unusually thorough and explicitly cross-reference the right sibling (e.g. judge_objective vs judge_objectives vs judge_imported_controls vs recompute_verdicts; refine_control vs remap_control vs regenerate_controls vs apply_control_changeset), which prevents most misselection. However, there are large overlapping clusters (start/pause/resume/discard control builds; get_controls/get_control_objectives/get_verification_report/get_sufficiency) and a real terminology collision where 'groups' means tags in list_groups/get_group/create_group while mitigation/assumption/satisfaction groups mean something else entirely. An agent can still pick correctly by reading carefully, but the boundaries are genuinely crowded.

Naming Consistency4/5

Nearly everything follows a predictable snake_case verb_noun pattern (add_asset, edit_attacker, remove_entity, list_findings, get_threat_model, submit_assertions). A few stative/odd names (assess_model, recompute_verdicts, judge_objective) and the decision to name tag tools 'group' while calling them tags in prose are minor deviations, but the style is internally consistent throughout.

Tool Count1/5

128 tools is an extreme count that far exceeds what an agent can reliably select from, and the surface contains visible redundancy (three judge_* tools, four control-build lifecycle tools, multiple verdict/report readers). Even granting the platform spans threat modeling, assurance, compliance, functional testing and composition, the number is a mismatch for practical tool routing.

Completeness5/5

Coverage is essentially exhaustive: full CRUD/lifecycle for models and every core entity (assets, attackers, components, trust boundaries, assumptions), control build/undo/versioning, assertions and sufficiency, findings, mitigation and satisfaction groups, risk acceptances and dispositions, compliance frameworks, functional tests, proposals/decisions, and composition lift/split. No obvious domain operation is missing, and import/export/archive round-trips close the loop.

Maintenance

ActivityActive
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • F
    license
    Not graded
    quality
    D
    maintenance
    Enables AI-powered threat modeling with tools for creating threat models, analyzing security threats, generating security controls, and validating architecture against best practices.
    -
  • A
    license
    A
    quality
    A
    maintenance
    Enables AI agents to interact with the SCF Controls Platform for security compliance, including browsing controls, tracking implementation, managing evidence, assessing risks, and monitoring vendors via natural language.
    196
    242 npm
    2
    MIT