Skip to main content
Glama

assess_ai_initiative

Read-onlyIdempotent

Evaluate an AI proposal in plain business language, request unresolved inputs once, and return an Accelerate, Fix, or Stop verdict with evidence gaps.

Instructions

The front door for one AI investment decision. CALL THIS FIRST when the user describes an AI idea in ordinary language or asks whether it should proceed. It resolves industry, revenue, business function, AI tier and organisational readiness, then returns one clarification covering every unresolved input or an Accelerate, Fix or Stop verdict. Ask that clarification once and call this tool again with the answers in the explicit fields. Use work_architecture to test whether the end-to-end workflow, affected roles, human decision rights and performance measures have been redesigned. A stated gap or missing work architecture evidence blocks Accelerate and stays visible in the audit trail. Pillar scores and work architecture evidence remain optional inputs, but unresolved values are never guessed and cannot unlock Accelerate. Use score_initiative when the canonical fields are already known, score_portfolio for several initiatives, and diagnose_process for measured waste in a running process. Deterministic calculation with no authentication. Anonymous usage telemetry may be sent; set AIBVF_TELEMETRY_DISABLE=1 to opt out.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
scoresNoOPTIONAL, and each pillar inside it is optional. The four AI BVF pillars, each an honest 0–100 self-assessment, combining deterministically into the verdict: governance_risk ≥ 70 OR financial_return ≤ 20 returns Stop; strategic_alignment, financial_return and change_enablement all ≥ 60 with governance_risk ≤ 40 returns Accelerate; everything else returns Fix. Pass ONLY the pillars the user has real evidence for — do NOT invent numbers for the rest. Missing pillars are estimated deterministically by the engine from disclosed AI BVF planning assumptions, the response reports which via pillar_basis and scores_used, decision score is haircut by how much was estimated, and a fully-estimated pass can never return Accelerate (it returns Fix pending confirmation). So call immediately with whatever the user gave you, then ask for evidence on the estimated pillars and re-call to firm the verdict up.
ai_tierNoOptional correction or answer: automation/RPA, GenAI/copilot, or agentic/autonomous. Overrides anything inferred from proposal.
functionNoOptional correction or answer in canonical or everyday language, for example customer service, procurement, finance or risk. Overrides anything inferred from proposal.
industryNoOptional correction or answer in canonical or everyday language, for example retail, hospital, bank or public sector. Overrides anything inferred from proposal.
proposalYesThe AI initiative in ordinary business language. Include the organisation, industry, approximate annual revenue, business function, AI ambition and how the organisation works today when known. The resolver extracts what it can and, when several inputs are missing, asks for all of them in one clarification; it never guesses an unresolved taxonomy value.
readinessNoOptional correction or answer: agile, traditional, or siloed, including everyday descriptions such as cross-functional, hierarchical or bureaucratic. Overrides anything inferred from proposal.
revenue_eurNoOptional approximate annual revenue in EUR. Overrides any EUR amount extracted from proposal. No currency conversion is performed.
work_architectureNoEvidence that workflows, roles, decision rights and measures are ready. Pass only what is known. Explicit gaps and omitted checks block Accelerate until all four checks are evidenced.
signal_completenessNoOptional 0 to 1 input-quality factor. The default ranges from 0.5 when all pillars are estimated to 1 when all are supplied. Supplied values still require evidence review. Lower this factor when the supplied pillars rest on weak evidence.

Output Schema

TableJSON Schema
NameRequiredDescriptionDefault
statusYesneeds_input when one or more required decision inputs remain unresolved; verdict when scoring completed.
verdictNoThe AI BVF score. Present only when status is verdict.
proposalYesThe supplied proposal, returned so the next call can preserve it verbatim.
bvf_versionYes
resolutionsYesEvery deterministic resolution, naming the field, canonical value, source and matched phrase.
suggestionsNoAccepted values for an explicitly supplied field that could not be resolved.
next_questionNoOne clarification covering every unresolved input. Present only when status is needs_input; ask it once, then call the tool again with the answers in the explicit fields.
missing_fieldsYes
resolved_inputsYesCanonical fields resolved so far. Explicit corrections override proposal inference.

Schema Changelog

Changes observed during successful MCP inspections.

  1. Addedv0.14.14

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond the readOnly/idempotent/no-open-world annotations, it discloses deterministic calculation, no authentication requirement, telemetry with an opt-out env var, that unresolved values are never guessed, that stated gaps block Accelerate and stay in the audit trail, and that a fully-estimated pass cannot return Accelerate. That is substantive process and side-effect context the annotations do not carry.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the trigger and verdict semantics, and almost every sentence earns its place. It is quite long and repeats the 'never guesses unresolved values' point, but the density is justified by the tool's complexity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be described, and the description still covers what matters: blocking conditions, estimation fallbacks, sibling routing, auth and telemetry. Complete for a 9-parameter tool with nested objects and verdict logic.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3, but the description adds cross-parameter behavior: pass only evidence-backed pillars, missing pillars are estimated and reported via pillar_basis/scores_used, decision score is haircut by estimation, and the intended call-then-re-call cadence. It does not restate every field, which is appropriate given the schema already documents them.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('assess an AI initiative', 'front door for one AI investment decision') and specifies what it resolves (industry, revenue, function, AI tier, readiness) and what it returns (one clarification or an Accelerate/Fix/Stop verdict). It explicitly distinguishes itself from score_initiative, score_portfolio, diagnose_process and work_architecture.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives an explicit trigger ('CALL THIS FIRST when the user describes an AI idea in ordinary language'), the re-call workflow once the clarification is answered, and routes each alternative by condition: score_initiative for known canonical fields, score_portfolio for several initiatives, diagnose_process for measured waste, work_architecture for redesign evidence. When-to-use and when-not-to-use are both covered.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.