Skip to main content
Glama

AI BVF: review one AI investment decision

Turn an AI proposal into a decision brief: the verdict, the evidence gaps and the next action for the person who owns the work.

Try one proposal in your browser

Open aibvf.com/start. Your first assessment needs no account or installation.

Paste this synthetic example, or describe a proposal you are reviewing:

We are a EUR 2.4bn manufacturer considering GenAI for predictive maintenance in our supply chain. Our operating model is traditional. We have a sponsor, but the affected roles, human override rights and performance measures still need review.

Run the assessment, check how the proposal was interpreted, and correct the assumptions. The assessment asks for unresolved inputs and marks estimated pillar scores. Keep the action list, name an owner and return with evidence when the work changes.

The decision brief

The result brings the following questions together. This is an illustrative reading guide for the example, with the engine's field names documented below.

Part of the brief

What to review

Verdict

Stop, Fix or Accelerate, with the rule that produced the call. An initial proposal with estimated pillars remains provisional.

Evidence gaps

Which inputs were supplied, which were estimated, and what remains unknown about workflow, roles, decision rights and measures.

Planning benefit

A readiness-adjusted EUR scenario. Build costs, operating costs, change costs and financial timing still need a separate business case.

Decision score

A 0 to 100 summary of pillar inputs and input completeness. It has no calibrated probability interpretation.

Next action

The evidence or work change needed before another review, with an accountable owner assigned by the team.

For the example above, the first useful action is to evidence the proposed workflow and its owners. An Accelerate verdict requires the pillar thresholds to clear and all four work architecture checks to be met.

See the reproducible worked example for exact inputs, calculations and a re-score that stays at Fix until the work architecture is evidenced.

Related MCP server: Intentions MCP Server

Bring the assessment into your existing workflow

The browser is the first-use route. The hosted connector and local MCP package let an assistant repeat the assessment while you work on a proposal.

Hosted connector on claude.ai

Open Settings, then Connectors, then Add custom connector, and paste:

https://mcp.aibvf.com/api/mcp

Then ask:

Assess this AI initiative using AI BVF: we are a EUR 2.4bn manufacturer considering GenAI predictive maintenance, with a traditional operating model. Resolve the inputs, show the assumptions and identify the evidence needed for the next decision.

Local with Claude Desktop, Claude Code or Cursor

Add this configuration to your MCP client:

{
  "mcpServers": {
    "aibvf": { "command": "npx", "args": ["-y", "aibvf-mcp"] }
  }
}

Quit and restart the client. Windows uses cmd /c; the client setup and troubleshooting guide contains the complete configuration.

Start with assess_ai_initiative. Use recommend_improvements for a Fix or Stop, review the proposed actions with the people who own the work, and re-score after the evidence changes.

How to read a result

AI BVF is a deterministic planning model with disclosed assumptions. The source and formulas can be inspected, and identical inputs produce identical outputs for a given engine version.

The four pillars are Strategic Alignment, Financial Return, Change Enablement and Governance Risk. GR >= 70 or FR <= 20 returns Stop; SA >= 60, FR >= 60, CE >= 60 and GR <= 40 clear the pillar test for Accelerate. A gap, partial assessment or missing work architecture evidence holds an otherwise green initiative at Fix.

Keep these boundaries with the result:

  • Decision score. Existing API fields confidence, decision_confidence and projected_confidence retain their names for compatibility. These are rule-based scores, with no measured probability of project success or prediction accuracy.

  • Planning benefit. The scorer's net_low_eur, net_high_eur and MCP net_value_eur apply a readiness capture assumption to a revenue-based benefit scenario. Project costs, margins, timing and overlap are outside that calculation.

  • Research context. External research informs the questions. It does not publish or validate the AI BVF function rates, industry multipliers, readiness capture percentages or drag rates.

  • Module labels. applied_modules records implementation context. Labels such as healthcare_clinical_validation and financial_dora_module do not perform clinical validation or certify regulatory compliance.

  • Supplied evidence. The engine records supplied values and work architecture checks. The organisation remains responsible for reviewing the evidence behind them.

Read the scoring formulas and worked example before using the outputs in a funding decision.

Tools for developers

Thirteen tools are exposed through local stdio and the hosted Streamable HTTP connector.

Tool

Purpose

assess_ai_initiative

Resolve a plain-English proposal, request missing decision inputs and return the assessment.

score_initiative

Score explicit inputs, apply the work architecture gate, and return reasoning, audit and sensitivity.

recommend_improvements

Propose pillar actions and work redesign steps for a Fix or Stop. Re-score evidence before accepting a projected outcome.

assemble_portfolio

Structure loose portfolio inputs, resolve aliases and disclose estimated pillars.

validate_portfolio

Validate a portfolio document against the published JSON Schema.

score_portfolio

Score a portfolio and return its aggregate shape. Review benefit overlap before using a total.

sequence_portfolio

Produce rollout waves with change-capacity constraints and named gates.

diagnose_process

Evaluate observed process signals and return an intervention with its modelled effect.

infer_readiness

Infer a readiness classification from supplied process signals and report coverage.

calculate_pace_layer_drag

Return a directional scenario for operating-model friction using disclosed rates.

get_benchmark

Return AI BVF planning rates with evidence status and use guidance.

list_taxonomy

List the accepted industries, functions, AI tiers and readiness levels.

map_to_taxonomy

Map everyday business terms to the supported taxonomy and expose unresolved terms.

For portfolios, use assemble_portfolio, validate_portfolio, score_portfolio, then sequence_portfolio. An aggregate modelled range needs a separate review of overlapping work and shared benefits.

Packages and public specification

Package

Version

Purpose

aibvf-mcp

0.14.14

MCP server, 13 tools, stdio plus hosted Streamable HTTP at mcp.aibvf.com.

aibvf-check

0.1.1

Policy checks for a declared AI initiative manifest in CI.

@aibvf/core

0.10.6

TypeScript assessment and scoring engine.

aibvf

0.2.2

Python scoring engine and validator. Check its documented feature coverage before substituting it for the TypeScript implementation.

The public portfolio specification is version 1.0. That document format has a separate version from the packages implementing it; the package version identifies the code and behaviour used for a particular assessment.

Protocol page · npm package · MCP registry · Release history

Anonymous usage telemetry

The MCP server can report tool calls and a server_connect event. Events include protocol and package versions, entry route, assessment stage, work architecture status, taxonomy fields, a daily-rotated caller hash, and classification plus confidence where supplied.

Local stdio calls also include a stable one-way install_hash for repeat-use measurement, derived from a random local seed. Hosted calls send no stable install hash, and a broad user_role is sent only when a local user explicitly sets AIBVF_USAGE_ROLE.

No portfolio content, revenue figures, numeric pillar scores or personal identifiers are included. Set AIBVF_TELEMETRY_DISABLE=1 to prevent events and creation of the local install-id file. Point at your own backend with AIBVF_TELEMETRY_URL and AIBVF_TELEMETRY_KEY.

Package downloads include repeat installs, dependencies and automation. Use completed assessments and subsequent decision reviews to evaluate adoption.

Contribute a case or a correction

Bring a reproducible counterexample: the inputs, actual output, expected decision, supporting evidence and engine version. The contribution guide includes a template and explains the review and licensing boundaries.

The ten-team pilot pack defines the first-use and return-use checks, interview prompts and tracker. Examples in this repository are synthetic unless a case explicitly records consent and its evidence.

If AI BVF helped you review a decision, star the repository or share a counterexample. Both give the project useful feedback.

License

Repository source code is MIT licensed under LICENSE. The specification and JSON Schema under spec/ are CC-BY-4.0, and the AI BVF names and logo are trademarks, as set out in NOTICE.

Private benchmark material and certification marks are outside the source-code contribution route. The contribution guide explains how to discuss those materials without changing the rights granted by the repository licenses.

About the author

Craig Horton is an independent transformation lead based in Amsterdam and the author of the AI Business Value Framework. His work connects AI investment decisions with organisational readiness and the redesign of work.

The Transformation Brief · Craig Horton on LinkedIn

Available Tools

13 tools
assemble_portfolioA
Read-onlyIdempotent
Inspect

Assemble a valid AI BVF v1.0 portfolio document from loose inputs, deterministically. Agents arrive with initiative names, plain-language functions and half the pillar scores, then hand-build the portfolio JSON and get the shape wrong; this tool builds it right. Give it the organisation (name plus industry in canonical or everyday language) and one entry per initiative (name, function, ai_tier, plus whatever pillar scores you actually have as bare numbers) and it returns the finished document: aliases resolved through the same mapping as map_to_taxonomy, ids generated from names and deduplicated, missing pillars estimated from readiness, tier, function and disclosed AI BVF planning assumptions with the estimation reported per initiative in estimated_pillars, and the whole document validated before it is returned. CALL THIS when the user lists several AI initiatives in conversation and you need a portfolio document for validate_portfolio, score_portfolio or sequence_portfolio, instead of composing the JSON by hand. Do NOT invent pillar scores to fill it: pass only the numbers the user gave you and let the estimation carry the rest honestly, the estimated pillars carry low confidence and scoring haircuts accordingly. Unresolvable inputs come back as issues with suggestions; ask the user to choose rather than guessing. Every default the assembler applies is named in plain language in assumptions: surface them to the user, the assembler structures inputs and never makes hidden business judgements. This tool creates a document in the response only: nothing is stored, nothing is edited, no state exists between calls. Deterministic calculation with no authentication. Anonymous usage telemetry may be sent; set AIBVF_TELEMETRY_DISABLE=1 to opt out.

ParametersJSON Schema
NameRequiredDescriptionDefault
readinessNoOrganisational readiness, canonical or plain language (bureaucratic resolves to siloed). Drives estimation of missing pillars. Defaults to traditional.
initiativesYesOne entry per initiative, from whatever the user gave you. Only name, function and ai_tier are required.
organizationYesOrganisation identity and context shared by every initiative in the assembled portfolio.

Output Schema

ParametersJSON Schema
NameRequiredDescription
auditYesReproducibility record: engine version, the rules that fired, and the resolved inputs. Deterministic, no timestamps. If the verdict is challenged months later, the same inputs on the same engine version reproduce it exactly.
issuesYesUnresolved inputs, each with path, message and suggestions where the taxonomy has them.
guidanceYes
portfolioNoThe assembled BVF v1.0 document, ready for validate_portfolio, score_portfolio and sequence_portfolio. Null when assembly is blocked on issues.
validationNovalidate() run on the assembled document.
assumptionsYesEvery default the assembler applied, in plain language. Surface these to the user: what was not given is named here.
bvf_versionYes
resolutionsYesEvery alias resolution performed, in plain language.
readiness_usedYes
estimated_pillarsYesInitiative id to the pillars the assembler estimated. Gather evidence for these, or expect scoring to haircut confidence.

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond the annotations (readOnly, idempotent, non-destructive, closed-world), the description discloses determinism, the estimation policy for missing pillars, per-initiative reporting via estimated_pillars, validation before return, plain-language assumptions, no hidden business judgements, error-as-issues behavior, statelessness ('nothing is stored, nothing is edited, no state exists between calls'), no authentication, and a telemetry opt-out env var. This is unusually complete behavioral disclosure.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the purpose, then usage, then behavior — a sensible order. It is dense but nearly every sentence carries distinct information (estimation policy, assumptions, statelessness, telemetry). A couple of clauses about assumptions and hidden business judgements restate each other and could be merged.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a nested, 3-parameter tool with a rich input schema and an output schema, the description covers what the schema cannot: alias resolution, id generation, estimation fallbacks, assumption reporting and failure mode. Return values are not re-explained, correctly leaving that to the output schema.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% so the baseline is 3, but the description adds real semantics the schema does not: alias resolution uses the same mapping as map_to_taxonomy, ids are generated from names and deduplicated, missing pillars are estimated from readiness, tier and function, and readiness defaults to traditional. It stops short of describing every field, which the schema already handles.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The opening sentence gives a precise verb+resource+guarantee: 'Assemble a valid AI BVF v1.0 portfolio document from loose inputs, deterministically.' It immediately frames the problem it solves (agents hand-building JSON and getting the shape wrong), which separates it from validate_portfolio, score_portfolio and sequence_portfolio.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicit trigger ('CALL THIS when the user lists several AI initiatives in conversation and you need a portfolio document for validate_portfolio, score_portfolio or sequence_portfolio'), explicit alternative ('instead of composing the JSON by hand'), and explicit when-not ('Do NOT invent pillar scores to fill it: pass only the numbers the user gave you'). Unresolvable inputs are also routed back to the user rather than guessed.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

assess_ai_initiativeA
Read-onlyIdempotent
Inspect

The front door for one AI investment decision. CALL THIS FIRST when the user describes an AI idea in ordinary language or asks whether it should proceed. It resolves industry, revenue, business function, AI tier and organisational readiness, then returns one clarification covering every unresolved input or an Accelerate, Fix or Stop verdict. Ask that clarification once and call this tool again with the answers in the explicit fields. Use work_architecture to test whether the end-to-end workflow, affected roles, human decision rights and performance measures have been redesigned. A stated gap or missing work architecture evidence blocks Accelerate and stays visible in the audit trail. Pillar scores and work architecture evidence remain optional inputs, but unresolved values are never guessed and cannot unlock Accelerate. Use score_initiative when the canonical fields are already known, score_portfolio for several initiatives, and diagnose_process for measured waste in a running process. Deterministic calculation with no authentication. Anonymous usage telemetry may be sent; set AIBVF_TELEMETRY_DISABLE=1 to opt out.

ParametersJSON Schema
NameRequiredDescriptionDefault
scoresNoOPTIONAL, and each pillar inside it is optional. The four AI BVF pillars, each an honest 0–100 self-assessment, combining deterministically into the verdict: governance_risk ≥ 70 OR financial_return ≤ 20 returns Stop; strategic_alignment, financial_return and change_enablement all ≥ 60 with governance_risk ≤ 40 returns Accelerate; everything else returns Fix. Pass ONLY the pillars the user has real evidence for — do NOT invent numbers for the rest. Missing pillars are estimated deterministically by the engine from disclosed AI BVF planning assumptions, the response reports which via pillar_basis and scores_used, decision score is haircut by how much was estimated, and a fully-estimated pass can never return Accelerate (it returns Fix pending confirmation). So call immediately with whatever the user gave you, then ask for evidence on the estimated pillars and re-call to firm the verdict up.
ai_tierNoOptional correction or answer: automation/RPA, GenAI/copilot, or agentic/autonomous. Overrides anything inferred from proposal.
functionNoOptional correction or answer in canonical or everyday language, for example customer service, procurement, finance or risk. Overrides anything inferred from proposal.
industryNoOptional correction or answer in canonical or everyday language, for example retail, hospital, bank or public sector. Overrides anything inferred from proposal.
proposalYesThe AI initiative in ordinary business language. Include the organisation, industry, approximate annual revenue, business function, AI ambition and how the organisation works today when known. The resolver extracts what it can and, when several inputs are missing, asks for all of them in one clarification; it never guesses an unresolved taxonomy value.
readinessNoOptional correction or answer: agile, traditional, or siloed, including everyday descriptions such as cross-functional, hierarchical or bureaucratic. Overrides anything inferred from proposal.
revenue_eurNoOptional approximate annual revenue in EUR. Overrides any EUR amount extracted from proposal. No currency conversion is performed.
work_architectureNoEvidence that workflows, roles, decision rights and measures are ready. Pass only what is known. Explicit gaps and omitted checks block Accelerate until all four checks are evidenced.
signal_completenessNoOptional 0 to 1 input-quality factor. The default ranges from 0.5 when all pillars are estimated to 1 when all are supplied. Supplied values still require evidence review. Lower this factor when the supplied pillars rest on weak evidence.

Output Schema

ParametersJSON Schema
NameRequiredDescription
statusYesneeds_input when one or more required decision inputs remain unresolved; verdict when scoring completed.
verdictNoThe AI BVF score. Present only when status is verdict.
proposalYesThe supplied proposal, returned so the next call can preserve it verbatim.
bvf_versionYes
resolutionsYesEvery deterministic resolution, naming the field, canonical value, source and matched phrase.
suggestionsNoAccepted values for an explicitly supplied field that could not be resolved.
next_questionNoOne clarification covering every unresolved input. Present only when status is needs_input; ask it once, then call the tool again with the answers in the explicit fields.
missing_fieldsYes
resolved_inputsYesCanonical fields resolved so far. Explicit corrections override proposal inference.

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond the readOnly/idempotent/no-open-world annotations, it discloses deterministic calculation, no authentication requirement, telemetry with an opt-out env var, that unresolved values are never guessed, that stated gaps block Accelerate and stay in the audit trail, and that a fully-estimated pass cannot return Accelerate. That is substantive process and side-effect context the annotations do not carry.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the trigger and verdict semantics, and almost every sentence earns its place. It is quite long and repeats the 'never guesses unresolved values' point, but the density is justified by the tool's complexity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be described, and the description still covers what matters: blocking conditions, estimation fallbacks, sibling routing, auth and telemetry. Complete for a 9-parameter tool with nested objects and verdict logic.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3, but the description adds cross-parameter behavior: pass only evidence-backed pillars, missing pillars are estimated and reported via pillar_basis/scores_used, decision score is haircut by estimation, and the intended call-then-re-call cadence. It does not restate every field, which is appropriate given the schema already documents them.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('assess an AI initiative', 'front door for one AI investment decision') and specifies what it resolves (industry, revenue, function, AI tier, readiness) and what it returns (one clarification or an Accelerate/Fix/Stop verdict). It explicitly distinguishes itself from score_initiative, score_portfolio, diagnose_process and work_architecture.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives an explicit trigger ('CALL THIS FIRST when the user describes an AI idea in ordinary language'), the re-call workflow once the clarification is answered, and routes each alternative by condition: score_initiative for known canonical fields, score_portfolio for several initiatives, diagnose_process for measured waste, work_architecture for redesign evidence. When-to-use and when-not-to-use are both covered.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

calculate_pace_layer_dragA
Read-onlyIdempotent
Inspect

Quantify the annual EUR cost of an AI ambition outrunning the operating model: queues, hand-offs and slow decisions that prevent the organisation capturing the value already assumed in the case. CALL THIS when the user needs the cost of waiting for the organisation to change, or when a Fix plan needs a cost-of-waiting figure. Do not use it to score an AI initiative, estimate the implementation cost, or calculate a process saving: use score_initiative for the investment verdict, diagnose_process for a running process, and recommend_improvements for the change plan. revenue_eur sets the absolute EUR range; ai_tier and readiness together set the drag rate and pace_gap, so gen3 in a siloed organisation costs more than gen1 in an agile one. industry is accepted for a consistent interface and defaults to universal, but does not change this calculation yet. Returns a low/high EUR range, drag rate, pace-gap severity, drivers and source. Deterministic calculation with no authentication. Anonymous usage telemetry may be sent; set AIBVF_TELEMETRY_DISABLE=1 to opt out.

ParametersJSON Schema
NameRequiredDescriptionDefault
ai_tierYesAmbition of the AI operating model: gen1 = automation/RPA, gen2 = GenAI, gen3 = agentic. Paired with readiness to set pace_gap severity — gen3 on any readiness below agile, or gen2 on siloed, is severe; a higher tier against a slower operating model widens the gap and raises the drag.
industryNoOptional; defaults to universal if omitted. Reserved for future vertical drag-rate adjustments — does not change the result today. Call list_taxonomy for accepted values.
readinessYesOrganisational readiness, honest self-assessment: agile = cross-functional, fast decisions; traditional = functional hierarchy; siloed = rigid, hand-off heavy. Agile readiness yields minimal drag at any tier; the mismatch between a fast AI tier and a slower operating model is what generates the Organisational Drag Cost.
revenue_eurYesApproximate annual revenue in EUR (must be ≥ 0). The result scales with this: annual_drag_eur is returned as an absolute range and as drag_rate, a fraction of this revenue (e.g. 0.02 = 2%).

Output Schema

ParametersJSON Schema
NameRequiredDescription
sourceYesCitation for the drag-rate model applied.
driversYesNamed factors contributing to the drag.
pace_gapYesSeverity of the tier↔readiness mismatch.
drag_rateYesDrag as a fraction of revenue (e.g. 0.02 = 2%), low/high.
bvf_versionYesAI BVF protocol version used.
annual_drag_eurYesEstimated annual Organisational Drag Cost in EUR, low/high.

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnly/idempotent/non-destructive/closed-world, and the description adds genuinely non-redundant behavior: deterministic calculation, no authentication required, anonymous telemetry with an explicit opt-out env var (AIBVF_TELEMETRY_DISABLE=1), and the shape of the result (low/high EUR range, drag rate, pace-gap severity, drivers, source). It also discloses the caveat that industry is accepted but does not affect the result yet.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with purpose, then triggers, exclusions, parameter mechanics, return shape and operational notes in that order. It is dense rather than padded, though the parameter-mechanics sentence overlaps somewhat with the already-detailed schema descriptions.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a deterministic read-only calculator with an output schema, the description covers everything an agent needs: when to call it, when not to, what drives the result, what industry does (nothing yet), what comes back, and the auth/telemetry profile. No material gap remains.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% and the enum descriptions are already rich, so the baseline is 3. The description adds value by stating how parameters interact ('ai_tier and readiness together set the drag rate and pace_gap', 'revenue_eur sets the absolute EUR range') and by flagging that industry is interface-only and does not change the calculation — a meaningful caveat an agent would otherwise mis-assume.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource: quantifying the annual EUR cost of an AI ambition outrunning the operating model (organisational drag / cost of waiting). It also scopes itself against siblings by naming score_initiative, diagnose_process and recommend_improvements as the tools for adjacent questions, so an agent can route without opening another schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicit trigger ('CALL THIS when the user needs the cost of waiting... or when a Fix plan needs a cost-of-waiting figure') followed by explicit exclusions mapped to named alternatives ('Do not use it to score an AI initiative... use score_initiative for the investment verdict, diagnose_process for a running process, recommend_improvements for the change plan'). This is textbook when/when-not/alternative guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

diagnose_processA
Read-onlyIdempotent
Inspect

Diagnose a single existing business process from operational evidence and return the intervention, modelled net EUR saving, efficiency gain, verdict and confidence. CALL THIS when the user can describe a process already running, including volume, touch time, waiting, hand-offs, rework, automation and cost. instances_per_year × fte_hours_per_instance × loaded_hourly_rate_eur builds the labour baseline, direct_spend_eur adds the non-labour baseline, and readiness caps the saving that the organisation can realise. The friction signals select the intervention: low automation points to Automate, many hand-offs or wait to Consolidate & re-sequence, rework to Quality controls, low-volume heavy work to Eliminate / insource. signal_completeness must fall when inputs are estimated, because it directly reduces decision score. Use score_initiative for a proposed AI investment and infer_readiness when the question is the organisation’s change capacity. Effectiveness bands are benchmark-cited and figures are directional, not audited. Deterministic calculation with no authentication. Anonymous usage telemetry may be sent; set AIBVF_TELEMETRY_DISABLE=1 to opt out.

ParametersJSON Schema
NameRequiredDescriptionDefault
functionYesBusiness function the process belongs to. See list_taxonomy.
handoffsYesDistinct owners/systems an instance passes through. Weighed against the per-function median; many handoffs make handoff drag dominant and point to Consolidate & re-sequence.
readinessNoOptional. Org change-absorption capacity — agile / traditional / siloed — which caps the realised (net) saving below the gross potential. Defaults to traditional.
process_idYesStable identifier for the process.
rework_rateYesFraction of instances reopened/reworked (0–1). When rework is the dominant drag factor the intervention becomes Quality controls, and it also sets the addressable share for that path.
touch_ratioYesTouch-time ÷ cycle-time (0–1). The remainder is wait; a low value means the process is mostly waiting, which pushes the intervention toward Consolidate & re-sequence.
cycle_time_daysYesMedian wall-clock days per instance, end to end. Long cycles relative to touch-time signal wait/latency drag.
automation_levelYesShare already automated (0–1). Low automation makes manual effort the dominant drag and selects Automate; the un-automated remainder is the addressable share.
direct_spend_eurYesAnnual licence/vendor/tooling spend on the process in EUR. Added to the labour baseline and shifts how much of the saving is labour- vs spend-addressable.
instances_per_yearYesProcess volume: how many times it runs per year. Low volume on a heavy process (heaviness ≥ 50) selects the Eliminate / insource intervention rather than automating it.
signal_completenessNoOptional 0–1. How much of the above was measured versus defaulted. Governs decision_confidence proportionally — lower it when you estimated inputs so the verdict stays honest. Defaults to 0.7.
fte_hours_per_instanceYesHuman touch-time in hours per instance. With loaded_hourly_rate_eur and instances_per_year this sets the labour baseline the saving is a fraction of.
loaded_hourly_rate_eurYesFully-loaded labour cost per hour in EUR (salary + on-costs). Multiplies fte_hours_per_instance × instances_per_year into the annual labour baseline.

Output Schema

ParametersJSON Schema
NameRequiredDescription
verdictYesThe call on the intervention.
functionYesBusiness function diagnosed.
heavinessYesProcess heaviness index, 0–100.
disclaimerYesDirectional decision aid, not an audited figure.
process_idYesEcho of the input process id.
assumptionsYesThe assumptions behind the figure — never a naked number.
bvf_versionYesAI BVF protocol version used.
interventionYesRecommended move.
brain_versionYesAdvisor Brain model version used.
net_saving_eurYesModelled net annual saving in EUR after readiness capture, low/high.
offer_to_executeYesTrue when the verdict warrants offering to action it (Accelerate).
baseline_cost_eurYesCurrent annual cost: labour + direct spend.
evidence_maturityYesStrength of the benchmark evidence behind the effectiveness band.
advisory_next_stepNoOptional CTA, present only for Fix/Stop verdicts.
drag_decompositionYesShare of heaviness from each friction factor (sums to ~1).
decision_confidenceYesConfidence in the verdict, 0–100.
efficiency_gain_pctYesEfficiency improvement on the targeted slice, percent.

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnly/idempotent/non-destructive, so the safety profile is covered; the description goes well beyond that by disclosing that it is a deterministic calculation with no authentication, that anonymous telemetry may be sent with an opt-out env var, and that effectiveness bands are benchmark-cited but figures are directional rather than audited. It also explains the internal model (baseline construction, readiness capping) which is genuine behavioral context.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

It is on the long side for a description, but it is front-loaded with purpose and return values, followed by trigger, model logic, sibling routing, and caveats in a sensible order. Nearly every sentence carries distinct information; the friction-signal mapping is dense but earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 13-parameter, 11-required analytical tool with an output schema, the description supplies what the schema cannot: the calculation model, the friction-to-intervention decision logic, sibling routing, and accuracy/telemetry caveats. Since an output schema exists, the brief enumeration of return values is sufficient and nothing an agent needs to call it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the schema already documents all 13 parameters, meaning the baseline is 3. The description still adds value by explaining how inputs combine (instances_per_year × fte_hours_per_instance × loaded_hourly_rate_eur builds the labour baseline; direct_spend_eur adds non-labour; readiness caps realised saving) and how friction signals map to interventions, which is causal meaning not present in the field descriptions alone.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Diagnose a single existing business process from operational evidence') and enumerates the outputs (intervention, modelled net EUR saving, efficiency gain, verdict, confidence). It actively distinguishes itself from siblings by naming score_initiative and infer_readiness with the conditions that select each. An agent can route without opening any schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

'CALL THIS when the user can describe a process already running, including volume, touch time, waiting, hand-offs, rework, automation and cost' gives an explicit trigger condition. It then names two alternative tools (score_initiative for a proposed AI investment, infer_readiness for change capacity), covering both when-to-use and when-not-to-use.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_benchmarkA
Read-onlyIdempotent
Inspect

Look up the disclosed AI BVF planning rates behind the value model for one business function and industry. CALL THIS when the user wants to inspect the revenue-uplift and cost-takeout assumptions before scoring, or to compare the value drivers across functions. function selects the base rate range and named drivers; industry applies the multiplier, while universal returns the unadjusted base rate. External research in the evidence register frames the adoption and value problem but does not publish these function rates. The output is a rate, expressed as a fraction of revenue, not an initiative verdict or EUR business case. Replace it with measured organisation evidence before funding. Use score_initiative for an Accelerate/Fix/Stop decision, score_portfolio for several initiatives and diagnose_process for measured operational waste. Deterministic lookup with no authentication. Anonymous usage telemetry may be sent; set AIBVF_TELEMETRY_DISABLE=1 to opt out.

ParametersJSON Schema
NameRequiredDescriptionDefault
functionYesBusiness function to benchmark — must be one of the list_taxonomy function values. Selects the base revenue-uplift and cost-reduction rate ranges (returned as fractions of revenue) and the value drivers.
industryYesIndustry whose multiplier to apply — must be one of the list_taxonomy industry values. The returned industry_multiplier is applied to the function base rates; pass "universal" for the un-adjusted rates.

Output Schema

ParametersJSON Schema
NameRequiredDescription
sourceYesCitation for the benchmark figures.
driversYesNamed value drivers behind the benchmark.
functionYesBusiness function the rates apply to.
industryYesIndustry whose multiplier was applied.
cost_takeout_rangeYesCost take-out as a fraction of revenue, lo/hi.
industry_multiplierYesMultiplier applied to the base rates for this industry.
revenue_uplift_rangeYesRevenue uplift as a fraction of revenue, lo/hi.

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond the readOnly/idempotent/openWorld annotations it discloses that the lookup is deterministic and unauthenticated, that anonymous telemetry is sent with an AIBVF_TELEMETRY_DISABLE=1 opt-out, and that the result is a planning rate rather than a verdict or business case. That is genuinely additive operational context.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads purpose, then triggers, then alternatives, then caveats. Seven sentences is on the long side and the telemetry/evidence-register lines could be trimmed, but each sentence carries distinct decision-relevant information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Covers the full picture for a 2-param read-only lookup: what it returns (a rate as a fraction of revenue), what it is not, when to use alternatives, and the operational caveats. With an output schema present, no further return-value detail is required.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3, but the description adds the interaction semantics: 'function selects the base rate range and named drivers; industry applies the multiplier, while universal returns the unadjusted base rate.' This clarifies how the two enum parameters combine rather than just restating their types.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Look up the disclosed AI BVF planning rates behind the value model') with explicit scope ('for one business function and industry'). It also draws a clear boundary against siblings by naming score_initiative, score_portfolio and diagnose_process, so an agent can route without opening schemas.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives an explicit trigger ('CALL THIS when the user wants to inspect the revenue-uplift and cost-takeout assumptions before scoring, or to compare the value drivers across functions') plus named alternatives for the adjacent jobs (verdicts, portfolio scoring, measured waste). It even pre-empts a likely confusion by stating that external research does not publish these rates.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

infer_readinessA
Read-onlyIdempotent
Inspect

Measure organisational readiness from process data, so the investment case does not depend on an untested maturity claim. CALL THIS before score_initiative, score_portfolio or calculate_pace_layer_drag when the user can provide at least two of five signals: hand-offs, rework, touch ratio, automation level and cycle time. function selects the comparison medians for hand-offs and cycle time; more signals increase confidence and disagreement between them reduces it. claimed_readiness is optional, but pass it when the organisation has declared itself agile, traditional or siloed, because the returned gap exposes where its self-image runs ahead of the process data. Fewer than two signals produces a refusal, not a guess. Pass the measured readiness into the downstream tool, then use diagnose_process when the next question is what to change in that process. Deterministic calculation with no authentication. Anonymous usage telemetry may be sent; set AIBVF_TELEMETRY_DISABLE=1 to opt out.

ParametersJSON Schema
NameRequiredDescriptionDefault
functionYesBusiness function the process belongs to. Selects the disclosed AI BVF cycle-time and hand-off reference points used to interpret the signals. Call list_taxonomy if unsure.
handoffsNoDistinct owners or systems an instance passes through. Read against the function median: 1.5x or more the median reads siloed, at or above the median reads traditional, below it reads agile.
rework_rateNoFraction of instances reopened or reworked (0-1). 15% or more reads siloed, 5-15% traditional, under 5% agile.
touch_ratioNoTouch-time divided by cycle-time (0-1); the remainder is waiting. Under 0.15 reads siloed (the process lives in queues), 0.15-0.4 traditional, above 0.4 agile.
cycle_time_daysNoMedian wall-clock days per instance. Read against the function median, same bands as handoffs.
automation_levelNoShare of the process already automated (0-1). Under 0.2 reads siloed, 0.2-0.5 traditional, above 0.5 agile.
claimed_readinessNoOptional. What the organisation says about itself. The measured result is compared against it and the gap returned as readiness_gap plus a gap_finding, because an organisation whose self-image runs ahead of its process data has just told you where the change work starts.

Output Schema

ParametersJSON Schema
NameRequiredDescription
auditNoReproducibility record: engine version, the rules that fired, and the resolved inputs. Deterministic, no timestamps. If the verdict is challenged months later, the same inputs on the same engine version reproduce it exactly.
guidanceYesHow to use the result downstream, including what a gap between measured and self-reported readiness means.
readinessYesThe readiness classification the measured signals support.
confidenceYesConfidence 0-100, set by signal coverage (2 signals ~45, 5 signals ~90) and discounted when signals disagree.
bvf_versionYesAI BVF protocol version used.
gap_findingNoThe claimed-versus-measured gap read as a change-readiness finding. Surface verbatim when present.
disagreementNoPresent when signals point in opposing directions: readiness is uneven across the process, read the per-signal detail.
signal_readsYesPer-signal read: the value, which readiness it leans toward, and why in plain language. Show these to the user.
signals_usedYesHow many of the five signals were provided.
readiness_gapNoOrdinal distance claimed-to-measured. Positive: the organisation claims better than it measures.
readiness_basisYesAlways measured: this came from process data, not self-report.
claimed_readinessNoEcho of the claim, when supplied.

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnly/idempotent/non-destructive, and the description adds substantial extra behavior: refusal thresholds, confidence scaling with signal count, disagreement lowering confidence, determinism, no authentication, and telemetry with an explicit opt-out env var. This is exactly the kind of context annotations cannot carry.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with purpose and call condition, and most sentences carry operational weight. It is dense and slightly long, with mild overlap between the determinism/no-auth sentence and the telemetry sentence, but no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With an output schema present it correctly avoids describing return values, while covering everything else an agent needs: prerequisites, refusal behavior, confidence mechanics, optional-parameter guidance, downstream chaining, and telemetry disclosure.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3, but the description adds meaning beyond the schema by explaining that 'function' selects comparison medians and by giving conditional guidance for the optional claimed_readiness (pass it when the org has self-declared agile/traditional/siloed because the gap is diagnostic).

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('Measure organisational readiness from process data') and immediately frames the business purpose. It also differentiates itself from siblings by naming score_initiative, score_portfolio and calculate_pace_layer_drag as downstream consumers.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicit invocation conditions ('CALL THIS before ... when the user can provide at least two of five signals'), an explicit exclusion ('Fewer than two signals produces a refusal, not a guess'), and a routing hint to diagnose_process for the follow-up question. Nothing is left to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_taxonomyA
Read-onlyIdempotent
Inspect

Return the exact industry, function, AI-tier and readiness values every AI BVF calculation accepts. CALL THIS when the caller needs the complete allowed list or when a free-text value is not obvious. It returns taxonomy only, no score, verdict or language mapping. Use map_to_taxonomy when the user has said customer service, banking, RPA or bureaucratic and you need the one canonical value; use this tool when they need the whole menu of values to choose from. Takes no parameters. Deterministic lookup with no authentication. Anonymous usage telemetry may be sent; set AIBVF_TELEMETRY_DISABLE=1 to opt out.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription
ai_tiersYesAll accepted ai_tier values (gen1/gen2/gen3).
functionsYesAll accepted business-function values.
readinessYesAll accepted organisational-readiness values.
industriesYesAll accepted industry values.
bvf_versionYesAI BVF protocol version these enums belong to.

TDQS

A4.6/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint and non-destructive. The description adds what the tool does NOT return (no score, verdict or language mapping), that it takes no parameters, is deterministic, requires no authentication, and has an opt-out telemetry behavior with the env var name. That is real behavioral context beyond annotations, though rate limits and output format are not detailed (output schema exists).

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded purpose followed by call guidance, routing rule, and behavioral notes. Well organized, though the telemetry sentence is somewhat tangential to selecting the tool.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-param, annotated, read-only lookup with an output schema, the description covers purpose, routing, exclusions, and operational notes. Nothing an agent needs to select or invoke it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Zero parameters, so baseline is 4. The description reinforces this with 'Takes no parameters.' Nothing further is needed or possible.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (Return) and resource (taxonomy values), and precisely scopes what is returned: 'the exact industry, function, AI-tier and readiness values every AI BVF calculation accepts.' The sibling map_to_taxonomy is named and distinguished.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicit when-to-use: 'CALL THIS when the caller needs the complete allowed list or when a free-text value is not obvious.' Explicit routing to the alternative with examples of inputs that select map_to_taxonomy. No inference required.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

map_to_taxonomyA
Read-onlyIdempotent
Inspect

Map everyday business language to the canonical AI BVF values required by the scoring tools. CALL THIS when the user says customer service, procurement, banking, GenAI copilot or bureaucratic and the matching enum is not certain. Pass only the fields written in free text; each returns the canonical value, what it matched on, or null with suggestions. A null result requires the user to choose from the suggestions, because a plausible guess would change the score. Use list_taxonomy when the user needs every permitted value, then pass the mapped values into score_initiative, diagnose_process, get_benchmark or the portfolio tools. Deterministic lookup with no authentication. Anonymous usage telemetry may be sent; set AIBVF_TELEMETRY_DISABLE=1 to opt out.

ParametersJSON Schema
NameRequiredDescriptionDefault
ai_tierNoEveryday AI language, e.g. RPA, GenAI copilot, autonomous agents. Resolved to gen1/gen2/gen3.
functionNoEveryday function language, e.g. customer service, procurement, legal, people. Resolved to cx, supply, risk, hr and so on.
industryNoEveryday industry language, e.g. banking, ecommerce, pharma. Resolved to the canonical enum.
readinessNoEveryday culture language, e.g. bureaucratic, cross-functional, hierarchical. Resolved to agile/traditional/siloed.

Output Schema

ParametersJSON Schema
NameRequiredDescription
ai_tierNo
functionNo
guidanceYes
industryNoinput, resolved and matched_on; or resolved null with suggestions when no confident match.
readinessNo
bvf_versionYes

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations cover the safety profile (readOnly, idempotent, non-destructive, closed-world), and the description adds rich context beyond them: deterministic lookup, no authentication, null-with-suggestions behavior, the rationale that a guess would change the score, and telemetry disclosure with an opt-out env var. This is well beyond what the structured fields provide.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core purpose, then trigger, then result semantics, then downstream usage. Five dense sentences, all earning their place, though the telemetry note could arguably be trimmed for a mapping tool.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Covers purpose, trigger, alternatives, return/error semantics, downstream routing, auth, and telemetry. With an output schema present, the description need not explain return values further, yet it already handles the null case that matters most for correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3, but the description adds genuine meaning: instruct to pass only free-text fields and explains per-field return semantics (canonical value, what it matched on, or null with suggestions). It complements rather than repeats the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource: map everyday business language to canonical AI BVF values required by scoring tools. It clearly distinguishes itself from the sibling list_taxonomy, which serves a different need (enumerating every permitted value).

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives an explicit trigger ('CALL THIS when the user says customer service, procurement, banking... and the matching enum is not certain') and names the alternative list_taxonomy with the condition that selects it. It also tells the agent where the outputs flow next (score_initiative, diagnose_process, get_benchmark, portfolio tools).

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

recommend_improvementsA
Read-onlyIdempotent
Inspect

Turn a Fix or Stop verdict into the change plan that could earn a re-score, with pillar targets, named plays, owners, stop conditions, cost of waiting and a deadline. CALL THIS after score_initiative returns Fix or Stop, using the same five context fields and any scores or work-architecture evidence from that call. Do not use it to produce the initial verdict, sequence several initiatives or diagnose measured process waste; use score_initiative, sequence_portfolio or diagnose_process for those jobs. Do not call it for Accelerate unless a specific delivery risk needs testing before commitment. resistance_type selects the will or skill route, risk_type selects the regulatory, reputational or operational route, and omitted diagnostics remain provisional with the next question returned. Lead with binding_constraint, surface honest_stop when present, and use rescore_gate to decide whether this remains Fix or becomes Stop. Deterministic calculation with no authentication. Anonymous usage telemetry may be sent; set AIBVF_TELEMETRY_DISABLE=1 to opt out.

ParametersJSON Schema
NameRequiredDescriptionDefault
scoresNoOPTIONAL, and each pillar inside it is optional. The four AI BVF pillars, each an honest 0–100 self-assessment, combining deterministically into the verdict: governance_risk ≥ 70 OR financial_return ≤ 20 returns Stop; strategic_alignment, financial_return and change_enablement all ≥ 60 with governance_risk ≤ 40 returns Accelerate; everything else returns Fix. Pass ONLY the pillars the user has real evidence for — do NOT invent numbers for the rest. Missing pillars are estimated deterministically by the engine from disclosed AI BVF planning assumptions, the response reports which via pillar_basis and scores_used, decision score is haircut by how much was estimated, and a fully-estimated pass can never return Accelerate (it returns Fix pending confirmation). So call immediately with whatever the user gave you, then ask for evidence on the estimated pillars and re-call to firm the verdict up.
ai_tierYesAmbition of the AI being deployed: gen1 = automation/RPA, gen2 = GenAI, gen3 = agentic. Interacts with readiness — a more ambitious tier running on lower readiness widens the pace-layer gap, which discounts the modelled EUR value even when the four pillar scores are strong.
functionYesBusiness function where the AI will operate, as one of the accepted enum values — selects which benchmark value drivers and rate ranges apply. Call list_taxonomy for the exact strings if unsure.
industryYesYour industry, as one of the accepted enum values — used to select the benchmark rate multiplier applied to the modelled EUR value. Call list_taxonomy for the exact strings if unsure.
readinessYesOrganisational readiness, honest self-assessment: agile = cross-functional, fast decisions; traditional = functional hierarchy; siloed = rigid, hand-off heavy. Sets the value-capture rate and, paired with ai_tier, the pace-layer drag — lower readiness against a higher tier reduces the captured value. Self-report is gameable: when the user has real process numbers, call infer_readiness first and pass its measured classification here instead.
risk_typeNoOptional. The nature of a high governance-risk score: "regulatory" = statute applies (EU AI Act, GDPR Article 22, DORA), "reputational" = the risk is how failure looks and lands publicly, "operational" = the system failing quietly inside a process. Selects between a regulatory remediation sequence, visible trust guardrails, and a proportionate governance review. If you do not know, omit it: the engine infers (gen3 tier, or a regulated function/industry, infers regulatory) and marks the play provisional.
revenue_eurYesApproximate annual revenue in EUR (must be ≥ 0). Scales the whole output: the disclosed AI BVF planning rates are applied as fractions of this figure, so the modelled EUR value range grows with it. A rough order-of-magnitude estimate is fine.
resistance_typeNoOptional. What sits behind a low change-enablement score: "will" = people do not want the change (power shifts, fear, no case for change), "skill" = people cannot yet do it (capability and capacity gap). Selects between a coalition-building play (Kotter 1-2 + ADKAR Awareness/Desire) and an owner-and-capability play (ADKAR Knowledge/Ability). If you do not know, omit it: the engine infers from readiness (agile infers skill, traditional/siloed infers will) and marks the play provisional. Ask the user "is the resistance about not wanting this, or not being able to do it yet?" and re-call to sharpen.
work_architectureNoEvidence that workflows, roles, decision rights and measures are ready. Pass only what is known. Explicit gaps and omitted checks block Accelerate until all four checks are evidenced.

Output Schema

ParametersJSON Schema
NameRequiredDescription
auditNoReproducibility record: engine version, the rules that fired, and the resolved inputs. Deterministic, no timestamps. If the verdict is challenged months later, the same inputs on the same engine version reproduce it exactly.
notesYesCaveats or context on the recommendation set.
feasibleYesWhether the target is reachable via the listed pillar moves.
feedbackNoOptional three-question feedback route, present only for Fix/Stop verdicts. The page records the response anonymously only when the user chooses an answer; no assessment data is attached.
bvf_versionYesAI BVF protocol version used.
change_planNoThe change-leader layer: a specific, sequenced route from Fix or Stop toward Go, aimed at the organisation. Present for Fix/Stop, absent when the initiative is already Accelerate. Present this to the user as the plan, not as raw data.
interpretationNoMeaning and limits of compatibility fields. Present these labels when explaining the result.
recommendationsYesPer-pillar improvement actions.
advisory_next_stepNoOptional CTA, present only for Fix/Stop verdicts.
target_classificationYesVerdict the recommendations aim to reach.
current_classificationYesVerdict as the initiative stands today.
projected_decision_confidenceYesProjected heuristic decision score, 0 to 100. The target remains conditional on work architecture and evidence review.

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations cover safety (readOnly, idempotent, non-destructive), and the description adds genuinely new behavioral facts they cannot convey: deterministic calculation, no authentication, anonymous telemetry with an AIBVF_TELEMETRY_DISABLE=1 opt-out, and the inference/provisional-marking behavior for omitted diagnostics.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core purpose and the gating call before any exclusion or parameter detail. It is long and repeats some schema-level routing facts, but nearly every sentence carries an instruction the agent needs.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 9-parameter nested-schema tool with an output schema present, this is complete: it covers the required context-match constraint, what may be copied from score_initiative, how omissions are handled, and leaves return format to the output schema.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% and the per-parameter descriptions are already rich, so the baseline is 3; the description nonetheless adds cross-parameter routing semantics (resistance_type selects will/skill, risk_type selects the regulatory/reputational/operational route, omissions stay provisional) that help an agent reason about which to supply.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a precise verb+resource+input condition+output contents: turn a Fix/Stop verdict into a change plan with pillar targets, named plays, owners, stop conditions, cost of waiting and a deadline. An agent can immediately tell this apart from score_initiative (which produces the verdict) and sequence_portfolio.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicit when-to-use ('after score_initiative returns Fix or Stop'), when-not-to ('do not use it to produce the initial verdict, sequence several initiatives or diagnose measured process waste'), and names the exact alternative for each excluded job—score_initiative, sequence_portfolio, diagnose_process—plus the Accelerate caveat.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

score_initiativeA
Read-onlyIdempotent
Inspect

Canonical-field scorer for one AI initiative. CALL THIS when industry, revenue_eur, function, ai_tier and readiness are already known, or when re-scoring with measured pillar evidence. For a proposal written in ordinary business language, call assess_ai_initiative first; it resolves these fields and asks for anything missing. Pillar scores remain optional: missing pillars are estimated deterministically, reported through pillar_basis, and reduce decision score, while a fully estimated pass can never return Accelerate. Returns Accelerate, Fix or Stop, modelled gross and readiness-adjusted EUR benefit ranges, decision score, sensitivity, assumptions and an audit trail. Use score_portfolio for several initiatives and diagnose_process for measured waste in an existing process. Deterministic calculation with no authentication. Anonymous usage telemetry may be sent; set AIBVF_TELEMETRY_DISABLE=1 to opt out.

ParametersJSON Schema
NameRequiredDescriptionDefault
scoresNoOPTIONAL, and each pillar inside it is optional. The four AI BVF pillars, each an honest 0–100 self-assessment, combining deterministically into the verdict: governance_risk ≥ 70 OR financial_return ≤ 20 returns Stop; strategic_alignment, financial_return and change_enablement all ≥ 60 with governance_risk ≤ 40 returns Accelerate; everything else returns Fix. Pass ONLY the pillars the user has real evidence for — do NOT invent numbers for the rest. Missing pillars are estimated deterministically by the engine from disclosed AI BVF planning assumptions, the response reports which via pillar_basis and scores_used, decision score is haircut by how much was estimated, and a fully-estimated pass can never return Accelerate (it returns Fix pending confirmation). So call immediately with whatever the user gave you, then ask for evidence on the estimated pillars and re-call to firm the verdict up.
ai_tierYesAmbition of the AI being deployed: gen1 = automation/RPA, gen2 = GenAI, gen3 = agentic. Interacts with readiness — a more ambitious tier running on lower readiness widens the pace-layer gap, which discounts the modelled EUR value even when the four pillar scores are strong.
functionYesBusiness function where the AI will operate, as one of the accepted enum values — selects which benchmark value drivers and rate ranges apply. Call list_taxonomy for the exact strings if unsure.
industryYesYour industry, as one of the accepted enum values — used to select the benchmark rate multiplier applied to the modelled EUR value. Call list_taxonomy for the exact strings if unsure.
readinessYesOrganisational readiness, honest self-assessment: agile = cross-functional, fast decisions; traditional = functional hierarchy; siloed = rigid, hand-off heavy. Sets the value-capture rate and, paired with ai_tier, the pace-layer drag — lower readiness against a higher tier reduces the captured value. Self-report is gameable: when the user has real process numbers, call infer_readiness first and pass its measured classification here instead.
revenue_eurYesApproximate annual revenue in EUR (must be ≥ 0). Scales the whole output: the disclosed AI BVF planning rates are applied as fractions of this figure, so the modelled EUR value range grows with it. A rough order-of-magnitude estimate is fine.
work_architectureNoEvidence that workflows, roles, decision rights and measures are ready. Pass only what is known. Explicit gaps and omitted checks block Accelerate until all four checks are evidenced.
signal_completenessNoOptional 0 to 1 input-quality factor. The default ranges from 0.5 when all pillars are estimated to 1 when all are supplied. Supplied values still require evidence review. Lower this factor when the supplied pillars rest on weak evidence.

Output Schema

ParametersJSON Schema
NameRequiredDescription
auditNoReproducibility record: engine version, the rules that fired, and the resolved inputs. Deterministic, no timestamps. If the verdict is challenged months later, the same inputs on the same engine version reproduce it exactly.
caveatNoPresent only when signal_completeness was low: warns the verdict rests on soft inputs and confidence was reduced.
reasonYesOne-line justification for the classification.
driversYesNamed value drivers behind the estimate.
feedbackNoOptional three-question feedback route, present only for Fix/Stop verdicts. The page records the response anonymously only when the user chooses an answer; no assessment data is attached.
bvf_versionYesAI BVF protocol version used.
multipliersYesFactors applied to the base rates.
scores_usedNoThe four pillar values the verdict was actually computed on, whether given by the caller or estimated by the engine. Show these to the user when any pillar was estimated.
sensitivityNoWhat moves this verdict, computed deterministically: the value if readiness were one notch worse, the value at revenue minus 20 percent, and the nearest single-pillar movements that flip the classification. Boards trust ranges with visible assumptions over point estimates; show this.
pillar_basisNoPer pillar: "given" (caller supplied it) or "estimated" (deterministic prior). When any pillar is estimated, tell the user which, and ask for evidence on those to firm up the verdict.
net_value_eurYesReadiness-adjusted benefit hypothesis before project build, run and change costs. Replace planning rates with scoped economics before funding.
classificationYesThe verdict for this initiative.
interpretationNoMeaning and limits of compatibility fields. Present these labels when explaining the result.
applied_modulesYesScoring-rule and sector-context labels. Sector labels do not certify clinical validation or regulatory compliance.
gross_value_eurYesModelled gross value in EUR before capture, low/high.
benchmark_sourceYesProvenance and evidence status for the AI BVF planning rates applied.
work_architectureYesThe work architecture gate across workflow, roles, human decision rights and performance measures. A stated gap or missing evidence blocks Accelerate.
advisory_next_stepNoOptional CTA, present only for Fix/Stop verdicts.
decision_confidenceYesHeuristic decision score, 0 to 100, adjusted for input quality. This score has no probability calibration.

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only declare readOnly/idempotent/non-destructive; the description adds substantially more: deterministic calculation with no authentication, anonymous telemetry with an AIBVF_TELEMETRY_DISABLE=1 opt-out, estimation behavior for missing pillars (reported via pillar_basis, decision score haircut, fully-estimated pass can never return Accelerate), and the returned artifacts. This is exactly the kind of context that goes beyond structured hints.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with purpose and the CALL THIS trigger, then alternatives, then return values and behavioral notes. It is dense and somewhat long, repeating pillar/estimation logic that the schema also covers, but nearly every sentence carries routing or behavioral information, so it is appropriately sized for a complex 8-param tool.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a high-complexity tool with nested objects, four enums, and a rich output schema, the description covers the trigger, the alternatives, the estimation fallback, the verdict logic, and the trust/telemetry model. Nothing an agent needs to call it correctly is missing; return details it summarizes are already backed by the output schema.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% and the schema is extremely rich (enum meanings, pillar thresholds), so the baseline is 3. The description adds genuine operational meaning on top: pass only pillars with real evidence, call immediately with partial data then re-call as evidence arrives, and the verdict-combination logic. It does not add per-parameter syntax beyond the schema, so it lands just above the baseline.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Opens with a specific verb+resource: "Canonical-field scorer for one AI initiative," and distinguishes itself from siblings by name (assess_ai_initiative, score_portfolio, diagnose_process). An agent can tell exactly what this tool produces (a verdict plus benefit ranges) versus the alternatives without opening any schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives an explicit trigger ("CALL THIS when industry, revenue_eur, function, ai_tier and readiness are already known, or when re-scoring with measured pillar evidence") and routes away to alternatives: assess_ai_initiative for ordinary business language, score_portfolio for several initiatives, diagnose_process for measured waste, infer_readiness when process numbers exist. When/when-not/alternatives are all present.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

score_portfolioA
Read-onlyIdempotent
Inspect

Score several AI initiatives as one AI BVF v1.0 portfolio and return the board-level position: counts of Accelerate / Fix / Stop, aggregate modelled EUR value range, mean decision score, the highest-value initiative, the highest-risk initiative, and every individual result. CALL THIS when the user has a portfolio document and needs to know what it contains before deciding funding or order, instead of looping score_initiative one initiative at a time. The single readiness value applies across every initiative: it changes capture rates and the pace-layer drag, so measure it with infer_readiness first when process data exists. The portfolio must carry organization.revenue_eur for EUR values; initiatives with missing revenue or invalid taxonomy are reported as skipped, never silently counted. Run validate_portfolio first only when the document shape is uncertain, then call sequence_portfolio when the verdicts need turning into a 90-day order. Deterministic calculation with no authentication. Anonymous usage telemetry may be sent; set AIBVF_TELEMETRY_DISABLE=1 to opt out.

ParametersJSON Schema
NameRequiredDescriptionDefault
portfolioYesAI BVF v1.0 portfolio with organization and initiatives. Each initiative carries id, name, function, ai_tier and four scores, supplied as numbers or { value, confidence } objects. Retain each initiative's work_architecture, pillar_basis and optional signal_completeness. pillar_basis marks given or estimated values; estimated pillars are recalculated for the current context and retain an input-quality reduction. Supplied values without provenance are caller-provided, with no claim of evidence verification. Accelerate requires the pillar thresholds and all four work-architecture checks. Missing work-design evidence returns Fix for an otherwise green case. organization.revenue_eur is required for benefit modelling. Validate unfamiliar documents first.
readinessYesOrganisational readiness applied to every initiative in the portfolio. Honest self-assessment: agile = cross-functional, fast decisions; traditional = functional hierarchy; siloed = rigid, hand-off heavy. The portfolio schema does not carry per-initiative readiness; this single value sets the capture rate for the whole portfolio and, paired with the ai_tier of each initiative, its pace-layer drag — lower readiness against a higher tier discounts the modelled EUR value.

Output Schema

ParametersJSON Schema
NameRequiredDescription
totalYesTotal initiatives in the portfolio (scored + skipped).
validYesTrue when the portfolio passed schema validation. False means no initiatives were scored.
summaryYes
feedbackNoOptional three-question feedback route, present only when any initiative was Fix or Stop. The page records the response anonymously only when the user chooses an answer; no assessment data is attached.
readinessYesReadiness value applied across all initiatives.
bvf_versionYesAI BVF protocol version used.
organizationYesEcho of the portfolio organisation fields applied to scoring.
interpretationNoMeaning and limits of compatibility fields. Present these labels when explaining the result.
validation_errorsNoEmpty when valid; otherwise one entry per schema violation.
advisory_next_stepNoOptional CTA, present only when any initiative was Fix or Stop.
scored_initiativesYesPer-initiative scoring result.
skipped_initiativesYesInitiatives that could not be scored, with the reason. Empty when all initiatives scored.
aggregate_net_value_eurYesArithmetic sum of benefit hypotheses. Reconcile overlapping scope, double counting and project costs before using this as a portfolio business case.
highest_risk_initiativeNoScored initiative most at risk: worst classification (Stop > Fix > Accelerate), tie-broken by lowest decision_confidence. Omitted when none were scored.
top_initiative_by_valueNoScored initiative with the highest mid-point readiness-adjusted EUR benefit. Omitted when none were scored.
mean_decision_confidenceYesMean decision score across scored initiatives (0–100); 0 when none were scored.

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With readOnlyHint=true and idempotentHint=true already declaring safety, the description still adds meaningful behavior: readiness 'changes capture rates and the pace-layer drag,' missing revenue or invalid taxonomy are 'reported as skipped, never silently counted,' and anonymous telemetry can be disabled via AIBVF_TELEMETRY_DISABLE=1. It does not spell out the full Accelerate/Fix/Stop decision thresholds, but the underlying rubric is beyond reasonable description scope.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the purpose and the output set, then execution guidance, then caveats and telemetry. Mostly dense and every sentence carries operational weight, though the telemetry sentence is the least relevant to invoking the tool correctly and could be trimmed.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a nested two-parameter portfolio tool with an output schema present and annotations covering safety, the description supplies the cross-cutting constraints an agent needs: pipeline ordering relative to infer_readiness/validate_portfolio/sequence_portfolio, the single readiness applying portfolio-wide, and what happens to bad data. An output schema exists so return values need not be explained. Only minor gaps (e.g., expected range or edge-case behavior of the skipped set) remain.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the parameter entries already document both fields in depth, establishing a baseline of 3. The description goes beyond that by emphasizing what is easy to miss operationally: 'The single readiness value applies across every initiative' and that organization.revenue_eur is required for EUR values, warning that missing revenue yields skipped rather than silent counts.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific verb and resource ('Score several AI initiatives as one AI BVF v1.0 portfolio') and immediately enumerates the exact outputs (counts of Accelerate/Fix/Stop, EUR range, mean score, top-value and top-risk initiative, individual results). It explicitly distinguishes itself from the sibling score_initiative by telling the agent to use this instead of looping that tool one initiative at a time.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives an explicit when-to-use trigger ('when the user has a portfolio document and needs to know what it contains before deciding funding or order'), names the anti-pattern (looping score_initiative), and routes to the right adjacent tools: infer_readiness first when process data exists, validate_portfolio only when the document shape is uncertain, and sequence_portfolio when verdicts need a 90-day order.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

sequence_portfolioA
Read-onlyIdempotent
Inspect

Turn a scored AI portfolio into three waves with gates over a configurable horizon, so the roadmap respects the change capacity of each business function. CALL THIS after score_portfolio when the user asks what to stop, fund first, defer or fit into the next 90 days. It does not change any verdict or re-score the business case. Stops enter wave 1 to reclaim budget and attention, quicker Accelerates enter wave 2, complex Accelerates and Fixes enter wave 3 behind their re-score gates. Pass the portfolio returned by score_portfolio directly through portfolio, or pass organization plus initiatives; both score shapes are accepted and nested values are flattened. readiness sets capture rates and pacing, max_parallel_per_function caps simultaneous change in one function per wave, and horizon_days divides the plan into three equal windows. Capacity overflow is reported as a conflict or a deferral beyond the horizon, never hidden. Run recommend_improvements for a Fix before treating its wave placement as permission to proceed. Deterministic calculation with no authentication. Anonymous usage telemetry may be sent; set AIBVF_TELEMETRY_DISABLE=1 to opt out.

ParametersJSON Schema
NameRequiredDescriptionDefault
portfolioNoAlternative input: the same AI BVF v1.0 portfolio document score_portfolio accepts (organization + initiatives with nested {value} pillar scores). Pass either this OR the top-level organization + initiatives; nested score values are flattened automatically, and missing pillars are estimated honestly.
readinessYesOrganisational readiness applied across the portfolio; sets capture rates and pacing. Measure it with infer_readiness when process numbers exist.
constraintsNoChange-capacity constraints. The defaults encode the core principle: no function absorbs unlimited concurrent change.
initiativesNoThe portfolio to sequence. Each initiative carries flat 0-100 pillar numbers (not the nested value objects of the portfolio wire format).
organizationNoOrganisation context used when initiatives are passed at the top level. Required with top-level initiatives and ignored when portfolio is supplied.

Output Schema

ParametersJSON Schema
NameRequiredDescription
auditYesReproducibility record: engine version, the rules that fired, and the resolved inputs. Deterministic, no timestamps. If the verdict is challenged months later, the same inputs on the same engine version reproduce it exactly.
wavesYesThree waves with named gates: Stops first (free the budget), quick Accelerates second (buy trust), complex Accelerates plus Fixes third (spend the trust). Present this to the user as the rollout plan.
totalsYesCounts: stopped, quick_wins, complex_or_fix, deferred.
skippedNo
bvf_versionYes
capacity_conflictsYesWhere more initiatives land on one function than it can absorb per wave, with the deferral applied. Surface these: an overloaded function is how good portfolios fail.
sequencing_principlesYes
deferred_beyond_horizonNoInitiatives that did not fit the horizon under the capacity constraint; they need their own decision.
aggregate_accelerate_value_eurNoSum of modelled net EUR for the sequenced Accelerates, low and high.

TDQS

A5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnly/idempotent/non-destructive, and the description adds substantially more: it explains that no verdict is changed or re-scored, that capacity overflow is surfaced as a conflict or deferral 'never hidden', and that the computation is deterministic with no authentication plus a telemetry opt-out variable (AIBVF_TELEMETRY_DISABLE=1).

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

It is long but front-loaded: the purpose and the call-timing trigger come first, then placement rules, then input mechanics, then caveats. Each sentence carries distinct, non-redundant information, so the length is earned rather than padding.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 5-parameter, nested, dual-input-path tool with a rich output schema, the description covers purpose, invocation ordering, input variants, assignment heuristics, constraint behavior, failure reporting and telemetry. Since an output schema exists, it correctly does not spend words on return-value shape.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Although schema coverage is 100%, the description adds semantics beyond the schema: it describes the wave-placement rules (stops to wave 1, quicker Accelerates to wave 2, complex Accelerates and Fixes to wave 3 behind gates), the dual portfolio/organization+initiatives input paths with flattening, and what readiness, max_parallel_per_function and horizon_days actually control.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The opening sentence names a specific verb and resource ('Turn a scored AI portfolio into three waves with gates over a configurable horizon') and states the goal of respecting change capacity per function. It clearly distinguishes itself from siblings by positioning downstream of score_portfolio and upstream of per-initiative Fix work.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly states 'CALL THIS after score_portfolio when the user asks what to stop, fund first, defer or fit into the next 90 days', giving both ordering and trigger conditions. It also names a prerequisite for a sibling ('Run recommend_improvements for a Fix before treating its wave placement as permission to proceed'), leaving nothing to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

validate_portfolioA
Read-onlyIdempotent
Inspect

Check whether a supplied AI BVF v1.0 portfolio document has the shape the portfolio tools require, before scoring, sequencing, storing or sharing it. CALL THIS when the document came from a file, another system or hand-built JSON and its structure is uncertain. It checks required fields, taxonomy values and 0–100 pillar ranges only; it does not judge the evidence or calculate a verdict. Pillars may be bare numbers or { value, confidence } objects, both are valid. Use assemble_portfolio when the user has a list of initiatives in conversation and needs the document built for them, score_portfolio when the document is already ready for verdicts, and sequence_portfolio only after its initiatives are scoreable. Returns valid=true or one error per failing JSON path. Deterministic validation with no authentication. Anonymous usage telemetry may be sent; set AIBVF_TELEMETRY_DISABLE=1 to opt out.

ParametersJSON Schema
NameRequiredDescriptionDefault
portfolioYesThe portfolio document as a JSON object following the AI BVF v1.0 schema: a top-level object with bvf_version, organization, and a non-empty "initiatives" array, each initiative carrying the same fields score_initiative expects (industry, revenue_eur, function, ai_tier, readiness, and a scores object with the four 0–100 pillars, each either a bare number or an object { value, confidence? }; both shapes pass). Checked structurally only — required fields present, correct types, enum values valid, pillar numbers in range; the pillar values are NOT scored or judged here (use score_initiative or score_portfolio for that). On failure, errors[] names each failing JSON path and the rule it broke.

Output Schema

ParametersJSON Schema
NameRequiredDescription
validYesTrue when the portfolio conforms to the schema.
errorsYesEmpty when valid; otherwise one entry per schema violation.
bvf_versionYesAI BVF protocol version validated against.

TDQS

A4.6/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover the safety profile (readOnly, idempotent, non-destructive, closed-world), but the description adds real operational context the annotations cannot: deterministic validation, no authentication required, the exact failure output shape ('valid=true or one error per failing JSON path'), and a telemetry notice with an opt-out env var. It stops short of documenting any size/time limits on large portfolios.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Dense but well ordered: what it does, when to call it, what it does not do, tolerated input shapes, sibling routing, then return/telemetry details. It is long for a single-parameter validator, but nearly every sentence carries a distinct fact, with only minor overlap between the description body and the schema prose.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, yet the description still tells the agent the shape of success and failure results (valid=true vs. one error per failing JSON path), which is exactly what a validation tool's caller needs to branch on. Combined with the enum/range scope statement and the telemetry disclosure, nothing material is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and there is only one parameter, so the schema carries the shape documentation; the baseline is 3. The description does add outcome-relevant semantics beyond raw shape by stating that both pillar encodings ('bare numbers or { value, confidence } objects') are accepted, which tells the caller what will and won't fail validation rather than just what the field looks like.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource (validate a BVF v1.0 portfolio document) and immediately scopes it: 'it checks required fields, taxonomy values and 0–100 pillar ranges only; it does not judge the evidence or calculate a verdict.' That negative scoping cleanly separates it from score_portfolio and score_initiative.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicit trigger ('CALL THIS when the document came from a file, another system or hand-built JSON and its structure is uncertain') plus named alternatives with their selecting conditions: assemble_portfolio for in-conversation lists, score_portfolio when ready for verdicts, sequence_portfolio only after scoreable. Nothing is left to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 12 tool updatesv0.14.14
    • Addedassemble_portfolio
    • Addedassess_ai_initiative
    • Changedcalculate_pace_layer_drag4 fields changed
      • changedInput schema / properties / ai_tier / description
        Previous value: -"Ambition of the AI being deployed: gen1=automation/RPA, gen2=GenAI, gen3=agentic."New value: +"Ambition of the AI operating model: gen1 = automation/RPA, gen2 = GenAI, gen3 = agentic. Paired with readiness to set pace_gap severity — gen3 on any readiness below agile, or gen2 on siloed, is severe; a higher tier against a slower operating model widens the gap and raises the drag."
      • changedInput schema / properties / industry / description
        Previous value: -"Optional; defaults to universal. Reserved for future vertical adjustments."New value: +"Optional; defaults to universal if omitted. Reserved for future vertical drag-rate adjustments — does not change the result today. Call list_taxonomy for accepted values."
      • changedInput schema / properties / readiness / description
        Previous value: -"Organisational readiness, honest self-assessment: agile = cross-functional, fast decisions; traditional = functional hierarchy; siloed = rigid, hand-off heavy."New value: +"Organisational readiness, honest self-assessment: agile = cross-functional, fast decisions; traditional = functional hierarchy; siloed = rigid, hand-off heavy. Agile readiness yields minimal drag at any tier; the mismatch between a fast AI tier and a slower operating model is what generates the Organisational Drag Cost."
      • changedInput schema / properties / revenue_eur / description
        Previous value: -"Approximate annual revenue in EUR."New value: +"Approximate annual revenue in EUR (must be ≥ 0). The result scales with this: annual_drag_eur is returned as an absolute range and as drag_rate, a fraction of this revenue (e.g. 0.02 = 2%)."
    • Changeddiagnose_process11 fields changed
      • changedInput schema / properties / automation_level / description
        Previous value: -"Share already automated (0–1)."New value: +"Share already automated (0–1). Low automation makes manual effort the dominant drag and selects Automate; the un-automated remainder is the addressable share."
      • changedInput schema / properties / cycle_time_days / description
        Previous value: -"Median wall-clock days per instance, end to end."New value: +"Median wall-clock days per instance, end to end. Long cycles relative to touch-time signal wait/latency drag."
      • changedInput schema / properties / direct_spend_eur / description
        Previous value: -"Annual licence/vendor/tooling spend on the process in EUR."New value: +"Annual licence/vendor/tooling spend on the process in EUR. Added to the labour baseline and shifts how much of the saving is labour- vs spend-addressable."
      • changedInput schema / properties / fte_hours_per_instance / description
        Previous value: -"Human touch-time in hours per instance."New value: +"Human touch-time in hours per instance. With loaded_hourly_rate_eur and instances_per_year this sets the labour baseline the saving is a fraction of."
      • changedInput schema / properties / handoffs / description
        Previous value: -"Distinct owners/systems an instance passes through."New value: +"Distinct owners/systems an instance passes through. Weighed against the per-function median; many handoffs make handoff drag dominant and point to Consolidate & re-sequence."
      • changedInput schema / properties / instances_per_year / description
        Previous value: -"Process volume: how many times it runs per year."New value: +"Process volume: how many times it runs per year. Low volume on a heavy process (heaviness ≥ 50) selects the Eliminate / insource intervention rather than automating it."
      • changedInput schema / properties / loaded_hourly_rate_eur / description
        Previous value: -"Fully-loaded labour cost per hour in EUR."New value: +"Fully-loaded labour cost per hour in EUR (salary + on-costs). Multiplies fte_hours_per_instance × instances_per_year into the annual labour baseline."
      • changedInput schema / properties / readiness / description
        Previous value: -"Optional. Org change-absorption capacity (caps realised saving). Defaults to traditional."New value: +"Optional. Org change-absorption capacity — agile / traditional / siloed — which caps the realised (net) saving below the gross potential. Defaults to traditional."
      • changedInput schema / properties / rework_rate / description
        Previous value: -"Fraction of instances reopened/reworked (0–1)."New value: +"Fraction of instances reopened/reworked (0–1). When rework is the dominant drag factor the intervention becomes Quality controls, and it also sets the addressable share for that path."
      • changedInput schema / properties / signal_completeness / description
        Previous value: -"Optional. How much of the above was measured vs defaulted (0–1). Governs confidence; defaults to 0.7."New value: +"Optional 0–1. How much of the above was measured versus defaulted. Governs decision_confidence proportionally — lower it when you estimated inputs so the verdict stays honest. Defaults to 0.7."
      • changedInput schema / properties / touch_ratio / description
        Previous value: -"Touch-time ÷ cycle-time (0–1). The remainder is wait."New value: +"Touch-time ÷ cycle-time (0–1). The remainder is wait; a low value means the process is mostly waiting, which pushes the intervention toward Consolidate & re-sequence."
    • Changedget_benchmark2 fields changed
      • changedInput schema / properties / function / description
        Previous value: -"Business function to benchmark. Must be one of the list_taxonomy function values."New value: +"Business function to benchmark — must be one of the list_taxonomy function values. Selects the base revenue-uplift and cost-reduction rate ranges (returned as fractions of revenue) and the value drivers."
      • changedInput schema / properties / industry / description
        Previous value: -"Industry whose multiplier to apply. Must be one of the list_taxonomy industry values."New value: +"Industry whose multiplier to apply — must be one of the list_taxonomy industry values. The returned industry_multiplier is applied to the function base rates; pass \"universal\" for the un-adjusted rates."
    • Addedinfer_readiness
    • Addedmap_to_taxonomy
    • Changedrecommend_improvements21 fields changed
      • addedInput schema / description
        Added value: +"Inputs for a change plan after a Fix or Stop verdict. industry, revenue_eur, function, ai_tier and readiness must match the scoring call so the plan is built against the same case. scores and work_architecture may be copied from score_initiative; omitted pillars are estimated and make the plan provisional. resistance_type and risk_type are optional diagnostics that select the play, and omission triggers a named inference plus the next question to ask."
      • changedInput schema / properties / ai_tier / description
        Previous value: -"gen1=automation/RPA, gen2=GenAI, gen3=agentic."New value: +"Ambition of the AI being deployed: gen1 = automation/RPA, gen2 = GenAI, gen3 = agentic. Interacts with readiness — a more ambitious tier running on lower readiness widens the pace-layer gap, which discounts the modelled EUR value even when the four pillar scores are strong."
      • changedInput schema / properties / function / description
        Previous value: -"Business function where the AI will operate."New value: +"Business function where the AI will operate, as one of the accepted enum values — selects which benchmark value drivers and rate ranges apply. Call list_taxonomy for the exact strings if unsure."
      • changedInput schema / properties / industry / description
        Previous value: -"Your industry. See list_taxonomy if unsure."New value: +"Your industry, as one of the accepted enum values — used to select the benchmark rate multiplier applied to the modelled EUR value. Call list_taxonomy for the exact strings if unsure."
      • changedInput schema / properties / readiness / description
        Previous value: -"Organisational readiness. Honest self-assessment."New value: +"Organisational readiness, honest self-assessment: agile = cross-functional, fast decisions; traditional = functional hierarchy; siloed = rigid, hand-off heavy. Sets the value-capture rate and, paired with ai_tier, the pace-layer drag — lower readiness against a higher tier reduces the captured value. Self-report is gameable: when the user has real process numbers, call infer_readiness first and pass its measured classification here instead."
      • addedInput schema / properties / resistance_type
        Added value: +{
        +  "description": "Optional. What sits behind a low change-enablement score: \"will\" = people do not want the change (power shifts, fear, no case for change), \"skill\" = people cannot yet do it (capability and capacity gap). Selects between a coalition-building play (Kotter 1-2 + ADKAR Awareness/Desire) and an owner-and-capability play (ADKAR Knowledge/Ability). If you do not know, omit it: the engine infers from readiness (agile infers skill, traditional/siloed infers will) and marks the play provisional. Ask the user \"is the resistance about not wanting this, or not being able to do it yet?\" and re-call to sharpen.",
        +  "enum": [
        +    "will",
        +    "skill"
        +  ],
        +  "type": "string"
        +}
      • changedInput schema / properties / revenue_eur / description
        Previous value: -"Approximate annual revenue in EUR."New value: +"Approximate annual revenue in EUR (must be ≥ 0). Scales the whole output: the disclosed AI BVF planning rates are applied as fractions of this figure, so the modelled EUR value range grows with it. A rough order-of-magnitude estimate is fine."
      • addedInput schema / properties / risk_type
        Added value: +{
        +  "description": "Optional. The nature of a high governance-risk score: \"regulatory\" = statute applies (EU AI Act, GDPR Article 22, DORA), \"reputational\" = the risk is how failure looks and lands publicly, \"operational\" = the system failing quietly inside a process. Selects between a regulatory remediation sequence, visible trust guardrails, and a proportionate governance review. If you do not know, omit it: the engine infers (gen3 tier, or a regulated function/industry, infers regulatory) and marks the play provisional.",
        +  "enum": [
        +    "regulatory",
        +    "reputational",
        +    "operational"
        +  ],
        +  "type": "string"
        +}
      • addedInput schema / properties / scores / description
        Added value: +"OPTIONAL, and each pillar inside it is optional. The four AI BVF pillars, each an honest 0–100 self-assessment, combining deterministically into the verdict: governance_risk ≥ 70 OR financial_return ≤ 20 returns Stop; strategic_alignment, financial_return and change_enablement all ≥ 60 with governance_risk ≤ 40 returns Accelerate; everything else returns Fix. Pass ONLY the pillars the user has real evidence for — do NOT invent numbers for the rest. Missing pillars are estimated deterministically by the engine from disclosed AI BVF planning assumptions, the response reports which via pillar_basis and scores_used, decision score is haircut by how much was estimated, and a fully-estimated pass can never return Accelerate (it returns Fix pending confirmation). So call immediately with whatever the user gave you, then ask for evidence on the estimated pillars and re-call to firm the verdict up."
      • changedInput schema / properties / scores / properties / change_enablement / description
        Previous value: -"Sponsor, owner, funded change budget (0-100)."New value: +"Optional; when omitted, estimated from readiness (agile 55, traditional 45, siloed 32 — always below the 60 floor, because an unevidenced change capability is unproven). Sponsor in place, owner named, change budget funded (0–100, higher is better). Must be ≥ 60 for an Accelerate verdict."
      • changedInput schema / properties / scores / properties / financial_return / description
        Previous value: -"Strength of modelled return (0-100)."New value: +"Optional; when omitted, estimated from the disclosed AI BVF planning range for the function (40–52, never enough to clear 60 unmodelled, never low enough to force a Stop). Strength of the modelled return (0–100, higher is better). A value ≤ 20 forces a Stop on its own; ≥ 60 is one of the four conditions required for Accelerate."
      • changedInput schema / properties / scores / properties / governance_risk / description
        Previous value: -"Regulatory / reputational exposure. Higher = more risk (0-100)."New value: +"Optional; when omitted, estimated from tier and regulated context (gen1 30 / gen2 42 / gen3 55, +10 in a regulated function, +8 in a regulated industry — agentic AI in regulated finance estimates at 73 and forces a Stop until governance evidence exists). This pillar is INVERTED: higher means MORE risk. ≥ 70 forces a Stop on its own; must be ≤ 40 for Accelerate."
      • changedInput schema / properties / scores / properties / strategic_alignment / description
        Previous value: -"How clearly this moves a board-level KPI (0-100)."New value: +"Optional; estimated at 50 (unproven) when omitted, since alignment to a board KPI cannot be read from context. How clearly this moves a board-level KPI (0–100, higher is better). Must be ≥ 60 — together with financial_return ≥ 60, change_enablement ≥ 60 and governance_risk ≤ 40 — for an Accelerate verdict."
      • removedInput schema / properties / scores / required
        Removed value: -[
        -  "strategic_alignment",
        -  "financial_return",
        -  "change_enablement",
        -  "governance_risk"
        -]
      • addedInput schema / properties / work_architecture
        Added value: +{
        +  "description": "Evidence that workflows, roles, decision rights and measures are ready. Pass only what is known. Explicit gaps and omitted checks block Accelerate until all four checks are evidenced.",
        +  "properties": {
        +    "decision_rights_defined": {
        +      "description": "True only when decision, override and escalation rights have named human owners, false when authority remains unclear.",
        +      "type": "boolean"
        +    },
        +    "measures_updated": {
        +      "description": "True only when performance measures and incentives reflect the redesigned work, false when the old measures remain.",
        +      "type": "boolean"
        +    },
        +    "roles_redesigned": {
        +      "description": "True only when affected roles, accountabilities and capability expectations have been rewritten, false when roles remain unchanged.",
        +      "type": "boolean"
        +    },
        +    "workflow_redesigned": {
        +      "description": "True only when the end-to-end workflow has been redesigned around the AI and retained human judgement, false when the existing workflow remains.",
        +      "type": "boolean"
        +    }
        +  },
        +  "type": "object"
        +}
      • changedInput schema / required
        Previous value: -[
        -  "industry",
        -  "revenue_eur",
        -  "function",
        -  "ai_tier",
        -  "readiness",
        -  "scores"
        -]New value: +[
        +  "industry",
        +  "revenue_eur",
        +  "function",
        +  "ai_tier",
        +  "readiness"
        +]
      • addedOutput schema / properties / audit
        Added value: +{
        +  "description": "Reproducibility record: engine version, the rules that fired, and the resolved inputs. Deterministic, no timestamps. If the verdict is challenged months later, the same inputs on the same engine version reproduce it exactly.",
        +  "properties": {
        +    "bvf_version": {
        +      "type": "string"
        +    },
        +    "engine": {
        +      "type": "string"
        +    },
        +    "engine_version": {
        +      "type": "string"
        +    },
        +    "inputs_used": {
        +      "description": "The resolved inputs the result was computed on, including estimated pillar values.",
        +      "type": "object"
        +    },
        +    "note": {
        +      "type": "string"
        +    },
        +    "rules_fired": {
        +      "description": "The rules that actually fired, in order: estimation, gates, classification, value arithmetic.",
        +      "items": {
        +        "type": "string"
        +      },
        +      "type": "array"
        +    }
        +  },
        +  "type": "object"
        +}
      • addedOutput schema / properties / change_plan
        Added value: +{
        +  "description": "The change-leader layer: a specific, sequenced route from Fix or Stop toward Go, aimed at the organisation. Present for Fix/Stop, absent when the initiative is already Accelerate. Present this to the user as the plan, not as raw data.",
        +  "properties": {
        +    "binding_constraint": {
        +      "description": "The one thing standing between this initiative and a Go. Lead with this; a single named blocker gets acted on where a list of four gets skimmed.",
        +      "type": "string"
        +    },
        +    "cost_of_waiting": {
        +      "description": "The cost of delay framed as a decision rule: fix if the plays cost less than the waiting, stop if they cost more.",
        +      "type": "string"
        +    },
        +    "cost_of_waiting_eur": {
        +      "description": "Estimated organisational drag over the plan window, {low, high} in EUR, from the pace-layer model. The price of sitting in Fix.",
        +      "type": "object"
        +    },
        +    "honest_stop": {
        +      "description": "Present when the truthful call is Stop rather than Fix. Surface this verbatim; it is the most valuable sentence in the response when it appears.",
        +      "type": "string"
        +    },
        +    "plays": {
        +      "description": "Named change plays, worst pillar first, each selected from the failing pillar AND the organisational context. Every play works two altitudes: the organisation (Kotter) and the person (Prosci ADKAR).",
        +      "items": {
        +        "properties": {
        +          "diagnosis": {
        +            "description": "What is actually blocking, in plain language.",
        +            "type": "string"
        +          },
        +          "diagnostic_questions": {
        +            "description": "Questions to put to the organisation to sharpen or challenge the play. Ask these before executing.",
        +            "items": {
        +              "type": "string"
        +            },
        +            "type": "array"
        +          },
        +          "id": {
        +            "description": "Play identifier, e.g. coalition-first, regulatory-remediation, value-rescope.",
        +            "type": "string"
        +          },
        +          "org_move": {
        +            "description": "The organisation-level move (method + action), typically a Kotter step.",
        +            "type": "object"
        +          },
        +          "owner": {
        +            "description": "The role that owns the play, to be filled with a named individual.",
        +            "type": "string"
        +          },
        +          "person_move": {
        +            "description": "The individual-level move (method + action), typically an ADKAR stage.",
        +            "type": "object"
        +          },
        +          "pillar": {
        +            "description": "The pillar this play repairs, or pace_gap for the cross-cutting operating-model play.",
        +            "type": "string"
        +          },
        +          "provisional": {
        +            "description": "True when the play was inferred from readiness/tier/function rather than told via resistance_type or risk_type. When true, ask the diagnostic questions and re-call with the answer to sharpen the plan.",
        +            "type": "boolean"
        +          },
        +          "source": {
        +            "description": "The named method behind the play: Kotter, Prosci ADKAR, EU AI Act, benchmark sources.",
        +            "type": "string"
        +          },
        +          "steps": {
        +            "description": "Sequenced actions, in order. Order matters: e.g. desire before change budget.",
        +            "items": {
        +              "type": "string"
        +            },
        +            "type": "array"
        +          },
        +          "stop_condition": {
        +            "description": "Present when the honest escalation from this play is Stop, and the condition that triggers it.",
        +            "type": "string"
        +          },
        +          "timeline_weeks": {
        +            "description": "Expected duration range in weeks, [low, high].",
        +            "items": {
        +              "type": "number"
        +            },
        +            "type": "array"
        +          }
        +        },
        +        "type": "object"
        +      },
        +      "type": "array"
        +    },
        +    "position": {
        +      "description": "Where this Fix sits between Go and Stop. near_go = a funding decision waiting on evidence; near_stop = one adverse finding from Stop, run only the first play then re-score.",
        +      "enum": [
        +        "near_go",
        +        "contested",
        +        "near_stop"
        +      ],
        +      "type": "string"
        +    },
        +    "position_detail": {
        +      "description": "One-paragraph read of the position, written for the organisation.",
        +      "type": "string"
        +    },
        +    "rescore_gate": {
        +      "description": "What must be true, and by when, for the re-score to arbitrate. Fix is a decision with a deadline, not a limbo state.",
        +      "type": "object"
        +    }
        +  },
        +  "type": "object"
        +}
      • addedOutput schema / properties / feedback
        Added value: +{
        +  "description": "Optional three-question feedback route, present only for Fix/Stop verdicts. The page records the response anonymously only when the user chooses an answer; no assessment data is attached.",
        +  "properties": {
        +    "question": {
        +      "type": "string"
        +    },
        +    "url": {
        +      "format": "uri",
        +      "type": "string"
        +    }
        +  },
        +  "required": [
        +    "question",
        +    "url"
        +  ],
        +  "type": "object"
        +}
      • addedOutput schema / properties / interpretation
        Added value: +{
        +  "additionalProperties": {
        +    "type": "string"
        +  },
        +  "description": "Meaning and limits of compatibility fields. Present these labels when explaining the result.",
        +  "type": "object"
        +}
      • changedOutput schema / properties / projected_decision_confidence / description
        Previous value: -"Confidence in the verdict if the recommendations land, 0-100."New value: +"Projected heuristic decision score, 0 to 100. The target remains conditional on work architecture and evidence review."
    • Changedscore_initiative26 fields changed
      • changedInput schema / properties / ai_tier / description
        Previous value: -"gen1=automation/RPA, gen2=GenAI, gen3=agentic."New value: +"Ambition of the AI being deployed: gen1 = automation/RPA, gen2 = GenAI, gen3 = agentic. Interacts with readiness — a more ambitious tier running on lower readiness widens the pace-layer gap, which discounts the modelled EUR value even when the four pillar scores are strong."
      • changedInput schema / properties / function / description
        Previous value: -"Business function where the AI will operate."New value: +"Business function where the AI will operate, as one of the accepted enum values — selects which benchmark value drivers and rate ranges apply. Call list_taxonomy for the exact strings if unsure."
      • changedInput schema / properties / industry / description
        Previous value: -"Your industry. See list_taxonomy if unsure."New value: +"Your industry, as one of the accepted enum values — used to select the benchmark rate multiplier applied to the modelled EUR value. Call list_taxonomy for the exact strings if unsure."
      • changedInput schema / properties / readiness / description
        Previous value: -"Organisational readiness. Honest self-assessment."New value: +"Organisational readiness, honest self-assessment: agile = cross-functional, fast decisions; traditional = functional hierarchy; siloed = rigid, hand-off heavy. Sets the value-capture rate and, paired with ai_tier, the pace-layer drag — lower readiness against a higher tier reduces the captured value. Self-report is gameable: when the user has real process numbers, call infer_readiness first and pass its measured classification here instead."
      • changedInput schema / properties / revenue_eur / description
        Previous value: -"Approximate annual revenue in EUR."New value: +"Approximate annual revenue in EUR (must be ≥ 0). Scales the whole output: the disclosed AI BVF planning rates are applied as fractions of this figure, so the modelled EUR value range grows with it. A rough order-of-magnitude estimate is fine."
      • addedInput schema / properties / scores / description
        Added value: +"OPTIONAL, and each pillar inside it is optional. The four AI BVF pillars, each an honest 0–100 self-assessment, combining deterministically into the verdict: governance_risk ≥ 70 OR financial_return ≤ 20 returns Stop; strategic_alignment, financial_return and change_enablement all ≥ 60 with governance_risk ≤ 40 returns Accelerate; everything else returns Fix. Pass ONLY the pillars the user has real evidence for — do NOT invent numbers for the rest. Missing pillars are estimated deterministically by the engine from disclosed AI BVF planning assumptions, the response reports which via pillar_basis and scores_used, decision score is haircut by how much was estimated, and a fully-estimated pass can never return Accelerate (it returns Fix pending confirmation). So call immediately with whatever the user gave you, then ask for evidence on the estimated pillars and re-call to firm the verdict up."
      • changedInput schema / properties / scores / properties / change_enablement / description
        Previous value: -"Sponsor, owner, funded change budget (0-100)."New value: +"Optional; when omitted, estimated from readiness (agile 55, traditional 45, siloed 32 — always below the 60 floor, because an unevidenced change capability is unproven). Sponsor in place, owner named, change budget funded (0–100, higher is better). Must be ≥ 60 for an Accelerate verdict."
      • changedInput schema / properties / scores / properties / financial_return / description
        Previous value: -"Strength of modelled return (0-100)."New value: +"Optional; when omitted, estimated from the disclosed AI BVF planning range for the function (40–52, never enough to clear 60 unmodelled, never low enough to force a Stop). Strength of the modelled return (0–100, higher is better). A value ≤ 20 forces a Stop on its own; ≥ 60 is one of the four conditions required for Accelerate."
      • changedInput schema / properties / scores / properties / governance_risk / description
        Previous value: -"Regulatory / reputational exposure. Higher = more risk (0-100)."New value: +"Optional; when omitted, estimated from tier and regulated context (gen1 30 / gen2 42 / gen3 55, +10 in a regulated function, +8 in a regulated industry — agentic AI in regulated finance estimates at 73 and forces a Stop until governance evidence exists). This pillar is INVERTED: higher means MORE risk. ≥ 70 forces a Stop on its own; must be ≤ 40 for Accelerate."
      • changedInput schema / properties / scores / properties / strategic_alignment / description
        Previous value: -"How clearly this moves a board-level KPI (0-100)."New value: +"Optional; estimated at 50 (unproven) when omitted, since alignment to a board KPI cannot be read from context. How clearly this moves a board-level KPI (0–100, higher is better). Must be ≥ 60 — together with financial_return ≥ 60, change_enablement ≥ 60 and governance_risk ≤ 40 — for an Accelerate verdict."
      • removedInput schema / properties / scores / required
        Removed value: -[
        -  "strategic_alignment",
        -  "financial_return",
        -  "change_enablement",
        -  "governance_risk"
        -]
      • changedInput schema / properties / signal_completeness / description
        Previous value: -"Optional 0–1. How grounded the four pillar scores are in real evidence versus estimated from context. Defaults to 1 (treated as measured). If the organisation lacks formal change-readiness or risk metadata, estimate the pillars from what you know AND set this lower to say so — decision confidence is reduced proportionally and a caveat is attached, instead of returning a falsely confident verdict on soft inputs."New value: +"Optional 0 to 1 input-quality factor. The default ranges from 0.5 when all pillars are estimated to 1 when all are supplied. Supplied values still require evidence review. Lower this factor when the supplied pillars rest on weak evidence."
      • addedInput schema / properties / work_architecture
        Added value: +{
        +  "description": "Evidence that workflows, roles, decision rights and measures are ready. Pass only what is known. Explicit gaps and omitted checks block Accelerate until all four checks are evidenced.",
        +  "properties": {
        +    "decision_rights_defined": {
        +      "description": "True only when decision, override and escalation rights have named human owners, false when authority remains unclear.",
        +      "type": "boolean"
        +    },
        +    "measures_updated": {
        +      "description": "True only when performance measures and incentives reflect the redesigned work, false when the old measures remain.",
        +      "type": "boolean"
        +    },
        +    "roles_redesigned": {
        +      "description": "True only when affected roles, accountabilities and capability expectations have been rewritten, false when roles remain unchanged.",
        +      "type": "boolean"
        +    },
        +    "workflow_redesigned": {
        +      "description": "True only when the end-to-end workflow has been redesigned around the AI and retained human judgement, false when the existing workflow remains.",
        +      "type": "boolean"
        +    }
        +  },
        +  "type": "object"
        +}
      • changedInput schema / required
        Previous value: -[
        -  "industry",
        -  "revenue_eur",
        -  "function",
        -  "ai_tier",
        -  "readiness",
        -  "scores"
        -]New value: +[
        +  "industry",
        +  "revenue_eur",
        +  "function",
        +  "ai_tier",
        +  "readiness"
        +]
      • changedOutput schema / properties / applied_modules / description
        Previous value: -"BVF scoring modules that fired for this input."New value: +"Scoring-rule and sector-context labels. Sector labels do not certify clinical validation or regulatory compliance."
      • addedOutput schema / properties / audit
        Added value: +{
        +  "description": "Reproducibility record: engine version, the rules that fired, and the resolved inputs. Deterministic, no timestamps. If the verdict is challenged months later, the same inputs on the same engine version reproduce it exactly.",
        +  "properties": {
        +    "bvf_version": {
        +      "type": "string"
        +    },
        +    "engine": {
        +      "type": "string"
        +    },
        +    "engine_version": {
        +      "type": "string"
        +    },
        +    "inputs_used": {
        +      "description": "The resolved inputs the result was computed on, including estimated pillar values.",
        +      "type": "object"
        +    },
        +    "note": {
        +      "type": "string"
        +    },
        +    "rules_fired": {
        +      "description": "The rules that actually fired, in order: estimation, gates, classification, value arithmetic.",
        +      "items": {
        +        "type": "string"
        +      },
        +      "type": "array"
        +    }
        +  },
        +  "type": "object"
        +}
      • changedOutput schema / properties / benchmark_source / description
        Previous value: -"Citation for the benchmark rates applied."New value: +"Provenance and evidence status for the AI BVF planning rates applied."
      • changedOutput schema / properties / decision_confidence / description
        Previous value: -"Confidence in the verdict, 0-100."New value: +"Heuristic decision score, 0 to 100, adjusted for input quality. This score has no probability calibration."
      • addedOutput schema / properties / feedback
        Added value: +{
        +  "description": "Optional three-question feedback route, present only for Fix/Stop verdicts. The page records the response anonymously only when the user chooses an answer; no assessment data is attached.",
        +  "properties": {
        +    "question": {
        +      "type": "string"
        +    },
        +    "url": {
        +      "format": "uri",
        +      "type": "string"
        +    }
        +  },
        +  "required": [
        +    "question",
        +    "url"
        +  ],
        +  "type": "object"
        +}
      • addedOutput schema / properties / interpretation
        Added value: +{
        +  "additionalProperties": {
        +    "type": "string"
        +  },
        +  "description": "Meaning and limits of compatibility fields. Present these labels when explaining the result.",
        +  "type": "object"
        +}
      • changedOutput schema / properties / net_value_eur / description
        Previous value: -"Modelled net value in EUR after capture rate, low/high."New value: +"Readiness-adjusted benefit hypothesis before project build, run and change costs. Replace planning rates with scoped economics before funding."
      • addedOutput schema / properties / pillar_basis
        Added value: +{
        +  "description": "Per pillar: \"given\" (caller supplied it) or \"estimated\" (deterministic prior). When any pillar is estimated, tell the user which, and ask for evidence on those to firm up the verdict.",
        +  "properties": {
        +    "change_enablement": {
        +      "enum": [
        +        "given",
        +        "estimated"
        +      ],
        +      "type": "string"
        +    },
        +    "financial_return": {
        +      "enum": [
        +        "given",
        +        "estimated"
        +      ],
        +      "type": "string"
        +    },
        +    "governance_risk": {
        +      "enum": [
        +        "given",
        +        "estimated"
        +      ],
        +      "type": "string"
        +    },
        +    "strategic_alignment": {
        +      "enum": [
        +        "given",
        +        "estimated"
        +      ],
        +      "type": "string"
        +    }
        +  },
        +  "type": "object"
        +}
      • addedOutput schema / properties / scores_used
        Added value: +{
        +  "description": "The four pillar values the verdict was actually computed on, whether given by the caller or estimated by the engine. Show these to the user when any pillar was estimated.",
        +  "properties": {
        +    "change_enablement": {
        +      "type": "number"
        +    },
        +    "financial_return": {
        +      "type": "number"
        +    },
        +    "governance_risk": {
        +      "type": "number"
        +    },
        +    "strategic_alignment": {
        +      "type": "number"
        +    }
        +  },
        +  "type": "object"
        +}
      • addedOutput schema / properties / sensitivity
        Added value: +{
        +  "description": "What moves this verdict, computed deterministically: the value if readiness were one notch worse, the value at revenue minus 20 percent, and the nearest single-pillar movements that flip the classification. Boards trust ranges with visible assumptions over point estimates; show this.",
        +  "properties": {
        +    "readiness_one_notch_down": {
        +      "description": "Null when readiness is already siloed.",
        +      "type": [
        +        "object",
        +        "null"
        +      ]
        +    },
        +    "revenue_minus_20pct": {
        +      "type": "object"
        +    },
        +    "verdict_flips": {
        +      "items": {
        +        "type": "string"
        +      },
        +      "type": "array"
        +    }
        +  },
        +  "type": "object"
        +}
      • addedOutput schema / properties / work_architecture
        Added value: +{
        +  "description": "The work architecture gate across workflow, roles, human decision rights and performance measures. A stated gap or missing evidence blocks Accelerate.",
        +  "properties": {
        +    "blocks_accelerate": {
        +      "type": "boolean"
        +    },
        +    "checks": {
        +      "items": {
        +        "type": "object"
        +      },
        +      "type": "array"
        +    },
        +    "gaps": {
        +      "items": {
        +        "type": "string"
        +      },
        +      "type": "array"
        +    },
        +    "gate": {
        +      "type": "string"
        +    },
        +    "next_question": {
        +      "type": "string"
        +    },
        +    "status": {
        +      "enum": [
        +        "unknown",
        +        "partial",
        +        "gap",
        +        "ready"
        +      ],
        +      "type": "string"
        +    },
        +    "unknowns": {
        +      "items": {
        +        "type": "string"
        +      },
        +      "type": "array"
        +    }
        +  },
        +  "required": [
        +    "status",
        +    "blocks_accelerate",
        +    "checks",
        +    "gaps",
        +    "unknowns",
        +    "gate"
        +  ],
        +  "type": "object"
        +}
      • changedOutput schema / required
        Previous value: -[
        -  "bvf_version",
        -  "classification",
        -  "reason",
        -  "net_value_eur",
        -  "gross_value_eur",
        -  "decision_confidence",
        -  "multipliers",
        -  "drivers",
        -  "benchmark_source",
        -  "applied_modules"
        -]New value: +[
        +  "bvf_version",
        +  "classification",
        +  "reason",
        +  "net_value_eur",
        +  "gross_value_eur",
        +  "decision_confidence",
        +  "multipliers",
        +  "drivers",
        +  "benchmark_source",
        +  "applied_modules",
        +  "work_architecture"
        +]
    • Changedscore_portfolio18 fields changed
      • changedInput schema / properties / portfolio / description
        Previous value: -"A portfolio document conforming to the AI BVF v1.0 schema: bvf_version, organization (name, industry, optional revenue_eur), and a non-empty initiatives array. Each initiative carries id, name, function, ai_tier, and a scores object whose four pillars each carry a numeric value (0–100). organization.revenue_eur is required to model EUR value; initiatives that cannot be scored (missing revenue, unknown function/ai_tier) appear in skipped_initiatives rather than scored_initiatives. Schema: https://www.aibvf.com/protocol."New value: +"AI BVF v1.0 portfolio with organization and initiatives. Each initiative carries id, name, function, ai_tier and four scores, supplied as numbers or { value, confidence } objects. Retain each initiative's work_architecture, pillar_basis and optional signal_completeness. pillar_basis marks given or estimated values; estimated pillars are recalculated for the current context and retain an input-quality reduction. Supplied values without provenance are caller-provided, with no claim of evidence verification. Accelerate requires the pillar thresholds and all four work-architecture checks. Missing work-design evidence returns Fix for an otherwise green case. organization.revenue_eur is required for benefit modelling. Validate unfamiliar documents first."
      • changedInput schema / properties / readiness / description
        Previous value: -"Organisational readiness applied to every initiative in the portfolio. Honest self-assessment: agile = cross-functional, fast decisions; traditional = functional hierarchy; siloed = rigid, hand-off heavy. The portfolio schema does not carry per-initiative readiness; this single value sets the capture rate for the whole portfolio."New value: +"Organisational readiness applied to every initiative in the portfolio. Honest self-assessment: agile = cross-functional, fast decisions; traditional = functional hierarchy; siloed = rigid, hand-off heavy. The portfolio schema does not carry per-initiative readiness; this single value sets the capture rate for the whole portfolio and, paired with the ai_tier of each initiative, its pace-layer drag — lower readiness against a higher tier discounts the modelled EUR value."
      • changedOutput schema / properties / aggregate_net_value_eur / description
        Previous value: -"Sum of net EUR value across scored initiatives, low/high."New value: +"Arithmetic sum of benefit hypotheses. Reconcile overlapping scope, double counting and project costs before using this as a portfolio business case."
      • addedOutput schema / properties / feedback
        Added value: +{
        +  "description": "Optional three-question feedback route, present only when any initiative was Fix or Stop. The page records the response anonymously only when the user chooses an answer; no assessment data is attached.",
        +  "properties": {
        +    "question": {
        +      "type": "string"
        +    },
        +    "url": {
        +      "format": "uri",
        +      "type": "string"
        +    }
        +  },
        +  "required": [
        +    "question",
        +    "url"
        +  ],
        +  "type": "object"
        +}
      • addedOutput schema / properties / interpretation
        Added value: +{
        +  "additionalProperties": {
        +    "type": "string"
        +  },
        +  "description": "Meaning and limits of compatibility fields. Present these labels when explaining the result.",
        +  "type": "object"
        +}
      • changedOutput schema / properties / mean_decision_confidence / description
        Previous value: -"Mean decision confidence across scored initiatives (0–100); 0 when none were scored."New value: +"Mean decision score across scored initiatives (0–100); 0 when none were scored."
      • changedOutput schema / properties / scored_initiatives / items / properties / applied_modules / description
        Previous value: -"BVF scoring modules that fired for this initiative."New value: +"Scoring-rule and sector-context labels. Sector labels do not certify clinical validation or regulatory compliance."
      • addedOutput schema / properties / scored_initiatives / items / properties / audit
        Added value: +{
        +  "description": "Reproducibility record: engine version, the rules that fired, and the resolved inputs. Deterministic, no timestamps. If the verdict is challenged months later, the same inputs on the same engine version reproduce it exactly.",
        +  "properties": {
        +    "bvf_version": {
        +      "type": "string"
        +    },
        +    "engine": {
        +      "type": "string"
        +    },
        +    "engine_version": {
        +      "type": "string"
        +    },
        +    "inputs_used": {
        +      "description": "The resolved inputs the result was computed on, including estimated pillar values.",
        +      "type": "object"
        +    },
        +    "note": {
        +      "type": "string"
        +    },
        +    "rules_fired": {
        +      "description": "The rules that actually fired, in order: estimation, gates, classification, value arithmetic.",
        +      "items": {
        +        "type": "string"
        +      },
        +      "type": "array"
        +    }
        +  },
        +  "type": "object"
        +}
      • addedOutput schema / properties / scored_initiatives / items / properties / caveat
        Added value: +{
        +  "description": "Present only when signal_completeness was low: warns the verdict rests on soft inputs and confidence was reduced.",
        +  "type": "string"
        +}
      • changedOutput schema / properties / scored_initiatives / items / properties / decision_confidence / description
        Previous value: -"Confidence in the verdict (0–100)."New value: +"Heuristic decision score, 0 to 100, adjusted for input quality. This score has no probability calibration."
      • changedOutput schema / properties / scored_initiatives / items / properties / net_value_eur / description
        Previous value: -"Modelled net EUR value, low/high."New value: +"Modelled readiness-adjusted EUR benefit, low/high."
      • addedOutput schema / properties / scored_initiatives / items / properties / pillar_basis
        Added value: +{
        +  "description": "Per pillar: \"given\" (caller supplied it) or \"estimated\" (deterministic prior). When any pillar is estimated, tell the user which, and ask for evidence on those to firm up the verdict.",
        +  "properties": {
        +    "change_enablement": {
        +      "enum": [
        +        "given",
        +        "estimated"
        +      ],
        +      "type": "string"
        +    },
        +    "financial_return": {
        +      "enum": [
        +        "given",
        +        "estimated"
        +      ],
        +      "type": "string"
        +    },
        +    "governance_risk": {
        +      "enum": [
        +        "given",
        +        "estimated"
        +      ],
        +      "type": "string"
        +    },
        +    "strategic_alignment": {
        +      "enum": [
        +        "given",
        +        "estimated"
        +      ],
        +      "type": "string"
        +    }
        +  },
        +  "type": "object"
        +}
      • addedOutput schema / properties / scored_initiatives / items / properties / scores_used
        Added value: +{
        +  "description": "The four pillar values the verdict was actually computed on, whether given by the caller or estimated by the engine. Show these to the user when any pillar was estimated.",
        +  "properties": {
        +    "change_enablement": {
        +      "type": "number"
        +    },
        +    "financial_return": {
        +      "type": "number"
        +    },
        +    "governance_risk": {
        +      "type": "number"
        +    },
        +    "strategic_alignment": {
        +      "type": "number"
        +    }
        +  },
        +  "type": "object"
        +}
      • addedOutput schema / properties / scored_initiatives / items / properties / sensitivity
        Added value: +{
        +  "description": "What moves this verdict, computed deterministically: the value if readiness were one notch worse, the value at revenue minus 20 percent, and the nearest single-pillar movements that flip the classification. Boards trust ranges with visible assumptions over point estimates; show this.",
        +  "properties": {
        +    "readiness_one_notch_down": {
        +      "description": "Null when readiness is already siloed.",
        +      "type": [
        +        "object",
        +        "null"
        +      ]
        +    },
        +    "revenue_minus_20pct": {
        +      "type": "object"
        +    },
        +    "verdict_flips": {
        +      "items": {
        +        "type": "string"
        +      },
        +      "type": "array"
        +    }
        +  },
        +  "type": "object"
        +}
      • addedOutput schema / properties / scored_initiatives / items / properties / signal_completeness
        Added value: +{
        +  "description": "Input-quality factor after estimated provenance, pillar quality and any explicit factor are combined conservatively.",
        +  "maximum": 1,
        +  "minimum": 0,
        +  "type": "number"
        +}
      • addedOutput schema / properties / scored_initiatives / items / properties / work_architecture
        Added value: +{
        +  "description": "The work architecture gate across workflow, roles, human decision rights and performance measures. A stated gap or missing evidence blocks Accelerate.",
        +  "properties": {
        +    "blocks_accelerate": {
        +      "type": "boolean"
        +    },
        +    "checks": {
        +      "items": {
        +        "type": "object"
        +      },
        +      "type": "array"
        +    },
        +    "gaps": {
        +      "items": {
        +        "type": "string"
        +      },
        +      "type": "array"
        +    },
        +    "gate": {
        +      "type": "string"
        +    },
        +    "next_question": {
        +      "type": "string"
        +    },
        +    "status": {
        +      "enum": [
        +        "unknown",
        +        "partial",
        +        "gap",
        +        "ready"
        +      ],
        +      "type": "string"
        +    },
        +    "unknowns": {
        +      "items": {
        +        "type": "string"
        +      },
        +      "type": "array"
        +    }
        +  },
        +  "required": [
        +    "status",
        +    "blocks_accelerate",
        +    "checks",
        +    "gaps",
        +    "unknowns",
        +    "gate"
        +  ],
        +  "type": "object"
        +}
      • changedOutput schema / properties / top_initiative_by_value / description
        Previous value: -"Scored initiative with the highest mid-point net EUR value. Omitted when none were scored."New value: +"Scored initiative with the highest mid-point readiness-adjusted EUR benefit. Omitted when none were scored."
      • changedOutput schema / properties / top_initiative_by_value / properties / net_value_eur / description
        Previous value: -"Net EUR value range for the top initiative."New value: +"Readiness-adjusted EUR benefit range for the top initiative."
    • Addedsequence_portfolio
    • Changedvalidate_portfolio1 field changed
      • changedInput schema / properties / portfolio / description
        Previous value: -"The portfolio document as a JSON object following the AI BVF v1.0 schema: a top-level object with an \"initiatives\" array, each initiative carrying the same fields score_initiative expects (industry, revenue_eur, function, ai_tier, readiness, and a scores object). Validated structurally; values are not scored here."New value: +"The portfolio document as a JSON object following the AI BVF v1.0 schema: a top-level object with bvf_version, organization, and a non-empty \"initiatives\" array, each initiative carrying the same fields score_initiative expects (industry, revenue_eur, function, ai_tier, readiness, and a scores object with the four 0–100 pillars, each either a bare number or an object { value, confidence? }; both shapes pass). Checked structurally only — required fields present, correct types, enum values valid, pillar numbers in range; the pillar values are NOT scored or judged here (use score_initiative or score_portfolio for that). On failure, errors[] names each failing JSON path and the rule it broke."
  2. 1 tool updatev0.6.0
    • Changedscore_initiative2 fields changed
      • addedInput schema / properties / signal_completeness
        Added value: +{
        +  "description": "Optional 0–1. How grounded the four pillar scores are in real evidence versus estimated from context. Defaults to 1 (treated as measured). If the organisation lacks formal change-readiness or risk metadata, estimate the pillars from what you know AND set this lower to say so — decision confidence is reduced proportionally and a caveat is attached, instead of returning a falsely confident verdict on soft inputs.",
        +  "maximum": 1,
        +  "minimum": 0,
        +  "type": "number"
        +}
      • addedOutput schema / properties / caveat
        Added value: +{
        +  "description": "Present only when signal_completeness was low: warns the verdict rests on soft inputs and confidence was reduced.",
        +  "type": "string"
        +}
  3. 4 tool updatesv0.5.0
    • Addeddiagnose_process
    • Changedrecommend_improvements1 field changed
      • changedOutput schema / properties / projected_decision_confidence / description
        Previous value: -"Confidence in the verdict if the recommendations land, 0–1."New value: +"Confidence in the verdict if the recommendations land, 0-100."
    • Changedscore_initiative1 field changed
      • changedOutput schema / properties / decision_confidence / description
        Previous value: -"Confidence in the verdict, 0–1."New value: +"Confidence in the verdict, 0-100."
    • Addedscore_portfolio
  4. 6 tool updatesv0.4.0
    • Changedcalculate_pace_layer_drag4 fields changed
      • changedInput schema / properties / ai_tier / description
        Previous value: -"gen1=automation/RPA, gen2=GenAI, gen3=agentic."New value: +"Ambition of the AI being deployed: gen1=automation/RPA, gen2=GenAI, gen3=agentic."
      • changedInput schema / properties / industry / description
        Previous value: -"Optional, for future vertical adjustments."New value: +"Optional; defaults to universal. Reserved for future vertical adjustments."
      • changedInput schema / properties / readiness / description
        Previous value: -"Organisational readiness. Honest self-assessment."New value: +"Organisational readiness, honest self-assessment: agile = cross-functional, fast decisions; traditional = functional hierarchy; siloed = rigid, hand-off heavy."
      • changedOutput schema / (root)
        Previous value: -nullNew value: +{
        +  "properties": {
        +    "annual_drag_eur": {
        +      "description": "Estimated annual Organisational Drag Cost in EUR, low/high.",
        +      "properties": {
        +        "high": {
        +          "type": "number"
        +        },
        +        "low": {
        +          "type": "number"
        +        }
        +      },
        +      "required": [
        +        "low",
        +        "high"
        +      ],
        +      "type": "object"
        +    },
        +    "bvf_version": {
        +      "description": "AI BVF protocol version used.",
        +      "type": "string"
        +    },
        +    "drag_rate": {
        +      "description": "Drag as a fraction of revenue (e.g. 0.02 = 2%), low/high.",
        +      "properties": {
        +        "high": {
        +          "type": "number"
        +        },
        +        "low": {
        +          "type": "number"
        +        }
        +      },
        +      "required": [
        +        "low",
        +        "high"
        +      ],
        +      "type": "object"
        +    },
        +    "drivers": {
        +      "description": "Named factors contributing to the drag.",
        +      "items": {
        +        "type": "string"
        +      },
        +      "type": "array"
        +    },
        +    "pace_gap": {
        +      "description": "Severity of the tier↔readiness mismatch.",
        +      "enum": [
        +        "minimal",
        +        "moderate",
        +        "severe"
        +      ],
        +      "type": "string"
        +    },
        +    "source": {
        +      "description": "Citation for the drag-rate model applied.",
        +      "type": "string"
        +    }
        +  },
        +  "required": [
        +    "bvf_version",
        +    "annual_drag_eur",
        +    "drag_rate",
        +    "pace_gap",
        +    "drivers",
        +    "source"
        +  ],
        +  "type": "object"
        +}
    • Changedget_benchmark1 field changed
      • changedOutput schema / (root)
        Previous value: -nullNew value: +{
        +  "properties": {
        +    "cost_takeout_range": {
        +      "description": "Cost take-out as a fraction of revenue, lo/hi.",
        +      "properties": {
        +        "hi": {
        +          "type": "number"
        +        },
        +        "lo": {
        +          "type": "number"
        +        }
        +      },
        +      "required": [
        +        "lo",
        +        "hi"
        +      ],
        +      "type": "object"
        +    },
        +    "drivers": {
        +      "description": "Named value drivers behind the benchmark.",
        +      "items": {
        +        "type": "string"
        +      },
        +      "type": "array"
        +    },
        +    "function": {
        +      "description": "Business function the rates apply to.",
        +      "type": "string"
        +    },
        +    "industry": {
        +      "description": "Industry whose multiplier was applied.",
        +      "type": "string"
        +    },
        +    "industry_multiplier": {
        +      "description": "Multiplier applied to the base rates for this industry.",
        +      "type": "number"
        +    },
        +    "revenue_uplift_range": {
        +      "description": "Revenue uplift as a fraction of revenue, lo/hi.",
        +      "properties": {
        +        "hi": {
        +          "type": "number"
        +        },
        +        "lo": {
        +          "type": "number"
        +        }
        +      },
        +      "required": [
        +        "lo",
        +        "hi"
        +      ],
        +      "type": "object"
        +    },
        +    "source": {
        +      "description": "Citation for the benchmark figures.",
        +      "type": "string"
        +    }
        +  },
        +  "required": [
        +    "function",
        +    "industry",
        +    "revenue_uplift_range",
        +    "cost_takeout_range",
        +    "industry_multiplier",
        +    "drivers",
        +    "source"
        +  ],
        +  "type": "object"
        +}
    • Changedlist_taxonomy1 field changed
      • changedOutput schema / (root)
        Previous value: -nullNew value: +{
        +  "properties": {
        +    "ai_tiers": {
        +      "description": "All accepted ai_tier values (gen1/gen2/gen3).",
        +      "items": {
        +        "type": "string"
        +      },
        +      "type": "array"
        +    },
        +    "bvf_version": {
        +      "description": "AI BVF protocol version these enums belong to.",
        +      "type": "string"
        +    },
        +    "functions": {
        +      "description": "All accepted business-function values.",
        +      "items": {
        +        "type": "string"
        +      },
        +      "type": "array"
        +    },
        +    "industries": {
        +      "description": "All accepted industry values.",
        +      "items": {
        +        "type": "string"
        +      },
        +      "type": "array"
        +    },
        +    "readiness": {
        +      "description": "All accepted organisational-readiness values.",
        +      "items": {
        +        "type": "string"
        +      },
        +      "type": "array"
        +    }
        +  },
        +  "required": [
        +    "bvf_version",
        +    "industries",
        +    "functions",
        +    "ai_tiers",
        +    "readiness"
        +  ],
        +  "type": "object"
        +}
    • Changedrecommend_improvements1 field changed
      • changedOutput schema / (root)
        Previous value: -nullNew value: +{
        +  "properties": {
        +    "advisory_next_step": {
        +      "description": "Optional CTA, present only for Fix/Stop verdicts.",
        +      "type": "string"
        +    },
        +    "bvf_version": {
        +      "description": "AI BVF protocol version used.",
        +      "type": "string"
        +    },
        +    "current_classification": {
        +      "description": "Verdict as the initiative stands today.",
        +      "enum": [
        +        "Accelerate",
        +        "Fix",
        +        "Stop"
        +      ],
        +      "type": "string"
        +    },
        +    "feasible": {
        +      "description": "Whether the target is reachable via the listed pillar moves.",
        +      "type": "boolean"
        +    },
        +    "notes": {
        +      "description": "Caveats or context on the recommendation set.",
        +      "items": {
        +        "type": "string"
        +      },
        +      "type": "array"
        +    },
        +    "projected_decision_confidence": {
        +      "description": "Confidence in the verdict if the recommendations land, 0–1.",
        +      "type": "number"
        +    },
        +    "recommendations": {
        +      "description": "Per-pillar improvement actions.",
        +      "items": {
        +        "properties": {
        +          "action": {
        +            "description": "Concrete action to close the gap.",
        +            "type": "string"
        +          },
        +          "current": {
        +            "description": "Current pillar score (0–100).",
        +            "type": "number"
        +          },
        +          "delta": {
        +            "description": "Points of improvement required (target − current).",
        +            "type": "number"
        +          },
        +          "pillar": {
        +            "enum": [
        +              "strategic_alignment",
        +              "financial_return",
        +              "change_enablement",
        +              "governance_risk"
        +            ],
        +            "type": "string"
        +          },
        +          "rationale": {
        +            "description": "Why this action moves the pillar.",
        +            "type": "string"
        +          },
        +          "target": {
        +            "description": "Pillar score needed to flip classification (0–100).",
        +            "type": "number"
        +          }
        +        },
        +        "required": [
        +          "pillar",
        +          "current",
        +          "target",
        +          "delta",
        +          "action",
        +          "rationale"
        +        ],
        +        "type": "object"
        +      },
        +      "type": "array"
        +    },
        +    "target_classification": {
        +      "description": "Verdict the recommendations aim to reach.",
        +      "enum": [
        +        "Accelerate",
        +        "Fix",
        +        "Stop"
        +      ],
        +      "type": "string"
        +    }
        +  },
        +  "required": [
        +    "bvf_version",
        +    "current_classification",
        +    "target_classification",
        +    "feasible",
        +    "recommendations",
        +    "projected_decision_confidence",
        +    "notes"
        +  ],
        +  "type": "object"
        +}
    • Changedscore_initiative1 field changed
      • changedOutput schema / (root)
        Previous value: -nullNew value: +{
        +  "properties": {
        +    "advisory_next_step": {
        +      "description": "Optional CTA, present only for Fix/Stop verdicts.",
        +      "type": "string"
        +    },
        +    "applied_modules": {
        +      "description": "BVF scoring modules that fired for this input.",
        +      "items": {
        +        "type": "string"
        +      },
        +      "type": "array"
        +    },
        +    "benchmark_source": {
        +      "description": "Citation for the benchmark rates applied.",
        +      "type": "string"
        +    },
        +    "bvf_version": {
        +      "description": "AI BVF protocol version used.",
        +      "type": "string"
        +    },
        +    "classification": {
        +      "description": "The verdict for this initiative.",
        +      "enum": [
        +        "Accelerate",
        +        "Fix",
        +        "Stop"
        +      ],
        +      "type": "string"
        +    },
        +    "decision_confidence": {
        +      "description": "Confidence in the verdict, 0–1.",
        +      "type": "number"
        +    },
        +    "drivers": {
        +      "description": "Named value drivers behind the estimate.",
        +      "items": {
        +        "type": "string"
        +      },
        +      "type": "array"
        +    },
        +    "gross_value_eur": {
        +      "description": "Modelled gross value in EUR before capture, low/high.",
        +      "properties": {
        +        "high": {
        +          "type": "number"
        +        },
        +        "low": {
        +          "type": "number"
        +        }
        +      },
        +      "required": [
        +        "low",
        +        "high"
        +      ],
        +      "type": "object"
        +    },
        +    "multipliers": {
        +      "description": "Factors applied to the base rates.",
        +      "properties": {
        +        "capture_high": {
        +          "type": "number"
        +        },
        +        "capture_low": {
        +          "type": "number"
        +        },
        +        "industry": {
        +          "type": "number"
        +        },
        +        "tier": {
        +          "type": "number"
        +        }
        +      },
        +      "required": [
        +        "industry",
        +        "tier",
        +        "capture_low",
        +        "capture_high"
        +      ],
        +      "type": "object"
        +    },
        +    "net_value_eur": {
        +      "description": "Modelled net value in EUR after capture rate, low/high.",
        +      "properties": {
        +        "high": {
        +          "type": "number"
        +        },
        +        "low": {
        +          "type": "number"
        +        }
        +      },
        +      "required": [
        +        "low",
        +        "high"
        +      ],
        +      "type": "object"
        +    },
        +    "reason": {
        +      "description": "One-line justification for the classification.",
        +      "type": "string"
        +    }
        +  },
        +  "required": [
        +    "bvf_version",
        +    "classification",
        +    "reason",
        +    "net_value_eur",
        +    "gross_value_eur",
        +    "decision_confidence",
        +    "multipliers",
        +    "drivers",
        +    "benchmark_source",
        +    "applied_modules"
        +  ],
        +  "type": "object"
        +}
    • Changedvalidate_portfolio1 field changed
      • changedOutput schema / (root)
        Previous value: -nullNew value: +{
        +  "properties": {
        +    "bvf_version": {
        +      "description": "AI BVF protocol version validated against.",
        +      "type": "string"
        +    },
        +    "errors": {
        +      "description": "Empty when valid; otherwise one entry per schema violation.",
        +      "items": {
        +        "properties": {
        +          "msg": {
        +            "description": "The rule that was broken.",
        +            "type": "string"
        +          },
        +          "path": {
        +            "description": "JSON path to the failing field.",
        +            "type": "string"
        +          }
        +        },
        +        "required": [
        +          "path",
        +          "msg"
        +        ],
        +        "type": "object"
        +      },
        +      "type": "array"
        +    },
        +    "valid": {
        +      "description": "True when the portfolio conforms to the schema.",
        +      "type": "boolean"
        +    }
        +  },
        +  "required": [
        +    "bvf_version",
        +    "valid",
        +    "errors"
        +  ],
        +  "type": "object"
        +}
  5. 6 tool updatesv0.3.5
    • First observedcalculate_pace_layer_drag
    • First observedget_benchmark
    • First observedlist_taxonomy
    • First observedrecommend_improvements
    • First observedscore_initiative
    • First observedvalidate_portfolio

TDQS

A4.8/5.0

Scored across 13 tools

Disambiguation5/5

Every tool targets a distinct action+resource, and descriptions explicitly route between neighbors (e.g. 'use score_initiative when fields are known, assess_ai_initiative for plain language', 'use score_portfolio for several initiatives'). The assess/score_initiative pair both return verdicts but is cleanly separated by input state, and map_to_taxonomy vs list_taxonomy vs validate_portfolio vs assemble_portfolio are unambiguous.

Naming Consistency5/5

All 13 names are snake_case with a consistent verb_noun pattern (assess_ai_initiative, score_portfolio, validate_portfolio, map_to_taxonomy, get_benchmark, infer_readiness, diagnose_process). No mixed conventions or vague verbs like process/run/execute.

Tool Count5/5

13 tools is well within the sweet spot and each earns its place: lookup (taxonomy, benchmark), assembly/validation, scoring (single/portfolio), sequencing, and diagnosis/remediation are all distinct functions. No redundant or filler tools.

Completeness5/5

The surface covers the full lifecycle: free-text mapping, taxonomy listing, portfolio assembly and validation, single/portfolio scoring, wave sequencing, improvement planning, pace-layer cost, benchmark lookup, process diagnosis and readiness inference. No obvious dead ends for the stated domain.

Maintenance

ActivityActive
ResponsivenessWithin a week

Related MCP Connectors

Related MCP Servers

  • A
    license
    A
    quality
    D
    maintenance
    A pre-action risk gate for AI agents. Your agent calls the forecast tool before any irreversible action — send email, run SQL, make a payment, delete a file — and gets a risk score (0–100) and a GO / CONFIRM / STOP verdict in a few seconds.
    1
    53 npm
    MIT
  • F
    license
    Not graded
    quality
    D
    maintenance
    Provides AI agents with personalized timing intelligence by scoring decisions (0-100) against a user's energy profile and the Five Elements framework, enabling optimal scheduling for actions like trip planning, product launches, and meetings.
    1
    -
  • A
    license
    Not graded
    quality
    A
    maintenance
    Enables coding agents to scout, rank, and preflight software work before implementation, returning evidence-backed ACT, VERIFY, or SKIP decisions for issues and pull requests.
    98 npm
    2
    MIT
  • A
    license
    Not graded
    quality
    C
    maintenance
    Turns an unstructured business note into deterministic, explainable signals for revenue, automation, risk, urgency, and data, returning a score, priority, evidence, and a suggested next action. This lets an AI assistant assess and prioritize business opportunities through an auditable, locally computed tool.
    MIT