Skip to main content
Glama

powerbi-orchestrator-mcp

Un servidor MCP orquestrador que unifica modelado semántico, autoría de reportes, nube Fabric, validación y visualización/UX para Power BI / Fabric en 28 herramientas de alto nivel (no 500 primitivas).

El orquestrador delega a motores especializados (subprocess) y presenta al LLM una superficie coherente y de alto nivel.

Tests Coverage Python License MCP

Status: v1.11.0 — Beta. 28 tools (execute_dax_query added), 89% cobertura, product-readiness pass (CLI, Dockerfile, examples, plugin system). See release notes · changelog.


Quickstart (60 segundos)

# 1. Install from PyPI (v1.11.0).
pip install powerbi-orchestrator-mcp

# 2. Configure your MCP client (Claude Desktop shown).
#    Edit claude_desktop_config.json:
{
  "mcpServers": {
    "powerbi-orchestrator-mcp": {
      "command": "powerbi-orchestrator-mcp",
      "args": ["--start"],
      "env": {"PBI_AUTH_MODE": "interactive"}
    }
}
# 3. Verify the orchestrator itself works.
powerbi-orchestrator-mcp --start  # runs over stdio
# Or in another terminal, run the smoke test:
python scripts/verify_mcp_server.py

That's it — your LLM now sees 28 tools for Power BI / Fabric. See the examples/ directory for 3 reproducible workflows.


Related MCP server: Microsoft Fabric RTI MCP Server

Arquitectura de 3 capas (importante)

El proyecto NO es un wrapper sobre los MCP servers existentes. Es un servidor MCP propio que consume otros MCP servers como subprocess. Esto es lo que permite presentar al LLM 28 tools coherentes en lugar de 500 primitivas dispersas.

┌─────────────────────────────────────────────────────────────────┐
│ Capa 1: MCP Client (Claude Desktop, VS Code, Copilot, Cursor)    │
│         Habla JSON-RPC sobre stdio con el orquestrador.          │
│         El LLM ve 28 tools de alto nivel.                        │
└────────────────────────────┬────────────────────────────────────┘
                             │ stdio + JSON-RPC
┌────────────────────────────▼────────────────────────────────────┐
│ Capa 2: powerbi-orchestrator-mcp (ESTE PAQUETE — Python)        │
│         Distribuido via PyPI: pip install powerbi-orchestrator-mcp│
│         Console script: powerbi-orchestrator-mcp                 │
└────────────────────────────┬────────────────────────────────────┘
                             │ subprocess + JSON-RPC sobre stdio
┌────────────────────────────▼────────────────────────────────────┐
│ Capa 3: Engines individuales (heterogéneos)                       │
│         powerbi-modeling-mcp → npm: npx @microsoft/...          │
│         te (Tabular Editor)    → .NET binary                     │
│         superbi-mcp             → npm: npx superbi-mcp            │
│         dscmd (DAX Studio)      → Windows binary                  │
│         pbip-validator          → pip: pip install pbip-validator │
│         python_report           → built-in (parte del orquestrador)│
└─────────────────────────────────────────────────────────────────┘

Por qué npm NO es necesario para instalar el orquestrador (sí para correr operaciones reales): npm es una dependencia RUNTIME de los engines, no del orquestrador. El paquete powerbi-orchestrator-mcp se publica solo en PyPI.


Qué es

Un servidor Model Context Protocol (stdio) que expone 26 herramientas de alto nivel para que un agente IA pueda trabajar end-to-end con Power BI:

  • Diseñar y validar modelos semánticos (TMDL/TOM).

  • Crear, editar y auditar reportes (.pbix, PBIP/PBIR).

  • Operar en la nube (Fabric / Power BI Service): workspaces, datasets, refresh, deployment pipelines, RLS, Git integration.

  • Auditar calidad (BPA, lint DAX, accesibilidad WCAG, star-schema).

  • Diseñar visualizaciones con razonamiento de UX/storytelling.

Qué problema resuelve

Los MCPs existentes cubren partes:

  • powerbi-modeling-mcp (oficial MS): solo modelo semántico, no toca reportes.

  • superbi-mcp (cyphonica, 490 tools): local, Windows-only, FSL.

  • powerbi-mcp (sulaiman013, 82 tools): cloud paths mock-tested, no live.

  • fabric-rti-mcp, Fabric Core MCP: solo nube, no autoría local.

Nadie entrega orquestación cross-engine + nube maduro + UX verificable. powerbi-orchestrator-mcp sí.

Instalación

1. Instalar el orquestrador (Python)

pip install powerbi-orchestrator-mcp

El paquete está publicado en PyPI: https://pypi.org/project/powerbi-orchestrator-mcp/

Para desarrollo local (con tests y dev dependencies):

git clone https://github.com/berriosb/powerbi-orchestrator-mcp.git
cd powerbi-orchestrator-mcp
pip install -e ".[dev]"

El comando powerbi-orchestrator-mcp queda disponible en el PATH.

2. Instalar engines opcionales (solo si vas a usar operaciones reales)

Los engines son dependencias runtime del orquestrador. Si solo vas a probar con python_report (built-in), no necesitas instalar nada más.

# Node.js + npm (para powerbi-modeling-mcp, superbi-mcp)
# macOS:   brew install node
# Linux:   apt install nodejs npm
# Windows: https://nodejs.org/

# Tabular Editor CLI (modeling fallback + BPA)
# Windows/macOS: https://github.com/TabularEditor/TabularEditor/releases
# Linux: dotnet tool install --global TabularEditor

# pbip-validator (Microsoft, cuando esté publicado)
pip install pbip-validator

# DAX Studio (Windows only)
# https://daxstudio.org/

Ver docs/engines-setup.md para detalles de instalación por engine y troubleshooting.

3. Configurar el MCP client

Edita la config de tu MCP client (ej. claude_desktop_config.json):

{
  "mcpServers": {
    "powerbi-orchestrator-mcp": {
      "command": "powerbi-orchestrator-mcp",
      "args": ["--start"],
      "env": {"PBI_AUTH_MODE": "interactive"}
    }
  }
}

Compatible con VS Code + Copilot, Claude Desktop, OpenClaw, Hermes, Claude Code, Cursor y cualquier cliente MCP stdio.

4. Probar

En tu cliente MCP, el LLM ve 28 tools de alto nivel, agrupadas por capa:

Sesión y planificación

  • connect_target — abrir sesión contra un PBIP / Fabric workspace / PBI Desktop

  • plan_change — crear un plan versionable

  • apply_plan — ejecutar el plan con rollback

Modelado semántico

  • create_semantic_model_from_schema, add_measure_with_validation, refactor_to_calculation_groups, diff_models, generate_data_dictionary

Autoría de reportes

  • create_report_from_dataset, edit_report_visual, design_report_page_from_requirements, select_visuals_for_kpis, screenshot_report_pages, optimize_report_performance

Nube Fabric / Power BI Service

  • deploy_to_workspace, run_refresh, promote_in_pipeline, setup_rls_and_roles, set_sensitivity_labels, commit_workspace_to_git, sync_git_to_workspace, pre_deploy_check, execute_dax_query

Auditoría y calidad

  • audit_model_and_report, audit_report_ux_and_storytelling, apply_theme_and_accessibility_rules, run_dax_regression

Diagnóstico

  • powerbi_health

Con connect_target + plan_change + apply_plan solos, el LLM ya puede hacer safe_rename, audit, deploy y regression sobre cualquier PBIP local (sin engines externos) o cualquier Fabric workspace (con powerbi-modeling-mcp instalado).

Estado actual (v1.11.0)

  • ✅ 28 tools implementadas (modelado, reportes, nube, auditoría, UX, DAX)

  • ✅ Cross-engine rollback

  • ✅ Audit log con HMAC chain

  • ✅ PlanBuilder con templates versionables

  • ✅ Engine adapters: python_report (built-in), powerbi-modeling-mcp, superbi-mcp, te (Tabular Editor)

  • ✅ Story variance analysis (detección de regresiones visuales)

  • ✅ mypy --strict clean, ruff clean, CI matrix Linux/macOS/Windows

  • ✅ Backlog de hardening cerrado (0 items pendientes)

  • ✅ Publicado en PyPI: https://pypi.org/project/powerbi-orchestrator-mcp/

  • ⏳ Pendiente: tests E2E con binaries reales (te, dscmd)

Ver RELEASE-NOTES-v1.11.0.md para detalles completos.

Licencia

MIT.

Atribución

Available Tools

28 tools
add_measure_with_validationA

Add a DAX measure to a semantic model with automated linting and validation.

Use this tool when the user asks to:

  • Create or add a new DAX measure to a Power BI model.

  • Validate DAX syntax and best practices (preventing division by zero, unformatted measures, etc.).

  • Dry-run a measure to check for lint issues before committing to TMDL.

Args: target: Target PBIP directory or TMDL path. measure_name: Name of the measure to create. table: Target table where the measure will reside. expression: DAX formula for the measure (e.g. "DIVIDE([Total Sales], [Units], 0)"). format_string: Format string (e.g. "$#,##0.00", "0.0%"). description: Measure documentation or business description. is_hidden: Whether the measure should be hidden in report view. fail_on_severity: Minimum lint severity that blocks creation ("error", "warning", "info"). dry_run: If True, validate lint rules without writing to disk. runtime_check: Whether to execute the measure against an active engine if connected. measure_writer: Optional custom measure writer callable.

Returns: Dict with success status, lint findings, and modified file paths.

ParametersJSON Schema
NameRequiredDescriptionDefault
tableYes
targetYes
dry_runNo
is_hiddenNo
expressionYes
descriptionNo
measure_nameYes
format_stringNo
runtime_checkNo
measure_writerNo
fail_on_severityNowarning

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does well: it discloses dry_run semantics (validate without writing to disk), fail_on_severity blocking behavior, runtime_check requiring an active engine connection, and that file paths are modified. It omits important mutation behavior such as what happens when measure_name already exists (overwrite vs. error) and permission requirements, which keeps it below 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the one-line purpose, then a compact bulleted when-to-use block and an Args/Returns structure. It is slightly long, but since the schema has zero description coverage the per-parameter list earns its place rather than duplicating structured data.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For an 11-parameter write tool with no annotations, the description covers purpose, triggers, parameter meaning, and even return shape (which the output schema already provides). The remaining gap is edge-case behavior — collision with an existing measure of the same name, and what 'lint findings' look like when creation is blocked.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% and there are 11 parameters, so the description must compensate entirely — and it does, documenting every argument with concrete examples ('DIVIDE([Total Sales], [Units], 0)', '$#,##0.00', '0.0%') and enumerating fail_on_severity values (error/warning/info). This adds substantial meaning the schema does not provide.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

Clear specific verb+resource: 'Add a DAX measure to a semantic model', with the distinguishing scope 'automated linting and validation'. An agent can separate this from execute_dax_query or run_dax_regression by the write-plus-validate framing, but no sibling is named explicitly, so it falls short of full sibling differentiation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The bullet list gives concrete trigger conditions ('Create or add a new DAX measure', 'Dry-run a measure to check for lint issues before committing to TMDL'), which is real when-to-use guidance. However it never states when NOT to use it or which sibling to reach for instead (e.g. edit existing measure, execute_dax_query for read-only checks).

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

apply_planA

Execute an approved plan with automatic rollback on step failure.

Use this tool when the user asks to:

  • Execute or apply an approved plan created by plan_change.

  • Run plan steps in dry-run mode before modifying actual files.

Args: plan_id: ID of the plan previously created by plan_change. dry_run: If True, simulate execution without modifying files or cloud resources. confirm_each_step: Reserved hook for interactive step confirmations.

Returns: ApplyResult with execution status, executed steps, and rollback details if needed.

ParametersJSON Schema
NameRequiredDescriptionDefault
dry_runNo
plan_idYes
confirm_each_stepNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultNo
failed_stepNo
executed_stepsNo
rollback_handleNo
artifacts_changedNo
rollback_steps_executedNo

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden, and it does disclose meaningful behavior: automatic rollback on step failure, dry-run non-mutation, and that confirm_each_step is a reserved hook. It stops short of stating permissions/auth requirements or the risk profile of mutating cloud resources, leaving some gaps for a mutation tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core action and its rollback guarantee, then a compact usage list, then Args/Returns. Every section earns its place, though the Args/Returns restatements are somewhat redundant given the schema and output schema already cover them.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists so return values need not be explained (the description lightly summarizes ApplyResult anyway). The description covers triggers, params, and dry-run behavior, but omits auth/permission prerequisites expected of a tool that mutates files and cloud resources.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate, and it documents all three parameters (plan_id, dry_run, confirm_each_step) with meaning beyond the schema's bare titles. The descriptions are accurate and clarifying, though they add little syntax/format detail beyond what the names imply.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (Execute) and resource (an approved plan) plus a key behavioral trait (automatic rollback on failure). It explicitly ties the tool to plan_change, so an agent can distinguish it from the sibling that creates plans rather than running them.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The 'Use this tool when the user asks to' block gives concrete trigger conditions and the dry-run workflow, and it names plan_change as the source of the plan being applied. This routes the agent to the correct sibling and the correct mode without inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

apply_theme_and_accessibility_rulesA

Apply a colorblind-safe theme, backfill visual alt text, and re-audit WCAG.

Use this tool when the user asks to:

  • Make a report accessible and WCAG 2.1 compliant.

  • Apply a colorblind-safe palette (e.g. Okabe-Ito, ColorBrewer).

  • Automatically generate informative alt text for visuals lacking descriptions.

Args: pbip_path: Path to the .pbip directory. palette: Colorblind-safe palette name ("okabe_ito", "colorbrewer"). auto_backfill_alt_text: Whether to generate missing alt text on visuals. alt_text_template: Format template for generated alt text.

Returns: Dict with updated WCAG score, modified visual count, and theme update details.

ParametersJSON Schema
NameRequiredDescriptionDefault
paletteNookabe_ito
pbip_pathYes
alt_text_templateNo{visual_type} visualizing measure {first_measure}
auto_backfill_alt_textNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.8/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. It discloses that visuals are modified (a 'modified visual count' is returned) and that a WCAG re-audit occurs, which signals mutation. However, it says nothing about permissions required, whether changes to the .pbip directory are reversible, idempotency, or rate/scope limits.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core action, then clearly sectioned into usage triggers, args, and returns. Slightly padded by the repeated 'Use this tool when the user asks to' phrasing, but every line carries usable information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a multi-action mutation tool, the description covers purpose, triggers, all four parameters, and (redundantly, since an output schema exists) return shape. The remaining gap is behavioral: side effects on the .pbip directory and whether the operation is safe to re-run are not addressed, which matters more here because no annotations are supplied.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate, and it largely does: it documents all four parameters (pbip_path, palette with example values, auto_backfill_alt_text, alt_text_template). It gives the meaning and intent of each, though it omits the exact enum constraint for palette and the default values that the schema supplies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names three concrete actions (apply a colorblind-safe theme, backfill alt text, re-audit WCAG) on a specific artifact type, so the agent knows exactly what it does. It is clear but never explicitly differentiates itself from sibling audit tools like audit_report_ux_and_storytelling or audit_model_and_report, which is what a 5 would require.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The 'Use this tool when the user asks to:' block gives three explicit trigger scenarios (WCAG compliance, colorblind-safe palette, alt-text generation), which is strong positive guidance. It stops short of naming when-not to use it or pointing to an alternative sibling, so it does not reach a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

audit_model_and_reportA

Composite quality and compliance audit on a Power BI project (PBIP).

Use this tool when the user asks to:

  • Audit, inspect, or validate a Power BI project (.pbip) or semantic model.

  • Check Best Practice Analyzer (BPA) rules, DAX code quality, or naming conventions.

  • Validate WCAG 2.1 accessibility (contrast, missing alt text, chart readability).

Args: pbip_path: Path to the root .pbip directory or folder. bpa_ruleset: Best practice ruleset to run ("default", "strict", "lenient"). dax_measures_json: Optional JSON string or dictionary mapping measure names to DAX expressions. bpa: Whether to execute Tabular BPA checks. dax_lint: Whether to execute static DAX linting checks. accessibility: Whether to audit WCAG 2.1 accessibility on report pages. naming: Whether to validate column, measure, and table naming conventions.

Returns: Dict with overall score (0-100), pass/fail status, and categorized findings.

ParametersJSON Schema
NameRequiredDescriptionDefault
bpaNo
namingNo
dax_lintNo
pbip_pathYes
bpa_rulesetNodefault
accessibilityNo
dax_measures_jsonNo{}

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.8/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden. It usefully discloses that the audit is composite and that bpa/dax_lint/accessibility/naming are individually switchable, and that it returns a score/pass-fail with categorized findings. It does not state whether the operation is read-only, what permissions or runtime cost it incurs, or how the sub-checks interact.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core purpose, then structured Usage/Args/Returns sections with no filler. The Returns block is partly redundant given an output schema exists, but the overall structure is efficient and scannable.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 7-parameter, no-annotation tool with an output schema, it covers purpose, triggers, and every argument. The gaps are minor: no read-only/side-effect note and no interaction guidance among the boolean toggles.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate, and it does: all seven parameters are named and explained, including the bpa_ruleset enum values (default/strict/lenient) and the optional dax_measures_json. It omits declared defaults and how dax_lint interacts with dax_measures_json, so not a full 5.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (audit) and resource (Power BI PBIP project / semantic model / report) and enumerates the concrete dimensions checked: BPA rules, DAX quality, naming, WCAG accessibility. It does not explicitly differentiate itself from close siblings like audit_report_ux_and_storytelling or powerbi_health, so it stops short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides a clear 'Use this tool when the user asks to:' trigger list covering audit, BPA/DAX/naming, and WCAG validation, which gives an agent solid routing signal. However, it names no alternatives or when-not-to-use conditions against the many overlapping siblings.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

audit_report_ux_and_storytellingA

Evaluate report storytelling, visual hierarchy, cognitive load, and UX design.

Use this tool when the user asks to:

  • Audit report design quality, narrative flow, or visual hierarchy.

  • Check if a report follows dashboard best practices for a specific audience.

Args: pbip_path: Path to the .pbip directory. page_name: Optional specific page name to audit; audits all pages if omitted. audience_assumed: Target audience ("executive", "analytical", "operational"). strictness: Scoring strictness ("lenient", "standard", "strict").

Returns: Dict with UX score (0-100), category breakdowns (hierarchy, density, narrative), and suggestions.

ParametersJSON Schema
NameRequiredDescriptionDefault
page_nameNo
pbip_pathYes
strictnessNostandard
audience_assumedNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.1/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are supplied, so the description carries the full behavioral burden. It usefully discloses scope behavior ('audits all pages if omitted') and the return shape (0-100 score, category breakdowns, suggestions), but never states that this is a read-only, non-destructive local analysis with no writes to the .pbip, nor any permission or performance notes. Adequate but with a clear gap given zero annotation coverage.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Purpose is front-loaded in one line, followed by purpose-built trigger bullets and an Args/Returns block. Every section earns its place, though the Args and Returns lists restate information the schema/output schema already carry, adding some redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 4-parameter, single-required read/analysis tool with an output schema, the description covers what it does, when to use it, the meaning of each parameter, and the shape of the result. Given the output schema exists, the Returns summary is a bonus rather than a necessity; the only real gap is the absence of any behavioral/safety note.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate, and it does: all four parameters are documented in the Args block, including that page_name is optional and defaults to auditing all pages and that strictness/audience_assumed take named categorical values. It stops short of enumerating the allowed values as a closed set (schema declares no enums) or giving syntax examples for pbip_path.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The opening sentence names a specific verb (Evaluate/Audit) and concrete resources (report storytelling, visual hierarchy, cognitive load, UX design), which is far more informative than the tool name alone. The 'when the user asks to' bullets further scope it to design-quality and dashboard-best-practice auditing, distinguishing it from siblings such as audit_model_and_report (model+report) and design_report_page_from_requirements (generation).

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives explicit trigger conditions ('Audit report design quality, narrative flow, or visual hierarchy' and 'Check if a report follows dashboard best practices for a specific audience'), so an agent knows when to reach for it. It does not name an alternative or exclusion (e.g., when to prefer audit_model_and_report or optimize_report_performance instead), which keeps it short of a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

commit_workspace_to_gitA

Export and commit a Fabric workspace into a local Git repository.

Use this tool when the user asks to:

  • Back up or version-control a Fabric workspace into Git.

  • Snapshot reports and semantic models into local PBIP files with Git commits.

Args: workspace_id: Source Fabric workspace ID (UUID). output_repo_path: Local path to destination Git repository. branch: Git branch to commit into. commit_message: Commit message describing the snapshot. exclude_items: Optional list of item IDs to exclude. dry_run: If True, inspect items without creating git commits.

Returns: Dict with committed items, commit SHA, and repository status.

ParametersJSON Schema
NameRequiredDescriptionDefault
branchNo
dry_runNo
workspace_idYes
exclude_itemsNo
commit_messageNo
output_repo_pathYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.1/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden. It usefully discloses that dry_run defaults to True and returns committed items, commit SHA, and repo status, but says nothing about Git authentication/credentials, whether existing files in the target repo are overwritten or staged, or whether the operation is reversible.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Well front-loaded: purpose sentence first, then trigger bullets, then Args, then Returns. The trigger bullets partially restate the opening sentence, which is slight redundancy, but nothing is padded or wasted.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a six-parameter mutation tool with an output schema already present, the description covers purpose, triggers, parameters, and a return summary. The main remaining gap is operational prerequisites (Git auth/permissions and repo state preconditions), which an agent would want before committing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the Args block is the only parameter documentation and it covers all six parameters, including the important dry_run semantic ('inspect items without creating git commits'). It remains thin on defaults (branch, commit_message) and on what 'item IDs' in exclude_items actually refer to.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific verb pair (export and commit), the resource (Fabric workspace), and the destination (local Git repository). The 'into a local Git repository' direction cleanly separates it from the sibling sync_git_to_workspace, which flows the other way.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives an explicit 'Use this tool when the user asks to' block with two concrete trigger scenarios (backup/version-control a workspace, snapshot reports and semantic models into PBIP). It stops short of naming when NOT to use it or pointing to the reverse-direction sibling, so it is clear context without exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

connect_targetA

Connect to a Power BI target and initialize an orchestrator session.

Use this tool when the user asks to:

  • Connect to a Power BI Desktop session, Fabric workspace, PBIP folder, or PBIX file.

  • Start an interactive session to inspect or modify Power BI assets.

Args: target_type: One of "pbi_desktop", "fabric_workspace", "pbip_folder", "pbix_file". target_ref: Reference path (for PBIP/PBIX) or workspace UUID (for Fabric). auth_mode: "interactive" (default, browser login) or "service_principal". tenant_id: Azure AD tenant ID (required when auth_mode is service_principal).

Returns: ConnectResult with session_id, engines_available status, and any diagnostic warnings.

ParametersJSON Schema
NameRequiredDescriptionDefault
auth_modeNointeractive
tenant_idNo
target_refYes
target_typeYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
warningsNo
session_idYes
engines_availableYes

TDQS

A4.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. It discloses auth modes (interactive browser login vs service_principal) and that the result carries session_id, engines_available, and diagnostic warnings, which is useful. It is silent on session lifecycle: whether repeated calls create duplicate sessions, whether a session must be closed, what happens on connection failure, or whether the call is stateful/side-effecting — significant gaps for a session-initializing tool with zero annotation coverage.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the action and effect, then organized into use-when, Args, and Returns, so a scanning agent finds routing information first. Efficient overall, though the Returns block restates what the output schema already delivers and could be trimmed.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 4-parameter tool with an output schema present, the description covers every argument's meaning, the conditional requirement, the default, and the high-level result shape. The remaining omission is session lifecycle/cleanup behavior, which matters for a stateful connect call but is not fatal given a documented return type.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate, and it does: it enumerates the valid target_type values, explains that target_ref is a path for PBIP/PBIX but a workspace UUID for Fabric, states the auth_mode default ('interactive', browser login), and adds the conditional requirement that tenant_id is needed when auth_mode is service_principal. This adds semantics the schema does not carry.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource (connect to a Power BI target) plus the secondary effect (initialize an orchestrator session), which distinguishes it from the pipeline/modification siblings that all assume a live session. The target_type enumeration ('pbi_desktop', 'fabric_workspace', 'pbip_folder', 'pbix_file') further pins down scope. An agent can identify this as the session bootstrap step without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides an explicit 'Use this tool when the user asks to...' block covering both the connection request and the intent to start an interactive inspection/modification session. However, it names no alternatives or exclusions — it never says whether this must precede plan_change/apply_plan or how it relates to powerbi_health. Clear context, but no routing away from siblings.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

create_report_from_datasetA

Scaffold a PBIR report folder and layout from an existing semantic model.

Use this tool when the user asks to:

  • Create a new report (.Report folder) for an existing dataset.

  • Generate starter report pages with cards, charts, and an accessible theme.

Args: pbip_path: Path to the .pbip directory containing the dataset. page_name: Name of the initial report page (default: "Overview"). visual_count: Number of starter visuals to generate. theme: Theme name to apply (default: "okabe_ito"). include_card: Whether to generate a top-line KPI card visual. inspector: Optional model inspector.

Returns: Dict with created page path, visual IDs, and scaffolded report files.

ParametersJSON Schema
NameRequiredDescriptionDefault
themeNookabe_ito
inspectorNo
page_nameNoOverview
pbip_pathYes
include_cardNo
visual_countNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.8/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. It discloses what gets generated (pages, cards, charts, theme) and the return shape, but says nothing about whether an existing .Report folder is overwritten, whether the operation is idempotent, or what permissions/paths are required for a filesystem-writing tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded purpose sentence followed by a scannable use-case list and Args/Returns blocks. Slightly verbose in restating defaults that already live in the schema, but nothing is wasted or buried.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so the Returns block is a bonus rather than a necessity, and the generation behavior and parameters are well covered. The remaining gap is mutation semantics (overwrite/idempotency) for a tool that writes report scaffolding with no annotations to fall back on.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate, and it largely does: all six parameters (pbip_path, page_name, visual_count, theme, include_card, inspector) are explained with intent and stated defaults. The one weak spot is 'inspector', described only as 'Optional model inspector' with no indication of what object or behavior is expected.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (scaffold) and artifact (.Report folder plus layout) derived from an existing semantic model, which clearly separates it from model-creation siblings like create_semantic_model_from_schema. It does not explicitly distinguish itself from the arguably overlapping design_report_page_from_requirements, so it falls short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The 'Use this tool when the user asks to' list gives concrete trigger conditions (create a new .Report folder for an existing dataset, generate starter pages). However, it names no alternatives and gives no when-not guidance, e.g. versus design_report_page_from_requirements or edit_report_visual.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

create_semantic_model_from_schemaA

Generate a TMDL semantic model and PBIP project from a declarative schema spec.

Use this tool when the user asks to:

  • Create, scaffold, or generate a new Power BI semantic model from scratch.

  • Define tables, columns, data types, relationships, and hierarchies declaratively.

Args: spec_yaml: YAML specification of tables, columns, types, and relationships. spec_json: JSON specification (either as a string or a structured object/dict). output_pbip_path: Target directory to write the generated .pbip project. dry_run: If True, validate specification without writing files to disk.

Returns: Dict with tables created, relationships created, hierarchies created, and validation status.

ParametersJSON Schema
NameRequiredDescriptionDefault
dry_runNo
spec_jsonNo
spec_yamlNo
output_pbip_pathNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden. It does disclose that dry_run validates without writing files and that output lands as a .pbip project on disk, but it says nothing about overwrite behavior when output_pbip_path already contains a project, required permissions, or the odd empty-string default of output_pbip_path (where does it write then?).

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded purpose sentence, then a bulleted trigger list, then Args/Returns. Given 0% schema coverage the Args block earns its space; nothing is redundant filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so the Returns text is a redundant but harmless convenience. The description is nearly complete for a file-writing generator: params are fully explained. The residual gap is mutation safety semantics (overwrite/merge behavior) which is material for a tool that writes a project directory.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, but the description compensates by documenting all four parameters: spec_yaml, spec_json (string or structured object), output_pbip_path, and dry_run semantics. It omits the key relationship between spec_yaml and spec_json (mutual exclusivity / precedence) and does not flag that dry_run defaults to true, leaving the agent to infer it from the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Generate a TMDL semantic model and PBIP project') plus the input modality ('from a declarative schema spec'). The 'from scratch' framing distinguishes it from sibling mutation tools like refactor_to_calculation_groups or create_report_from_dataset, which operate on existing artifacts.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives clear when-to-use triggers ('Create, scaffold, or generate a new Power BI semantic model from scratch') and enumerates the declarative constructs it handles. It never names an alternative tool or states when NOT to use it (e.g., modifying an existing model), so it stops short of full routing guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

deploy_to_workspaceA

Deploy a PBIP project to a Fabric/Power BI workspace with pre-deploy gates.

Use this tool when the user asks to:

  • Deploy, publish, or release a Power BI project (.pbip) to Microsoft Fabric or Power BI Service.

  • Run pre-deployment quality gates before publishing.

  • Schedule automatic daily dataset refreshes upon publication.

Args: pbip_path: Local filesystem path to the root .pbip directory. workspace_id: Target Fabric / Power BI workspace ID (UUID). refresh_daily_hour: Daily UTC hour (0-23) for scheduled refresh (default: 6 AM UTC). findings_json: Optional list or JSON string of pre-existing audit findings to evaluate. gate_profile: Quality gate profile ("strict", "standard", "lenient"). auth_mode: "interactive" (default, browser login) or "service_principal". tenant_id: Azure AD tenant ID. client_id: Azure AD client ID. client_secret: Azure AD client secret. mock: If True, simulate deployment without calling external APIs.

Returns: Dict with deployment status, gate evaluation results, published item IDs, and refresh configuration.

ParametersJSON Schema
NameRequiredDescriptionDefault
mockNo
auth_modeNointeractive
client_idNo
pbip_pathYes
tenant_idNo
gate_profileNostandard
workspace_idYes
client_secretNo
findings_jsonNo[]
refresh_daily_hourNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the burden and does well: it discloses pre-deploy gate evaluation, published item IDs, scheduled daily refreshes, auth modes, and a mock/dry-run flag. It stops short of stating required permissions, reversibility, or side effects on an existing workspace deployment.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded purpose followed by scannable when-to-use bullets, then Args and Returns. Given 10 undocumented parameters, the length is justified and nothing repeats structured data.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 10-parameter deployment tool with no annotations and 0% schema coverage, the description supplies the missing parameter meaning, auth guidance, dry-run option, and return shape (with an output schema also present). Nothing an agent needs to invoke it correctly is absent.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, so the description must compensate and largely does — all 10 params are documented with types, defaults (6 AM UTC, 'standard', 'interactive'), and allowed values ('strict'/'standard'/'lenient', interactive/service_principal). It could go further by explaining what the gate profiles actually mean or the findings_json shape.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb + resource + scope ('Deploy a PBIP project to a Fabric/Power BI workspace with pre-deploy gates'), which lets an agent distinguish it from siblings like pre_deploy_check and promote_in_pipeline without opening a schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives an explicit when-to-use list covering deploy/publish, running quality gates, and scheduling refreshes. It does not, however, name the adjacent siblings (pre_deploy_check, promote_in_pipeline) or state exclusions, so an agent must infer boundaries between overlapping tools.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

design_report_page_from_requirementsA

Synthesize a complete PBIR report page layout from a natural language brief.

Use this tool when the user asks to:

  • Design or generate a new Power BI report page based on business requirements.

  • Automatically select, size, position, and format visuals matching an analytical goal.

  • Apply professional color schemes and visual hierarchy to a page.

Args: pbip_path: Path to the target .pbip directory. brief: Natural language description of what the report page should convey. page_name: Display name for the newly created report page (default: "Overview"). audience: Target audience ("executive", "analytical", "operational"). palette: Color palette name (default: "okabe_ito"). inspector: Optional model inspector for schema context.

Returns: Dict containing synthesized visual containers, positions, and page metadata.

ParametersJSON Schema
NameRequiredDescriptionDefault
briefYes
paletteNookabe_ito
audienceNoexecutive
inspectorNo
page_nameNoOverview
pbip_pathYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.5/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden for what is clearly a mutating, file-writing operation. It never states whether an existing page of the same name is overwritten, whether the pbip is modified on disk, what permissions or prerequisites apply, or whether the operation is reversible. The Returns line is the only behavioral disclosure and it duplicates the output schema.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded purpose sentence followed by tight, scannable usage bullets and a compact args block. The Returns section is redundant given an output schema exists and could be dropped, but overall there is little waste.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a six-parameter, zero-coverage-schema, unannotated mutation tool, the parameter documentation is adequate and the output schema covers return values. What is missing is the mutation/side-effect profile (overwrite behavior, disk writes) that the absent annotations leave entirely to the description.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate, and it largely does: all six parameters are named with meaning, including defaults for page_name and palette, the accepted audience values, and that inspector is an optional model inspector for schema context. Remaining gap is format detail for pbip_path and the inspector object.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The opening sentence gives a precise verb and resource: synthesize a complete PBIR report page layout from a natural-language brief. That is far more specific than sibling names like create_report_from_dataset or select_visuals_for_kpis, but the description never explicitly contrasts itself with those siblings, so an agent must infer the boundary.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The 'Use this tool when the user asks to' block gives three concrete triggering scenarios (new page from business requirements, auto visual selection/sizing, applying color schemes and hierarchy). That is clear usage context, but there is no when-not guidance and no named alternative for cases where an existing page should be edited (e.g. edit_report_visual).

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

diff_modelsA

Compare two semantic models and report structural differences.

Use this tool when the user asks to:

  • Compare two versions of a Power BI model or PBIP directory.

  • See what tables, columns, measures, or relationships changed between branches or releases.

Args: before: Path to the baseline PBIP directory or snapshot. after: Path to the target PBIP directory or snapshot. inspector: Optional model inspector callable.

Returns: Dict detailing added, removed, and modified tables, columns, measures, and relationships.

ParametersJSON Schema
NameRequiredDescriptionDefault
afterYes
beforeYes
inspectorNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.1/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden. It discloses the comparison inputs and the return shape (added/removed/modified tables, columns, measures, relationships), but never states that the operation is non-mutating or what permissions/paths are required.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded purpose sentence followed by structured usage bullets, Args, and Returns. Well-organized and every section earns its place, though the Returns section is slightly redundant given an output schema exists.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Covers purpose, usage, all three parameters, and the return structure. An output schema exists so return values needn't be explained, but the extra description is harmless and the definition is complete enough to invoke correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, so the description must compensate. It does: 'before' is the baseline and 'after' the target PBIP directory or snapshot, and 'inspector' is documented as an optional model inspector callable, giving real meaning beyond the bare string/any schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource: 'Compare two semantic models and report structural differences.' The purpose is unambiguous and no sibling tool performs a diff, so an agent can select it confidently.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides explicit triggering contexts via 'Use this tool when the user asks to: compare two versions... see what changed between branches or releases.' Clear positive guidance, though it names no alternatives or when-not conditions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

edit_report_visualA

Modify a visualContainer in a PBIR report page (type, fields, formatting, layout).

Use this tool when the user asks to:

  • Edit, reformat, or resize a specific chart/visual on a report page.

  • Change visual fields, bindings, alt text, or visibility.

Args: pbip_path: Path to the root .pbip directory. page_name: Name of the page containing the visual. visual_id: Unique ID of the visualContainer to edit. type: New visual type if changing (e.g. "barChart", "lineChart", "card"). fields_json: Optional dict or JSON string specifying field bindings to update. format_json: Optional dict or JSON string specifying formatting options. position_json: Optional dict or JSON string with x, y, width, height layout. alt_text: New alt text for accessibility. is_hidden: Whether to hide the visualContainer.

Returns: Dict with success status, changes applied, and page path.

ParametersJSON Schema
NameRequiredDescriptionDefault
typeNo
alt_textNo
is_hiddenNo
page_nameYes
pbip_pathYes
visual_idYes
fields_jsonNo
format_jsonNo
position_jsonNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. It discloses the return shape (success status, changes applied, page path), which is useful, but says nothing about permissions required, reversibility of edits, validation/error behavior, or what happens to fields not specified. Adequate but incomplete for a mutation tool with zero annotation coverage.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the core purpose, then offers a tight when-to-use block, an Args list, and a Returns line. Given nine undocumented parameters, the length is justified and every line earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists so return values need not be spelled out, yet the description still gives a concise summary. Combined with full parameter coverage, an agent has everything needed to call this tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With 0% schema description coverage, the description fully compensates by documenting all nine args: root .pbip path, page name, visual ID, new visual type with examples ('barChart', 'lineChart', 'card'), and the fields/format/position JSON blobs including accepted shapes (dict or JSON string; x/y/width/height). This adds substantial meaning the bare schema lacks.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (Modify) and resource (a visualContainer in a PBIR report page) and enumerates the scope (type, fields, formatting, layout). An agent can immediately see this is the visual-editing tool, though it doesn't explicitly name which siblings it is not (e.g. design_report_page_from_requirements).

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides an explicit 'Use this tool when the user asks to:' block with concrete triggers (edit/reformat/resize a chart, change fields/bindings/alt text/visibility). Clear usage context, but no explicit when-not or named alternative among the many sibling tools.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

execute_dax_queryA

Execute a DAX query against a published Power BI semantic model / Fabric dataset.

Use this tool when the user asks to:

  • Run, evaluate, or test a DAX query against a live dataset in Power BI or Fabric.

  • Inspect actual business data, measure outputs, KPI calculations, or table rows.

  • Verify Row-Level Security (RLS) filters by simulating a specific user principal name.

Args: workspace_id: Fabric / Power BI workspace ID (UUID). dataset_id: Published semantic model / dataset ID (UUID). dax_query: The DAX query expression (e.g. "EVALUATE TOPN(10, 'Sales')" or "EVALUATE ROW("Total", [Total Sales])"). impersonated_user_name: Optional User Principal Name (UPN) to test RLS rules as that user. auth_mode: "interactive" (default, browser login) or "service_principal". tenant_id: Azure AD tenant ID (required for service_principal). client_id: Azure AD client ID (for service_principal). client_secret: Azure AD client secret (for service_principal). fabric_client: Optional injected FabricClient instance (for testing).

Returns: Dict containing query execution results with tabular rows, columns, and execution metadata.

ParametersJSON Schema
NameRequiredDescriptionDefault
auth_modeNointeractive
client_idNo
dax_queryYes
tenant_idNo
dataset_idYes
workspace_idYes
client_secretNo
fabric_clientNo
impersonated_user_nameNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.8/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the burden. It discloses the auth modes, default behavior (interactive browser login), and that results are wrapped in a dict. But it lacks key behavioral details: whether the query is read-only, whether impersonation can mutate state, rate limits/timeouts, or error behavior — important for a tool that runs arbitrary DAX.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core purpose, then structured into use cases, Args, and Returns sections. The parameter list is verbose for the client_secret/tenant/client trio, but the structure is clear and scan-friendly.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Covers the two auth flows, the optional RLS impersonation path, the required parameters, and a brief return description, which is sufficient given a rich input schema with 9 parameters. The absence of documented error/timeout behavior and lack of differentiation from run_dax_regression are the main gaps.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate, and it does list every parameter with semantic detail: workspace_id/dataset_id are UUIDs, dax_query includes example expressions, auth_mode enum values, and tenant/client fields for service principals. Only missing: return-shape details per param and whether fabric_client is intended for production use.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (Execute) and resource (DAX query against a published Power BI semantic model / Fabric dataset). The use-case bullets make the domain clear. However, it doesn't explicitly distinguish from the sibling 'run_dax_regression', which also involves DAX execution — so the differentiation is partial.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides explicit when-to-use bullets: run/evaluate/test DAX, inspect business data, verify RLS. But no when-not-to-use or named alternatives (e.g. run_dax_regression for regression suites) are given, and the RLS bullet overlaps with sibling 'setup_rls_and_roles' without clarifying the difference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

generate_data_dictionaryA

Generate Markdown documentation and Mermaid ER diagram for a semantic model.

Use this tool when the user asks to:

  • Document a Power BI dataset or semantic model.

  • Generate a data dictionary listing all tables, columns, types, and descriptions.

  • Create a Mermaid entity-relationship (ER) diagram of the model.

Args: pbip_path: Path to the .pbip directory. output_path: Optional file path to save the generated Markdown. inspector: Optional model inspector.

Returns: Dict with data dictionary markdown content, table count, measure count, and output path.

ParametersJSON Schema
NameRequiredDescriptionDefault
inspectorNo
pbip_pathYes
output_pathNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.1/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description must carry the behavioral burden. It discloses that output_path optionally saves the Markdown (a file-write side effect) and describes the return payload, but says nothing about permissions, cost, or whether the operation touches the model. Adequate but thin for a tool with zero annotation coverage.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded one-line purpose followed by scannable trigger bullets and Args/Returns sections. The Returns block slightly duplicates the output schema, but overall it is efficient and well-organized.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 3-parameter documentation generator with an output schema present, the description covers purpose, triggers, parameters, and return shape. Nothing essential to a correct invocation is missing, though the inspector parameter's nature remains unspecified.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate, and it documents all three parameters: pbip_path as 'Path to the .pbip directory', output_path as an optional save path, and inspector as an optional model inspector. Only 'inspector' is somewhat circular/tautological, so a 4 rather than 5.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource (generate Markdown documentation and Mermaid ER diagram) for a semantic model, and the trigger bullets make the output artifacts concrete. It is clearly distinguishable from siblings like audit_model_and_report or create_report_from_dataset.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides an explicit 'Use this tool when the user asks to' list with three concrete trigger scenarios (document a dataset, generate a data dictionary, create a Mermaid ER diagram). It gives clear positive context but no when-not conditions or named alternatives.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

optimize_report_performanceA

Analyze PBIR report pages for visual performance bottlenecks and anti-patterns.

Use this tool when the user asks to:

  • Diagnose slow-loading report pages or improve visual performance.

  • Identify excessive visual density, expensive custom visuals, or unoptimized filters.

Args: pbip_path: Path to the .pbip directory. target_load_ms: Desired maximum page load latency in milliseconds (default: 5000 ms).

Returns: Dict with performance score (0-100), estimated load time, and actionable recommendations.

ParametersJSON Schema
NameRequiredDescriptionDefault
pbip_pathYes
target_load_msNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.8/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries full behavioral burden. 'Analyze' implies a read-only operation and the Returns block discloses a score/load-time/recommendations shape, but it never states prerequisites (e.g. that the .pbip must exist/be connected), that nothing is mutated, or any cost/latency characteristics of the analysis.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core purpose, then a scannable when-to-use list and short Args/Returns blocks. Minor redundancy in the Returns section given an output schema already exists, but nothing is wasted.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 2-parameter read-only analysis tool, the description covers purpose, triggers, both arguments, and the return shape. The Returns block partially duplicates the output schema, and no sibling routing is offered, but an agent has enough to call it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, so the description must compensate, and it does: pbip_path is explained as the .pbip directory path and target_load_ms as desired max page load latency with a 5000 ms default. Both parameters gain meaning beyond the bare schema titles.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (Analyze) plus resource (PBIR report pages) and scope (visual performance bottlenecks and anti-patterns). This is distinct from UX-focused siblings like audit_report_ux_and_storytelling, though the description never names a sibling to reinforce the boundary.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

An explicit 'Use this tool when the user asks to...' block with two concrete triggers (slow-loading pages, excessive visual density/expensive custom visuals). It gives clear positive context but no exclusions or named alternatives against siblings such as audit_model_and_report.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

plan_changeA

Create a versionable, multi-step execution plan from an intent template.

Use this tool when the user asks to:

  • Plan a complex or risky operation requiring atomic execution or rollback capability.

  • Safely rename a table or column across both model and visual bindings (intent="safe_rename").

  • Plan an audit, deployment, or DAX regression run.

Args: intent: One of the supported templates ("safe_rename", "audit", "deploy", "dax_regression"). options: Template-specific arguments (e.g. for safe_rename: old_path, new_path, scope). May also contain PlanOptions fields (auto_rollback, max_impact_threshold, dry_run_first).

Returns: PlanResult with plan_id, plan_yaml, steps, risk_score, estimated_changes.

ParametersJSON Schema
NameRequiredDescriptionDefault
intentYes
optionsNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
stepsNo
plan_idYes
plan_yamlYes
risk_scoreNo
estimated_changesNo

TDQS

A4.1/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are supplied, so the description carries the full disclosure burden. It usefully flags rollback/atomicity intent and points at options like auto_rollback and dry_run_first, but never states the critical fact that this only produces a plan and does not mutate the model (unlike apply_plan), nor any auth/prerequisite requirements.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded one-line summary, then a scannable usage list and short Args/Returns blocks. Slightly verbose in the Returns section (an output schema already exists), but every other sentence earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 2-parameter tool with an output schema, the description covers intent templates, option shapes, and the return fields the agent needs to chain into apply_plan. The main omission is the non-mutating/side-effect guarantee and any connectivity prerequisite.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% and the schema declares intent as a bare string with no enum, so the description is the only source of the four valid template values and the per-template options shape (old_path, new_path, scope). It adds substantial meaning beyond the schema, though options remains loosely specified.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and artifact ('Create a versionable, multi-step execution plan') plus its source ('from an intent template'). This clearly separates it from the sibling apply_plan, which presumably executes what this produces, without needing to open either schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives an explicit 'Use this tool when the user asks to' list covering complex/risky operations, safe_rename, and audit/deploy/DAX runs. It provides clear positive context but names no alternatives (e.g. apply_plan for execution) or when-not-to-use conditions, so it stops short of a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

powerbi_healthA

Diagnose orchestrator health, detected modeling engines, and storage readiness.

Use this tool when the user asks to:

  • Check if the Power BI MCP orchestrator is running properly.

  • See which external tools/engines are installed (Tabular Editor, DAX optimizer, etc.).

  • Get installation or setup instructions for missing components.

Args: include_engine_details: Whether to return full diagnostic info and remediation tips for each engine.

Returns: Dict with system status, engine availability matrix, active store counts, and remediation advice.

ParametersJSON Schema
NameRequiredDescriptionDefault
include_engine_detailsNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.1/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. 'Diagnose' plus the enumerated returns (system status, engine availability matrix, active store counts, remediation advice) implies a read-only operation, but the description never explicitly states that it has no side effects or what permissions it needs. The Returns section is largely redundant with the declared output schema.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Well front-loaded: purpose first, then use-when bullets, then args, then returns. The Args and Returns blocks are cleanly separated, but the Returns block duplicates the output schema and dilutes the otherwise tight structure.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a one-parameter read-only diagnostic with an output schema, the description covers purpose, triggers, parameter meaning, and return shape. The remaining gap is an explicit statement of read-only/non-mutating behavior and any auth or environment prerequisites for the engines it inspects.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the brief 'Whether to return full diagnostic info and remediation tips for each engine' is the only semantic explanation of the sole parameter, and it usefully clarifies the true/false tradeoff. It adequately compensates for the schema gap, though it could say more about verbosity/cost implications.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb+resource: 'Diagnose orchestrator health, detected modeling engines, and storage readiness.' This is a unique diagnostic function among the siblings (all of which are model/report operations), so an agent can immediately distinguish it without inspecting any schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives explicit trigger conditions ('Use this tool when the user asks to: check if the orchestrator is running, see which engines are installed, get setup instructions'), which is unusually clear for context selection. It stops short of a 5 only because it names no exclusions, though no sibling offers an overlapping capability that would need routing away from.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

pre_deploy_checkA

Evaluate audit findings against a pre-deployment quality gate profile.

Use this tool when the user asks to:

  • Check whether audit findings block deployment or meet quality criteria.

  • Validate findings against 'strict', 'standard', or 'lenient' governance gates.

Args: findings_json: List or JSON string of audit findings to evaluate. profile: Gate profile to evaluate against ("strict", "standard", "lenient").

Returns: Dict with gate decision (passed=True/False), blocking findings, and warnings.

ParametersJSON Schema
NameRequiredDescriptionDefault
profileNostandard
findings_jsonNo[]

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.8/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full behavioral burden. It discloses the return payload (gate decision, blocking findings, warnings), which is useful, but says nothing about side effects, permissions, or whether the check is purely read-only or has any deployment impact.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is front-loaded with the core purpose, then uses clean labeled sections (usage bullets, Args, Returns) that make it easy to scan. It is slightly longer than strictly necessary given the output schema, but nothing is wasteful.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a two-parameter, zero-required, non-destructive evaluation tool, both parameters are described and the return shape is sketched (even though a real output schema exists). The only gap is the absence of any behavioral/permission context, which is minor for this kind of gate-check tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, so the description must compensate, and it does: it documents both parameters, clarifying that findings_json accepts a list or JSON string, and enumerating the 'strict'/'standard'/'lenient' profile values that the schema leaves unspecified.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource ('Evaluate audit findings against a pre-deployment quality gate profile'), which is clear and actionable. It is reasonably distinguishable from siblings like audit_model_and_report, though it never explicitly contrasts itself with any of them.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives explicit 'use this tool when the user asks to' guidance with two concrete scenarios, which is stronger than most. However, it names no alternatives and gives no when-not-to-use conditions, so it falls short of full routing guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

promote_in_pipelineA

Promote artifacts across Microsoft Fabric Deployment Pipeline stages.

Use this tool when the user asks to:

  • Promote or move items between Fabric deployment stages (e.g. dev to test, test to prod).

  • Run pre-promotion quality gates before moving items.

Args: pipeline_id: Fabric deployment pipeline ID (UUID). source_stage: Source stage ("dev", "test", "prod"). target_stage: Target stage ("test", "prod"). items: Optional list of specific item IDs to promote. Promotes all if omitted. dry_run: If True, validate stages and gate checks without triggering actual promotion.

Returns: Dict with promotion status, gate outcomes, and affected items.

ParametersJSON Schema
NameRequiredDescriptionDefault
itemsNo
dry_runNo
pipeline_idYes
source_stageNodev
target_stageNotest

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.8/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations exist, so the description carries the full behavioral burden. It usefully discloses that dry_run validates stages and gate checks without executing, and that omitting items promotes everything, but it says nothing about required permissions, whether the target stage is mutated/overwritten, or reversibility — notable gaps for a promotion operation that can touch prod.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads purpose, then usage conditions, then args, then returns — a clean, skimmable order. The Args/Returns sections partially restate structure, but given 0% schema coverage that repetition earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 5-parameter mutation tool with no annotations, the description covers usage, parameters, and behavioral traits well, and an output schema exists so return detail is optional (it supplies it anyway). The main omission is any permission/side-effect guidance, which keeps it from being fully complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0% and the schema has only titles, so the description does the heavy lifting: it documents all five parameters with meaning (UUID, stage values dev/test/prod, optional item list defaulting to all, dry_run semantics). It even supplies enum-like stage values the schema lacks, though it could be more explicit about constraints.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (promote) and resource (artifacts) scoped to Microsoft Fabric Deployment Pipeline stages, which is unambiguous. It does not, however, name or distinguish itself from plausible siblings like deploy_to_workspace or pre_deploy_check, leaving the agent to infer the boundary.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The 'Use this tool when the user asks to' block gives concrete triggers (moving items dev→test→prod, running pre-promotion quality gates). It provides clear usage context but no explicit exclusions or named alternatives, so the agent must still infer when another tool is preferable.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

refactor_to_calculation_groupsA

Consolidate repetitive measures (e.g. YTD, QTD, PY) into calculation groups.

Use this tool when the user asks to:

  • Refactor or clean up redundant DAX measures using calculation groups.

  • Reduce model complexity and standardize time intelligence calculations.

Args: target: Target PBIP directory or TMDL path. min_candidates: Minimum measure patterns needed to trigger consolidation. reconcile_strategy: "strict" or "lenient". preserve_originals: Whether to keep original measures alongside the calculation group. auto_apply: If True, write calculation items immediately; if False, return proposed refactoring plan. inspector: Optional model inspector. measure_writer: Optional measure writer callable.

Returns: Dict with proposed or applied calculation items, candidate measures, and impact assessment.

ParametersJSON Schema
NameRequiredDescriptionDefault
targetYes
inspectorNo
auto_applyNo
measure_writerNo
min_candidatesNo
preserve_originalsNo
reconcile_strategyNostrict

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does disclose the key behavioral split: auto_apply=True writes immediately while False returns a proposed plan, and preserve_originals controls whether source measures survive. It does not state reversibility, permissions, or risk of modifying the model, leaving some mutation-safety gaps.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded one-line purpose followed by scannable trigger list and Args/Returns blocks. The Returns block largely restates what the existing output schema already provides, so it is slightly redundant but not bloated.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a mutation tool with no annotations and seven parameters, the definition covers purpose, triggers, every argument, and the plan-vs-apply distinction, and an output schema exists so return shape need not be re-explained. Missing only safety/reversibility context that annotations would normally supply.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, so the description must compensate, and it documents all seven parameters. Most add real meaning (auto_apply, preserve_originals, reconcile_strategy), though 'inspector: Optional model inspector' and 'measure_writer: Optional measure writer callable' are near-tautological and reconcile_strategy never explains strict vs lenient.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('consolidate') and resource ('repetitive measures ... into calculation groups') with a concrete example (YTD, QTD, PY), so the agent knows exactly what the tool does. It does not name or distinguish itself from any sibling, which keeps it just short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives explicit when-to-use triggers ('when the user asks to refactor or clean up redundant DAX measures', 'reduce model complexity'). There is no when-not guidance or named alternative (e.g. add_measure_with_validation), so it is clear but not exhaustive.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

run_dax_regressionA

Run DAX queries against a baseline and assert regression tolerance.

Use this tool when the user asks to:

  • Verify that measures or models return consistent results across changes.

  • Compare live DAX calculation outputs against a golden baseline file.

  • Check numerical tolerances on calculation outputs during CI/CD.

Args: baseline_path: Path to the JSON baseline file containing expected results. queries_json: Optional list or JSON string of DAX queries to execute. tolerance_pct: Maximum allowed percentage difference between actual and expected numeric values (default: 0.1%). query_executor: Optional custom query execution callable.

Returns: Dict containing diff summary, passed/failed queries, and variance details.

ParametersJSON Schema
NameRequiredDescriptionDefault
queries_jsonNo[]
baseline_pathYes
tolerance_pctNo
query_executorNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden. It discloses that the tool compares outputs against a baseline and returns a diff summary, passed/failed queries, and variance details. However, it does not explicitly state side-effect behavior, file-write semantics, permission requirements, or whether it is strictly read-only, leaving some behavioral gaps.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is front-loaded with purpose and usage, then follows a clean Args/Returns structure. The Returns section is somewhat redundant because an output schema exists, but otherwise the text is efficient and well-organized.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity, 0% schema coverage, and no annotations, the description is complete enough for an agent to invoke it correctly. It covers purpose, usage conditions, all parameter meanings, and return content, and the output schema provides additional return structure.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must fully compensate. It does: each of the four parameters is given clear semantics, including that baseline_path points to a JSON baseline, queries_json is an optional list or JSON string, tolerance_pct is a percentage with default 0.1%, and query_executor is an optional custom callable.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource ('Run DAX queries against a baseline and assert regression tolerance') and clearly scopes the operation to regression testing, distinguishing it from the sibling execute_dax_query. An agent can immediately tell this tool validates consistency rather than simply running queries.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides an explicit 'Use this tool when the user asks to' section with three concrete conditions: verifying consistency across changes, comparing live outputs against a golden baseline, and checking numerical tolerances during CI/CD. It lacks explicit when-not-to-use guidance or direct naming of alternatives like execute_dax_query.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

run_refreshA

Trigger, monitor, and optionally wait for a Power BI dataset refresh.

Use this tool when the user asks to:

  • Refresh data in a published Power BI semantic model.

  • Check the completion status of a refresh operation.

  • Perform full, automatic, or data-only refreshes.

Args: workspace_id: Fabric / Power BI workspace ID (UUID). dataset_id: Dataset / semantic model ID (UUID). refresh_type: "full", "automatic", "data_only", "calculate", or "clearValues". wait: Whether to poll and wait for the refresh to complete before returning. timeout_s: Maximum wait time in seconds (default: 1800). auth_mode: "interactive" or "service_principal". tenant_id: Azure AD tenant ID. client_id: Azure AD client ID. client_secret: Azure AD client secret.

Returns: Dict with refresh status, duration, error details, and rollback status if applicable.

ParametersJSON Schema
NameRequiredDescriptionDefault
waitNo
auth_modeNointeractive
client_idNo
tenant_idNo
timeout_sNo
dataset_idYes
refresh_typeNofull
workspace_idYes
client_secretNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.8/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden. It usefully discloses polling/wait behavior, the 1800s default timeout, both auth modes, and the shape of the return (status, duration, error, rollback). However, it never states that a refresh is a data-mutating operation, what permissions are required, or the cost/impact of triggering one, leaving significant behavioral gaps for a mutation tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is well front-loaded with purpose, then structured into 'Use this tool when', 'Args', and 'Returns' sections. The Args list is long but justified by nine parameters and the schema's lack of descriptions; a few entries are terse filler, but overall it earns its space.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be re-explained, and the description still summarizes them briefly. Parameters are fully enumerated and auth options covered; what is missing is the operational context an agent needs for a mutation (side effects, failure/rollback handling, permissions), so it is not fully complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate, and it does: all nine parameters are listed with meaning, and it supplies the refresh_type enum values ('full', 'automatic', 'data_only', 'calculate', 'clearValues') that are absent from the schema. The auth-related parameters are described only minimally (e.g., 'tenant_id: Azure AD tenant ID'), which keeps it from a 5.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The opening sentence gives a precise verb+resource: 'Trigger, monitor, and optionally wait for a Power BI dataset refresh,' which tells an agent exactly what the tool does. It does not, however, reference any sibling tool or explicitly disambiguate from alternatives, so it lands at 4 rather than 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The 'Use this tool when the user asks to:' block gives three concrete triggering conditions (refresh data, check completion status, perform specific refresh types). This is clear context for when to use it, but it offers no exclusions or alternative tools for cases like reading data instead of refreshing it.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

screenshot_report_pagesA

Capture screenshots or structural wireframes of Power BI report pages.

Use this tool when the user asks to:

  • Visually inspect or capture report pages for reviews or regression diffs.

  • Generate SVG wireframes or image snapshots of PBIR layouts.

Args: pbip_path: Path to the .pbip directory. pages: Optional list of specific page names to capture. format: Output format ("png", "svg", "pdf"). resolution: Target resolution ("desktop", "mobile", "tablet"). output_dir: Directory where captured images are written. wait_ms: Time in ms to wait for visual rendering.

Returns: Dict with output image paths, warnings, and rendering metadata.

ParametersJSON Schema
NameRequiredDescriptionDefault
pagesNo
formatNopng
wait_msNo
pbip_pathYes
output_dirNo./screenshots
resolutionNodesktop

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations exist, so the description carries the full burden. It does disclose that images are written to output_dir (a filesystem side effect), that wait_ms controls visual rendering, and that results include warnings and rendering metadata. However it says nothing about required permissions, whether a live connection (connect_target) must precede it, or how failures are surfaced.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded purpose sentence, then scannable use-case bullets, Args, and Returns. Given 0% schema coverage, the Args block earns its space rather than duplicating structured data.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a rendering tool with a large sibling set, the description covers purpose, invocation context, args, and return shape (even though an output schema also exists). The one real gap is prerequisite state — whether the report must be connected/open first — which the sibling connect_target implies but the description never states.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate, and it does: all six parameters are named with meaning, and it supplies value sets the schema lacks (format: png/svg/pdf; resolution: desktop/mobile/tablet; pages as an optional page-name list). It stops short of 5 only because units/defaults and interactions (e.g. format vs resolution) are not clarified.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (capture screenshots or structural wireframes) against a specific resource (Power BI report pages). No sibling tool does capture/snapshot work, so an agent can select it without ambiguity.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides explicit when-to-use triggers ('visually inspect or capture report pages for reviews or regression diffs', 'generate SVG wireframes or image snapshots of PBIR layouts'). No when-not conditions or named alternatives (e.g. audit_report_ux_and_storytelling), so it stops short of a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

select_visuals_for_kpisA

Recommend optimal visual types and chart configurations for given KPIs.

Use this tool when the user asks to:

  • Choose the best charts or visual types for a specific set of KPIs or metrics.

  • Tailor visual recommendations to an audience ('executive', 'analytical', 'operational').

  • Get primary and alternative chart suggestions with rationale based on data types.

Args: kpis_json: List of KPIs or JSON string (each with name, semantic_type, fields, etc.). audience: Target persona ("executive", "analytical", "operational"). max_results: Maximum number of alternative visual recommendations per KPI. inspector: Optional model inspector providing column cardinality and schema info.

Returns: Dict with recommended primary visual, alternatives, and rationale for each KPI.

ParametersJSON Schema
NameRequiredDescriptionDefault
audienceNoexecutive
inspectorNo
kpis_jsonYes
max_resultsNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.8/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. 'Recommend' implies a non-mutating helper, and the Returns section sketches the output shape, but it never states side-effect freedom, cost, or whether it requires a live model connection. Adequate but with clear gaps for a zero-annotation tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Structurally front-loaded with the purpose sentence followed by usage bullets and args. The Returns section partly duplicates the existing output schema, which is mild redundancy, but overall the text is tight and skimmable.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a four-parameter, zero-coverage, zero-annotation tool, the description covers purpose, triggers, and every argument. Since an output schema exists, the Returns prose is a bonus rather than a requirement, and only the absence of behavioral caveats keeps it from a 5.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate, and it does: it documents all four parameters, enumerates the audience personas ('executive', 'analytical', 'operational'), and explains inspector as supplying column cardinality/schema info. This is well beyond what the bare schema conveys.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource: 'Recommend optimal visual types and chart configurations for given KPIs.' An agent can tell this is a recommendation/design aid rather than a mutation. It does not explicitly differentiate itself from siblings like edit_report_visual or design_report_page_from_requirements, which keeps it from a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It provides explicit trigger conditions via three 'Use this tool when the user asks to:' bullets covering chart selection, audience tailoring, and primary/alternative suggestions. There is no statement of when NOT to use it or which sibling to prefer instead, so it falls short of a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

set_sensitivity_labelsA

Apply Microsoft Purview information protection sensitivity labels to items.

Use this tool when the user asks to:

  • Classify or protect Power BI items (reports, semantic models, dashboards).

  • Set Purview sensitivity labels (Confidential, General, Highly Confidential).

Args: items: List of dicts specifying item IDs and types (e.g. [{"id": "...", "type": "Report"}]). label_id: Microsoft Purview label GUID. label_name: Display name of the sensitivity label. admin_scopes: Optional list of administrative authorization scopes. redact_names: Whether to redact item names in returned logs for security. dry_run: If True, validate permissions without applying labels.

Returns: Dict with updated items, failed items, and compliance status.

ParametersJSON Schema
NameRequiredDescriptionDefault
itemsYes
dry_runNo
label_idYes
label_nameYes
admin_scopesNo
redact_namesNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.8/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden, and it does add real value: dry_run is explained as 'validate permissions without applying labels' and redact_names as log redaction. However, it omits whether labels are reversible, what permissions are actually required (admin_scopes is mentioned only as an arg), and whether dry_run defaults to True — schema shows default true, meaning an agent could believe labels were applied when nothing happened.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Purpose and trigger conditions are front-loaded, then Args/Returns blocks; nothing is wasted. The Args section inherently restates parameter names, and the Returns block is partly redundant given an output schema exists, but the utility description itself is tight.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists so return-value detail is not required, and the description still summarizes the shape (updated items, failed items, compliance status). For a 6-parameter mutation tool with zero annotations, it covers intent, params, dry-run safety, and error surface adequately; the remaining gap is permission/irreversibility context.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% (only titles like 'Label Id'), so the description must compensate, and it documents all six parameters with meaning beyond the schema: item dict shape with a concrete example, label_id as a Purview GUID, label_name as display name, admin_scopes as optional authorization scopes, and the semantics of redact_names and dry_run. It does not list valid label_name values or item type enumerations, which keeps it at 4.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The first sentence states a specific verb and resource ('Apply Microsoft Purview information protection sensitivity labels to items'), which is unambiguous and actionable. It does not name or contrast with any sibling tool, but no sibling in the list performs label application, so differentiation is implicit rather than stated.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

'Use this tool when the user asks to: classify or protect Power BI items... / set Purview sensitivity labels' gives explicit trigger conditions tied to user intent. It stops short of stating when NOT to use it or pointing to an alternative for adjacent tasks (e.g. RLS/roles via setup_rls_and_roles), so it earns 4 rather than 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

setup_rls_and_rolesA

Configure Row-Level Security (RLS) roles and validation rules in a TMDL model.

Use this tool when the user asks to:

  • Set up, add, or configure RLS roles and DAX table filter expressions.

  • Test and validate security rules against sample queries.

Args: target: Target PBIP directory or TMDL path. spec_yaml: YAML specification of security roles, members, and DAX filters. spec_json: JSON specification (either as a string or a structured object/dict). dry_run: If True, validate role specification without writing to disk. rollback_on_test_failure: Whether to revert modifications if test queries fail.

Returns: Dict with roles created, test query outcomes, and rollback status if applicable.

ParametersJSON Schema
NameRequiredDescriptionDefault
targetYes
dry_runNo
spec_jsonNo
spec_yamlNo
rollback_on_test_failureNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden. It usefully discloses that dry_run validates without writing to disk and that rollback_on_test_failure can revert changes, which are important safety semantics for a mutating setup tool. It still omits permissions, overwrite behavior, and broader side effects.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is front-loaded with purpose, then usage, args, and returns in a clear structure. The Returns section is somewhat redundant because an output schema exists, but the overall text remains efficient and easy to scan.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a five-parameter RLS setup tool with no annotations and an output schema, the description covers purpose, usage, arguments, and return shape. It is nearly complete, though it does not explain the expected schema of spec_yaml/spec_json beyond high-level contents, which would help an agent construct the call.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, but the description documents all five parameters: target, spec_yaml, spec_json, dry_run, and rollback_on_test_failure. It adds substantive meaning beyond the bare schema titles, including the YAML/JSON spec contents and the write/validation behavior of the boolean flags.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: 'Configure Row-Level Security (RLS) roles and validation rules in a TMDL model.' That is distinguishable from siblings such as set_sensitivity_labels or apply_theme_and_accessibility_rules, which address different Power BI concerns.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives explicit when-to-use guidance via bulleted user intents: setting up/adding/configuring RLS roles and DAX table filter expressions, plus testing and validating security rules. It does not name alternatives or exclusion conditions, so it stops short of the top score.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

sync_git_to_workspaceA

Deploy a local Git repository with PBIP projects to a Fabric workspace.

Use this tool when the user asks to:

  • Synchronize or publish a local Git repository or branch to a Fabric workspace.

  • Update workspace items based on version-controlled PBIP files.

Args: repo_path: Path to the local Git repository. workspace_id: Target Fabric workspace ID (UUID). branch_or_commit: Git ref to sync (default: "HEAD"). conflict_resolution: Conflict handling strategy ("manual", etc.). dry_run: If True, calculate changes without publishing.

Returns: Dict with synchronized items, skipped items, and deployment results.

ParametersJSON Schema
NameRequiredDescriptionDefault
dry_runNo
repo_pathYes
workspace_idYes
branch_or_commitNoHEAD
conflict_resolutionNomanual

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.1/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden; it does disclose that dry_run calculates changes without publishing and that conflict_resolution governs conflict handling, and it outlines the return shape. However, it never states that publishing overwrites existing workspace items, what permissions are required, or whether the operation is reversible, which are the key behavioral facts for a mutating deploy tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with a one-line purpose, then scannable 'when to use', Args, and Returns sections. Well organized and mostly earning its space, though the Returns block is largely redundant given that an output schema exists.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a mutating deploy tool with no annotations, the description covers the action, the trigger conditions, every parameter, and the dry-run safety path. Missing pieces are secondary details like overwrite semantics and auth prerequisites, and the return-values block is unnecessary since an output schema is present.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, so the description must compensate, and it documents all five parameters: repo_path, workspace_id (as a UUID), branch_or_commit (default HEAD), conflict_resolution (e.g. 'manual'), and dry_run (compute without publishing). It does not enumerate the allowed conflict_resolution values beyond 'manual', which is the only real gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource: deploying a local Git repo with PBIP projects into a Fabric workspace. This clearly distinguishes it from the reverse-direction sibling commit_workspace_to_git and from deploy_to_workspace. An agent can identify the operation without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides explicit 'Use this tool when the user asks to' bullets covering synchronization/publishing and updating items from version-controlled PBIP files. It gives clear positive triggers but never names alternatives (e.g. commit_workspace_to_git for the opposite direction) or states when NOT to use it.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 9 tool updatesv1.11.0
    • Changedaudit_model_and_report1 field changed
      • removedInput schema / properties / dax_measures_json / type
        Removed value: -"string"
    • Changedcreate_semantic_model_from_schema1 field changed
      • removedInput schema / properties / spec_json / anyOf
        Removed value: -[
        -  {
        -    "type": "string"
        -  },
        -  {
        -    "type": "null"
        -  }
        -]
    • Changeddeploy_to_workspace1 field changed
      • removedInput schema / properties / findings_json / type
        Removed value: -"string"
    • Changededit_report_visual3 fields changed
      • removedInput schema / properties / fields_json / anyOf
        Removed value: -[
        -  {
        -    "type": "string"
        -  },
        -  {
        -    "type": "null"
        -  }
        -]
      • removedInput schema / properties / format_json / anyOf
        Removed value: -[
        -  {
        -    "type": "string"
        -  },
        -  {
        -    "type": "null"
        -  }
        -]
      • removedInput schema / properties / position_json / anyOf
        Removed value: -[
        -  {
        -    "type": "string"
        -  },
        -  {
        -    "type": "null"
        -  }
        -]
    • Addedexecute_dax_query
    • Changedpre_deploy_check1 field changed
      • removedInput schema / properties / findings_json / type
        Removed value: -"string"
    • Changedrun_dax_regression1 field changed
      • removedInput schema / properties / queries_json / type
        Removed value: -"string"
    • Changedselect_visuals_for_kpis1 field changed
      • removedInput schema / properties / kpis_json / type
        Removed value: -"string"
    • Changedsetup_rls_and_roles1 field changed
      • removedInput schema / properties / spec_json / anyOf
        Removed value: -[
        -  {
        -    "type": "string"
        -  },
        -  {
        -    "type": "null"
        -  }
        -]
  2. 1 tool updatev1.9.0
    • Addedpowerbi_health
  3. 26 tool updatesv0.1.0
    • First observedadd_measure_with_validation
    • First observedapply_plan
    • First observedapply_theme_and_accessibility_rules
    • First observedaudit_model_and_report
    • First observedaudit_report_ux_and_storytelling
    • First observedcommit_workspace_to_git
    • First observedconnect_target
    • First observedcreate_report_from_dataset
    • First observedcreate_semantic_model_from_schema
    • First observeddeploy_to_workspace
    • First observeddesign_report_page_from_requirements
    • First observeddiff_models
    • First observededit_report_visual
    • First observedgenerate_data_dictionary
    • First observedoptimize_report_performance
    • First observedplan_change
    • First observedpre_deploy_check
    • First observedpromote_in_pipeline
    • First observedrefactor_to_calculation_groups
    • First observedrun_dax_regression
    • First observedrun_refresh
    • First observedscreenshot_report_pages
    • First observedselect_visuals_for_kpis
    • First observedset_sensitivity_labels
    • First observedsetup_rls_and_roles
    • First observedsync_git_to_workspace

TDQS

A3.8/5.0

Scored across 28 tools

Disambiguation4/5

Most tools target distinct operations (plan vs apply, run_refresh vs execute_dax_query, diff_models vs audit), and detailed descriptions clarify boundaries. There is notable overlap among the report-generation tools (create_report_from_dataset vs design_report_page_from_requirements) and among the multiple audit tools (audit_model_and_report, audit_report_ux_and_storytelling, optimize_report_performance), which could cause occasional misselection.

Naming Consistency4/5

The set overwhelmingly follows a snake_case verb_noun pattern (plan_change, apply_plan, connect_target, deploy_to_workspace, promote_in_pipeline). A couple of tools break the pattern with noun-first/phrase names like pre_deploy_check and powerbi_health, but these are minor deviations and still readable.

Tool Count3/5

28 tools is on the heavy side and pushes toward the borderline-heavy band. However, the server spans many genuinely distinct domains (connection, planning, DAX, auditing, report design, accessibility, Git sync, RLS, deployment, security labels), so most tools earn their place rather than being redundant.

Completeness4/5

The surface covers an unusually full lifecycle: connect/plan/apply, DAX execution and regression, model/report creation, editing, auditing, deployment, Git bidirectional sync, RLS, and sensitivity labeling. The main gap is destructive/removal operations (no delete_measure, delete_visual, delete_role, disconnect), which agents may need alongside the abundant create/edit tools.

Maintenance

ActivityActive
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • A
    license
    B
    quality
    D
    maintenance
    Enables AI assistants to interact with Microsoft Fabric and Power BI services through the Model Context Protocol. Users can manage workspaces, execute DAX queries, refresh datasets, and create Fabric notebooks using natural language.
    6
    19 npm
    2
    MIT
  • A
    license
    B
    quality
    A
    maintenance
    Enables AI agents to interact with Microsoft Fabric Real-Time Intelligence services, allowing for seamless data querying, analysis, and streaming capabilities.
    39
    5,623 PyPI
    132
    MIT
  • A
    license
    B
    quality
    A
    maintenance
    Enables AI agents to author and edit Power BI files locally (.pbix/PBIP), covering semantic models, reports, and Power Query M with 490 tools.
    100
    23 npm
    3
    Cryptographic Autonomy 1.0 (Combined Work Exception)