Skip to main content
Glama
josimarh

azure-mcp-pilot

by josimarh

IdenGraph

Identity & Access intelligence powered by Microsoft Graph and MCP.

A read-only MCP server that turns natural-language questions into audited answers about Microsoft Entra ID and Azure RBAC — answered by your own Copilot, inside VS Code.

VS Code Marketplace PyPI License: MIT

It never writes. Every create, update, delete, grant, or privilege-activation operation is blocked before routing.

What you can ask

  • Who has Owner on Azure? And who holds Global Administrator in Entra?

  • Which applications were created in my tenant, and which are Microsoft first-party?

  • Which users have no MFA, or rely on weak methods?

  • Which application secrets expire in the next 30 days?

  • Which objects have no owner assigned?

  • What is the blast radius of a given identity?

  • Are we paying for security features we don't use?

66 tools across users, groups, applications, service principals, managed identities, PIM, RBAC, authentication, conditional access, ownership, privilege timeline, licensing posture, and toxic combinations (SoD).

Related MCP server: Microsoft MCP Server for Enterprise

Design principles

These three rules define the behavior, and they matter more than the feature list.

Source separation

Microsoft Graph answers for identity and directory. Azure APIs answer for RBAC and resources. Graph is never treated as a source of truth for Azure RBAC.

No false zero

If a permission is missing, the answer is PERMISSION_DENIED with NOT_EVALUATED coverage — never 0.

"I could not evaluate this" and "this does not exist" are different answers. Conflating them in an audit is worse than not answering at all, because a false zero looks like a clean result.

This applies to licensing too: zero conditional access policies means something entirely different when the feature isn't licensed versus when it's licensed and unused.

Read-only by construction

Only GET, LIST, QUERY, ASSESS, and CORRELATE. The graph_query tool can receive read-only KQL generated from a user request, but it does not accept arbitrary endpoints: the capability registry selects the data source, and the executor blocks mutation operators and external-access constructs before sending the query.

Architecture

flowchart TB
    User(["You"]) -->|natural language| Copilot["Copilot Chat<br/><i>your own model</i>"]
    Copilot <-->|MCP / stdio| Server

    subgraph Server["IdenGraph MCP Server · 66 read-only tools"]
        direction TB
        Router["Capability router<br/><i>intent → capability</i>"]
        Guard{{"Write guard<br/><i>blocks mutations</i>"}}
        Registry[("Capability registry<br/>43 capabilities · 29 domains")]
        Executor["Read-only executor<br/><i>allowlist + validation</i>"]

        Router --> Guard --> Registry --> Executor
    end

    Executor -->|identity & directory| Graph["Microsoft Graph"]
    Executor -->|RBAC & resources| Azure["Azure Management<br/>+ Resource Graph"]

    Graph --> Normalizer["Normalizer & correlation<br/><i>preserves evidence and coverage</i>"]
    Azure --> Normalizer
    Normalizer -->|answer + provenance| Copilot

    style Guard fill:#c62828,color:#fff
    style Server fill:#0B1F3A,color:#fff
    style Normalizer fill:#1565c0,color:#fff

Two details worth highlighting:

The write guard sits before routing, not after. A mutation request is rejected before it can be interpreted as a query.

The normalizer preserves coverage, not just data. Every answer carries where it came from and whether the source could actually be evaluated — which is what makes the no-false-zero rule enforceable rather than aspirational.

How the pieces are distributed

Layer

Artifact

Role

Discovery

VS Code extension

One-click install, prerequisite checks, settings UI

Engine

idengraph on PyPI

The MCP server itself

Model

Your Copilot subscription

No LLM cost to this project or to you

The extension does not replace the Python package — it registers it. The engine runs the same way whether launched by the extension or configured by hand.

Install

Install IdenGraph from the Marketplace, then sign in to Azure:

az login

Open Copilot Chat in agent mode and ask. The extension verifies prerequisites and guides you if anything is missing.

Alternative: manual MCP configuration

Create .vscode/mcp.json:

{
  "servers": {
    "idengraph": {
      "type": "stdio",
      "command": "uvx",
      "args": ["idengraph"],
      "env": { "MOCK_MODE": "false" }
    }
  }
}

Prerequisites

  • Python 3.10+

  • uv

  • Azure CLI with an active session (az login)

Authentication uses DefaultAzureCredential, which reuses your Azure CLI session. There is no API key to manage, and no credential is stored by this project.

Configuration

Setting

Env var

Default

Purpose

idengraph.useMockData

MOCK_MODE

false (extension)

Query fictional data instead of your tenant

idengraph.sanitizeOutput

SANITIZE_FOR_LLM

true

Mask resource names, subscription IDs, and IPs before they reach the model

idengraph.subscriptions

AZURE_SUBSCRIPTIONS

all accessible

Restrict queries to specific subscriptions

The Python package defaults to MOCK_MODE=true so that nothing touches a real tenant without explicit intent. The extension sets it to false, since installing it is already that intent.

Permissions

On Azure: Reader on the subscriptions you want to audit.

On Microsoft Graph, delegated permissions vary by question:

Area

Permission

Users, groups, applications, service principals

Directory.Read.All

Directory roles and directory PIM

RoleManagement.Read.All

Authentication methods and MFA

UserAuthenticationMethod.Read.All

Conditional access

Policy.Read.All

License posture

Organization.Read.All

Missing a permission only marks the matching area as not evaluated. Everything else keeps working — and the affected area reports why it could not be evaluated.

Privacy

Queried data belongs to your tenant and travels between your machine, Microsoft APIs, and the Copilot model you already use. This project sends nothing to third-party servers and collects no telemetry.

By default, SANITIZE_FOR_LLM=true masks resource names, resource groups, subscription IDs, and IP addresses before content reaches the model.

Development

python -m venv .venv
.venv\Scripts\activate        # Windows
# source .venv/bin/activate   # Linux/macOS
pip install -r requirements.txt
cp .env.example .env
python test_smoke.py

All tests run in mock mode and never touch a tenant:

python test_smoke.py
python test_capability_layer.py
python test_tool_annotations.py
python test_tool_catalog_smoke.py

Repository layout

mcp_server.py                       MCP server and tool registration
services/azure_auth.py              credentials, tokens, caching
services/graph_capabilities.py      capability registry
services/capability_router.py       natural language → capability
services/capability_executor.py     validated read-only execution
services/azure_role_definitions.py  authoritative Azure role name resolution
services/azure_pim.py               resource PIM with confirmed coverage
services/entra_licenses.py          tenant licensing and feature availability
services/data/                      mock data used when MOCK_MODE=true
extension/                          VS Code extension (TypeScript)

The published package's surface is the MCP server only — there is no bundled UI or chat orchestrator; the model comes from your own Copilot/MCP client.

Known limitations

  • Agent Identity detection is heuristic where the directory exposes no dedicated type. Results are labeled as such, never presented as fact.

  • Public IP indicates a public address, which does not prove workload exposure.

  • Without RoleManagement.Read.All, directory PIM reports as not evaluated rather than empty.

  • License posture currently cross-references PIM and Conditional Access. Capabilities still marked not_integrated in the registry are not probed, and deliberately return no verdict.

License

MIT — see LICENSE.

Available Tools

66 tools
agent_natural_language_queryC
Read-onlyIdempotent

Interpreta perguntas sobre Agent Identities em linguagem natural.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNo
questionYes

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, destructiveHint=false, and openWorldHint=true, so the safety profile is fully covered structurally. The description adds nothing beyond that: it does not say what the natural-language interpretation does, whether it hits an LLM, how free-form the question may be, or how results are shaped.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single short sentence with the domain front-loaded and zero filler. It is efficient, though its brevity borders on under-specification rather than genuine conciseness.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, 0% parameter documentation, and no explanation of how free-form questions are handled or returned, the definition leaves an agent guessing about inputs and results. For a tool whose whole contract is natural-language interpretation, this is too thin.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% for both parameters. The word "perguntas" loosely maps to the required `question` parameter, but the description never explains the question format, and the `limit` parameter (default 20) is entirely unexplained in either schema or description.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ("Interpreta") and a scoped resource ("perguntas sobre Agent Identities"), which cleanly separates it from the sibling NL-query tools iam_natural_language_query, pim_natural_language_query, and timeline_natural_language_query. It never names those alternatives explicitly, but the Agent Identities domain is a real differentiator.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The domain scoping ("about Agent Identities") implies when this tool is the right choice versus the other natural-language query tools, but there is no explicit when-to-use, no prerequisites, and no named alternative. Usage must be inferred.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

assess_privileged_mfaB
Read-onlyIdempotent

Assessment de MFA para usuários privilegiados (sem MFA ou métodos fracos).

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnly, idempotent, non-destructive, and open-world behavior, so safety is covered. The description adds the assessed population and condition (privileged users without MFA or weak methods), but does not describe output, auth requirements, or rate limits beyond that.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

One front-loaded sentence with no filler; every phrase contributes to defining the assessment's scope.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-parameter, read-only assessment tool with no output schema, the description identifies the target population and assessment focus but does not say what the assessment returns or how it relates to overlapping authentication sibling tools. It is minimally adequate but leaves the agent to infer the output form.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Zero parameters, so per the rubric baseline is 4; the empty schema means there are no parameter semantics to document and the description appropriately does not introduce any.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource (MFA assessment) and scopes it to privileged users with no MFA or weak methods, which distinguishes it from broader authentication siblings. However, it does not name an alternative tool, so it stops short of full sibling differentiation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No explicit when-to-use or when-not-to-use guidance, and no alternatives named among overlapping siblings like list_users_without_mfa or list_users_with_weak_authentication. Usage must be inferred from the scope phrase.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

compute_identity_blast_radiusB
Read-onlyIdempotent

Estima o blast radius (alcance) de uma identidade (subscriptions, management groups, capacidade de conceder acesso).

ParametersJSON Schema
NameRequiredDescriptionDefault
identity_identifierYes

TDQS

B3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, openWorldHint and destructiveHint=false, so the safety profile is covered. The description adds useful behavioral context by enumerating what the radius covers (subscriptions, management groups, grant-access capability), but says nothing about cost, latency, or whether results are cached.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single, front-loaded sentence with no filler. It is well sized for the tool's simplicity, though the parenthetical enumeration slightly compresses information that could be structured more clearly.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a one-parameter read-only tool this is nearly adequate, and the parenthetical hints at what the estimate covers. However, with no output schema the description should say more about what the estimation returns, and it omits the input identifier format.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With schema description coverage at 0%, the description must carry the meaning of identity_identifier, yet it never states the expected format (UPN, object ID, service principal id). Only a vague reference to 'uma identidade' appears, leaving the single required parameter effectively undocumented.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('Estima') and resource ('blast radius de uma identidade') and clarifies the scope in parentheses (subscriptions, management groups, access-granting capability). It is distinguishable from the sibling list_top_blast_radius, though the description never explicitly contrasts single-identity computation with that list-style tool.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance, no prerequisites, and no mention of alternatives such as list_top_blast_radius or get_user_effective_azure_access. The agent must infer the appropriate context entirely on its own.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

detect_toxic_combinationsB
Read-onlyIdempotent

Detecta toxic combinations / violações de Separation of Duties com evidência.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNo

TDQS

B3.1/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, destructiveHint=false, and openWorldHint=true, so the safety and idempotency profile is clear. The description adds the notion of 'evidence,' implying findings are returned, but does not describe return format, pagination, or rate limits. With annotations carrying the main burden, a 3 is appropriate.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single concise phrase with no wasted words. It is front-loaded and easy to parse, though extremely short and perhaps too terse to be fully informative.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a detection/assessment tool in a rich identity context, the description is incomplete. There is no output schema to explain return values, no clarification of what 'evidence' entails, and no usage context. Annotations cover safety, but the agent still lacks sufficient info to invoke it confidently among many similar siblings.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

There is only one optional parameter (limit) and schema description coverage is 0%, so the schema does not explain its purpose. The description provides no parameter information at all. Baseline is not 4 because the single param lacks documentation, but the impact is limited due to low complexity.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb (detect) and resource (toxic combinations / Separation of Duties violations) and mentions evidence. It is distinguishable from siblings focused on listing, summarizing, or graphing, though it could be sharper about the exact output or scope.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No explicit when-to-use guidance or alternatives. The agent can infer it relates to SoD and toxic combinations, but there is no mention of prerequisites, when to choose this over run_iam_assessment or list_privilege_timeline_events, or what 'evidence' means in context.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

export_privilege_timeline_reportC
Read-onlyIdempotent

Gera relatório de timeline com top grants/revokes por identidade, role e escopo.

ParametersJSON Schema
NameRequiredDescriptionDefault
topNo
daysNo

TDQS

C2.4/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, destructiveHint=false, and openWorldHint, so safety is covered. The description adds no behavioral context beyond the annotations: it does not say where the export goes, what format it takes, whether it is synchronous, or whether large day ranges are bounded.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single, front-loaded sentence with no filler. It is efficient, though its brevity contributes to the coverage gaps noted elsewhere.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with two undocumented parameters, no output schema, and many timeline siblings, the description does not supply enough to call it correctly or to route to it over alternatives. Missing are the time-window semantics, output/export format, and any differentiation from sibling timeline tools.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description carries the full burden for the two parameters. It only loosely gestures at "top" (top grants/revokes) and says nothing about the "days" window or how the two interact, leaving both parameters effectively undocumented.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a concrete verb ("Gera relatório") and resource (timeline report) and specifies the report contents (top grants/revokes by identity, role, scope). However, it does not distinguish this tool from close siblings like summarize_privilege_timeline, get_identity_privilege_timeline, or list_privilege_timeline_events, leaving the agent to guess which timeline artifact it produces.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance, no prerequisites, and no mention of alternatives despite a crowded set of timeline-related siblings. The verb "export" is the only hint that this produces a file/report rather than an inline result, and that is left implicit.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_agent_relationshipsC
Read-onlyIdempotent

Resolve Agent → Owners → Identity → Blueprint → Graph Permissions → Azure RBAC.

ParametersJSON Schema
NameRequiredDescriptionDefault
agent_identifierYes

TDQS

C2/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, destructiveHint=false, and openWorldHint=true, so the safety profile is covered. The description adds nothing beyond the schema: no information on what each arrow step returns, whether it's a graph traversal, potential performance costs, or required permissions. For a composite resolution tool with annotations, the description should add context about the multi-hop nature but does not.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is extremely short—a single arrow chain. It is front-loaded and has no wasted words, but the cryptic notation may sacrifice clarity for brevity. Still, it earns credit for efficiency.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool that evidently resolves a multi-step ownership and permission path, the description is completely inadequate. It omits parameter details, expected outputs, usage context, and does not account for the lack of an output schema. An agent would not know how to invoke it correctly or interpret results.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% for the single required parameter 'agent_identifier'. The description does not explain what format or type of identifier is expected (e.g., object ID, name), leaving a critical gap. With low schema coverage, the description must compensate, and it fails to do so.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description is an arrow chain: 'Resolve Agent → Owners → Identity → Blueprint → Graph Permissions → Azure RBAC.' This does specify a verb (Resolve) and resource (Agent relationships/ownership/permissions), but the notation is cryptic and doesn't clearly distinguish from siblings like get_agents_by_owner or graph_permissions. An agent must guess that this is a composite traversal tool.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines1/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No indication of when to use this tool versus alternatives such as get_agents_by_owner, graph_permissions, or graph_role_assignments. No context, no exclusions, no prerequisites. The description provides zero usage guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_agents_by_ownerA
Read-onlyIdempotent

Lista Agents sob responsabilidade de um owner (UPN, nome ou objectId).

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNo
owner_identifierYes

TDQS

A3.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, openWorldHint, and destructiveHint=false, so safety behavior is covered. The description adds the accepted owner identifier formats but omits pagination, rate limits, or result-shape details. It does not contradict the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single, front-loaded sentence with no filler. The scoping condition and accepted identifier forms are packed efficiently without repetition.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple read-only list tool with rich annotations and no output schema, the description covers the essential purpose, owner scope, and identifier formats. It leaves minor gaps around the limit parameter and pagination behavior, but these are not critical for correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It adds meaningful semantics for owner_identifier by specifying UPN, name, or objectId, but it does not explain the limit parameter. This is partial compensation for a two-parameter tool.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a clear verb and resource ('Lista Agents') with an explicit scoping condition ('sob responsabilidade de um owner'). Sibling differentiation is not explicit; it does not name related tools like list_agent_identities or get_agent_relationships.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage context is implied: use this when you need agents owned by or under a specific owner. It does not state when to prefer alternatives or when not to use it. No explicit routing guidance is provided.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_authentication_methods_summaryB
Read-onlyIdempotent

Resumo de registro de MFA e capacidade passwordless do tenant.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

B3.3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true and destructiveHint=false, so the safety profile is covered without the description's help. The description adds the tenant-wide aggregation scope, but says nothing about what the summary returns, its granularity, or freshness, and there is no output schema to compensate.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single short sentence with no filler and the scope front-loaded. It is efficient, though the language mismatch with the rest of the tool surface slightly reduces clarity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a no-param, read-only aggregate the annotations carry the safety burden and no parameters need explaining. Still, with no output schema and no statement of what the MFA/passwordless summary contains, an agent must guess at the response shape.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so the schema and description cannot conflict and there is no parameter semantics to document. Baseline 4 applies for a no-argument tool.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific resource and scope: MFA enrollment plus passwordless capability at tenant level, which is clearer than a bare 'get summary'. However, it does nothing to distinguish itself from close siblings such as get_authentication_strength_summary or list_users_without_mfa, and it is written in Portuguese while the tool name and sibling names are English.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to call this versus the many overlapping authentication siblings (get_authentication_strength_summary, get_user_authentication_methods, list_users_without_mfa, assess_privileged_mfa). Usage is only implied by the phrase 'do tenant'.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_authentication_strength_summaryB
Read-onlyIdempotent

Resumo enterprise de força de autenticação (forte/misto/fraco) e adoção de passkey/FIDO2.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

B3.3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, openWorldHint, and destructiveHint=false, so the safety profile is covered. The description adds the summary categories and passkey/FIDO2 focus, but does not disclose scope (tenant-wide vs filtered), required permissions, or return shape; with annotations, this is moderate added value.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

One front-loaded sentence with no filler. It is appropriately sized for a parameterless summary tool, though it could have used the same space to clarify scope.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description carries more of the return-value burden. It identifies the summary dimensions but does not state tenant scope, timeframe, or the structure of the enterprise summary, leaving ambiguity relative to many authentication-related siblings.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has zero parameters, so per calibration the baseline is 4. There are no parameter semantics to add beyond the empty schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific summarization task and names the dimensions: authentication strength (strong/mixed/weak) and passkey/FIDO2 adoption. It does not differentiate this tool from sibling tools such as get_authentication_methods_summary, list_users_with_weak_authentication, or list_users_with_passkey.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No when-to-use guidance or alternatives are provided. The summary nature is implied by the tool name, but the sibling list contains many adjacent authentication tools and the description does not route the agent among them.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_environment_summaryB
Read-onlyIdempotent

Get a read-only high-level snapshot of the accessible Azure environment.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, and destructiveHint=false, covering the safety profile. The description adds that the snapshot is 'high-level' and limited to the 'accessible Azure environment,' but does not describe what the snapshot includes, its size, or any rate limits or auth requirements.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, front-loaded sentence with no wasted words, efficiently conveying the core idea.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given no parameters, full annotation coverage for safety, and no output schema, the description is minimally adequate. However, for a summary tool in a suite of many summary tools, it would benefit from hinting at the scope or content of the snapshot to help an agent choose it over more specific alternatives.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes no parameters, so there are no parameter semantics to document. Per the rubric, zero parameters baseline is 4.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('Get') and resource ('high-level snapshot of the accessible Azure environment'), making the tool's purpose clear. It does not explicitly differentiate from sibling summary tools like get_subscriptions_count or get_identity_access_summary, but its broad scope is implied.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no guidance on when to use this tool versus the many other summary tools available among its siblings. There are no exclusions, prerequisites, or alternative recommendations.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_identity_access_summaryA
Read-onlyIdempotent

Retorna resumo de identidade e RBAC:

  • total de usuários

  • usuários habilitados/desabilitados

  • usuários com permissões diretas

  • usuários desabilitados com roles ativas

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A3.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, destructiveHint=false and openWorldHint, so the safety profile is fully covered externally. The description adds the content breakdown, which is useful, but says nothing about cost, caching/freshness of the aggregated counts, or the shape of the response. With annotations carrying the behavioral load, a 3 is appropriate.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A one-line purpose statement followed by four terse bullets; the content is front-loaded and each bullet names a distinct returned metric, so nothing is wasted. Slightly over-fragmented for the amount of information, but efficient.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With zero parameters and no output schema, the description carries the burden of describing the return value — and it does, by enumerating the four aggregate metrics, which is exactly what an agent needs to decide to call it. What is missing is any indication of scope or freshness of the counts.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so the schema baseline for this dimension is 4. There is nothing for the description to clarify beyond what the empty argument object already communicates.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Retorna resumo de identidade e RBAC') and enumerates the four metrics returned (total users, enabled/disabled, direct permissions, disabled with active roles). It is clearly an aggregate-overview tool rather than a listing tool. However, it never names a sibling to differentiate itself from list_users or list_disabled_users_with_active_roles.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no explicit when-to-use or when-not-to-use guidance and no alternatives named, despite heavy sibling overlap (list_users, list_users_with_direct_permissions, list_disabled_users_with_active_roles). Usage is only implied by the enumerated contents: an agent can infer this is a quick overview rather than a detailed listing.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_identity_privilege_timelineC
Read-onlyIdempotent

Retorna timeline de privilégios para uma identidade específica (UPN, nome ou objectId).

ParametersJSON Schema
NameRequiredDescriptionDefault
daysNo
limitNo
identity_identifierYes

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, openWorldHint, and destructiveHint=false, so the safety profile is covered. The description adds no behavioral context of its own: nothing about default time windows, result volume, pagination, or what a 'timeline' entry contains. With annotations doing all the work, the description contributes essentially nothing here.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single compact sentence with no filler and the core action front-loaded. It is efficient, though the brevity is partly the problem rather than the virtue.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 3-parameter tool with no output schema and 0% schema description coverage, the description is too thin. It omits the meaning of the days/limit defaults and gives no sense of what the returned timeline looks like, so an agent cannot fully predict the call's behavior.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate, and it partially does by explaining that identity_identifier accepts a UPN, a name, or an objectId — genuinely useful since the schema only says 'string'. However, the days (default 90) and limit (default 200) parameters are never explained, so half the parameter surface remains undocumented.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: it returns a privilege timeline scoped to one identity. It is clear and unambiguous, but it does not differentiate itself from nearby siblings such as list_privilege_timeline_events, summarize_privilege_timeline, or timeline_natural_language_query, leaving the agent to infer the boundary.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no statement of when to use this tool versus the many timeline-related siblings (list_privilege_timeline_events, summarize_privilege_timeline, export_privilege_timeline_report, timeline_natural_language_query). Usage is only implied by the tool name.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_license_postureA
Read-onlyIdempotent

Cruza as licenças assinadas do tenant com o uso real dos recursos de segurança que elas habilitam.

Use para perguntas como:

  • Estamos pagando por recursos que não usamos?

  • Temos Entra ID P2? O PIM está sendo usado?

  • Quantas licenças estão ociosas?

Distingue explicitamente 'não licenciado', 'licenciado e não configurado' e 'licenciado mas não avaliado', porque as três levam a ações diferentes.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnly, idempotent, non-destructive, so safety burden is low. Description adds meaningful analytical context: it explicitly categorizes results into 'unlicensed', 'licensed but unconfigured', and 'licensed but unassessed', telling the agent these lead to different actions—genuinely useful beyond annotations. It doesn't mention data freshness or scope limits, keeping it below 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core operation, then example questions, then a key distinction. Every line earns its place. The three-category distinction sentence is slightly dense but well-placed. Minimal waste.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a no-param, read-only analytical tool with no output schema, the description gives enough to understand the return semantics (the three license posture categories) and scope. It could note result format or whether it covers all licensed products, but it is otherwise complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Tool has zero parameters, so baseline is 4. Schema coverage is 100% but trivially empty; the description correctly implies a tenant-wide operation with no inputs.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific analytical operation: crossing signed tenant licenses against actual usage of the security features they enable. This distinguishes it from sibling get_tenant_licenses (which presumably just lists licenses) by emphasizing the posture/usage-gap analysis. The example questions reinforce exactly what the tool answers.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides clear example questions that define the usage context ('are we paying for unused features', 'do we have Entra ID P2 and is PIM used', 'how many licenses are idle'). It does not explicitly exclude alternatives or name sibling tools like get_tenant_licenses vs this one, but the examples narrow intent well.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_management_group_inventoryC
Read-onlyIdempotent

Retorna inventário de Management Groups visíveis para auditoria de governança.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNo

TDQS

C2.8/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover the safety profile (readOnlyHint, idempotentHint, destructiveHint=false), so the description's burden is lower. It adds the scoping qualifier 'visíveis' (visible) which hints at permission-based filtering, but doesn't describe pagination behavior, the limit's effect, or the return shape. Modest added value over annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single efficient sentence with no waste. Front-loaded with the action and resource, though it could be more information-dense given the description's brevity leaves gaps elsewhere.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with no output schema, an undocumented pagination parameter, and many inventory-like siblings, the description is too thin. It should at least clarify what 'inventory' contains and how it differs from adjacent summary/count tools.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% for the single 'limit' parameter, so the description must compensate. It does not mention 'limit' or its default of 50 at all. Baseline of 3 is generous given the gap, but with only one self-evident pagination parameter the damage is limited.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a clear verb+resource (retorna inventário de Management Groups) and adds a purpose qualifier ('visíveis para auditoria de governança'). However, it does not differentiate from sibling tools like get_environment_summary or list_resource_groups, leaving ambiguity about its distinct scope.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description mentions an audit context but provides no explicit when-to-use, when-not-to-use, or alternative tool guidance. With many similar inventory/summary siblings, an agent has no routing information.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_owned_objectsC
Read-onlyIdempotent

Ownership 360: objetos (Groups, Applications, Service Principals, Agents, Blueprints) sob responsabilidade de um usuário.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNo
user_identifierYes

TDQS

C2.9/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, destructiveHint=false, so the safety profile is fully covered. The description adds the set of object categories that can be returned, which is useful scoping context, but says nothing about pagination, limits, or whether results are truncated.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single compact sentence with the scope front-loaded and no filler. It is perhaps too terse for the amount of undocumented behavior, but nothing is wasted.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema and two undocumented parameters, the description should at minimum explain the identifier format, the limit's effect, and whether results are paginated. None of this is present, leaving the agent under-equipped for correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must carry the parameter burden and does not: user_identifier's expected format (UPN vs. object ID) is unstated, and the limit parameter (default 100) is never mentioned at all. Only the vague phrase 'de um usuário' hints at user_identifier.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific resource set (Groups, Applications, Service Principals, Agents, Blueprints) and scopes it to objects owned by a user, so the agent knows what comes back. However, the verb is implicit (no 'list'/'retrieve') and it never distinguishes itself from close siblings like get_agents_by_owner or list_objects_without_owner.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance, no prerequisites, and no mention of alternatives such as get_agents_by_owner, which covers an overlapping subset of these objects. The agent must infer usage entirely from the scope phrase.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_pim_state_summaryB
Read-onlyIdempotent

Retorna comparativo de estados PIM (Active vs Eligible vs Permanent).

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

B3.3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint and destructiveHint=false, so the safety profile is covered. The description adds that the result is a comparative/aggregate rather than raw state rows, but says nothing about scope (tenant-wide?), which states are counted, or how the comparison is computed.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single compact sentence with no filler, front-loading the verb and the comparison dimensions. Nothing is wasted, though it is arguably under-specified rather than concise.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-parameter aggregation with no output schema, the description should say what the comparison contains (counts per state? breakdown by role?) and its scope. It leaves the return shape entirely unspecified.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so there is nothing for the description to disambiguate; baseline 4 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

Names a specific verb+resource: a comparative summary of PIM states across Active/Eligible/Permanent. It implies an aggregate view, distinguishing it somewhat from the sibling list_pim_role_states, though it never explicitly contrasts them.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No statement of when to use this versus list_pim_role_states or pim_natural_language_query. The agent must infer usage from the word 'comparativo' alone.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_resource_groups_countA
Read-onlyIdempotent

Retorna a quantidade de Resource Groups visíveis no escopo autenticado.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, openWorldHint=true, and destructiveHint=false. The description adds the authenticated-scope constraint, but does not describe return format, rate limits, or other behavior beyond what the annotations provide.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with no wasted words. It states the core action and scope immediately.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-parameter count tool with no output schema, the description is nearly complete: it says it returns a quantity within the authenticated scope. It could be slightly more explicit about the return type, but the meaning is clear.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so there is nothing for the description to clarify beyond the schema. Baseline 4 is appropriate because no parameter semantics are needed.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('Retorna') and resource ('quantidade de Resource Groups') along with scope ('escopo autenticado'). It clearly distinguishes this count tool from list_resource_groups and get_subscriptions_count.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No explicit when-to-use guidance or alternatives are named. The count semantics imply it is for aggregate counts rather than listing resources, but the description does not state that or route the agent from list_resource_groups.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_role_risk_scoreC
Read-onlyIdempotent

Calcula score de risco (0-100) para uma role privilegiada de Entra ID ou Azure RBAC.

ParametersJSON Schema
NameRequiredDescriptionDefault
roleYes
scopeNo/
stateNo
providerYes

TDQS

C2.9/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, destructiveHint=false and openWorldHint=true, so the safety profile is covered. The description adds useful context with the 0-100 score range and the two supported providers, but says nothing about what drives the score, caching, or latency.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with no filler; the score range and provider scope appear immediately. It is efficient, though the brevity comes partly from under-specification rather than tight editing.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, no annotation detail on return semantics, and 4 undocumented parameters, the description leaves an agent unable to determine valid provider values, how 'scope' and 'state' affect the result, or how to interpret the numeric score. For a tool with this parameter surface it does too little.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% across 4 parameters, and the description only indirectly implies 'role' and 'provider' ('role privilegiada de Entra ID ou Azure RBAC'). The optional 'scope' (default '/') and 'state' parameters — including valid provider values and scope path format — are entirely undocumented in both description and schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (Calcula) and resource (score de risco) and scopes it to privileged roles in Entra ID or Azure RBAC, with an explicit 0-100 range. It is clearly distinguishable from sibling listing/assessment tools, though it does not name a sibling it replaces.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance, no prerequisites, and no routing to alternatives such as detect_toxic_combinations or compute_identity_blast_radius. The agent must infer that this is a per-role scoring call from the description alone.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_subscription_direct_access_summaryC
Read-onlyIdempotent

Resumo de role assignments diretos em escopo de subscription (inclui privilegiados).

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNo

TDQS

C2.7/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, destructiveHint=false, and openWorldHint=true, so the safety profile is covered. The description adds that privileged assignments are included, which is useful scope context, but it does not describe return format, pagination, or any other behavioral trait beyond what annotations provide.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single efficient sentence that front-loads the resource and scope. It contains no wasted words, though its brevity contributes to gaps in other dimensions rather than being an issue of structure itself.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a summary tool with no output schema, the description should explain what the summary contains or how the limit parameter affects results. It only states the scope and inclusion of privileged assignments, leaving the agent without enough information to interpret or control the call fully.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% and the sole parameter ('limit', default 20) is not mentioned in the description at all. The description does nothing to compensate for the missing parameter documentation.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific scope ('role assignments diretos em escopo de subscription') and a qualifier ('inclui privilegiados'), making clear it is a subscription-level summary that includes privileged assignments. It does not name or contrast against siblings like list_azure_role_assignments or get_identity_access_summary, so it falls short of full sibling differentiation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no indication of when to use this tool versus alternatives such as list_azure_role_assignments, list_privileged_azure_role_assignments, or get_identity_access_summary. Usage is only implied by the scope, with no exclusions or prerequisites.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_subscriptions_countA
Read-onlyIdempotent

Retorna exclusivamente a quantidade de subscriptions Azure visíveis para a identidade autenticada.

Use esta ferramenta para perguntas como:

  • Quantas subscriptions existem?

  • Quantas assinaturas Azure eu tenho?

  • Quantas subscriptions consigo visualizar?

NÃO use search_official_guidance para descobrir quantidade de subscriptions.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnly, idempotent, non-destructive, open-world, so the safety profile is covered. The description adds meaningful context beyond that: results are scoped to what the authenticated identity can actually see (RBAC visibility), which explains why two callers get different numbers. No return format or rate-limit details.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The core behavior is front-loaded in the first sentence, followed by usage examples and an exclusion. The three example questions are largely paraphrases of one another, so the middle block is slightly redundant, but nothing is misleading or bloated.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-parameter counting tool with annotations covering safety, the description supplies the scope, the intended queries, and an anti-pattern tool to avoid. With no output schema present, a brief note on the exact return shape (a single integer/record) would make it fully self-contained, but everything needed to invoke it correctly is present.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so there is nothing for the description to disambiguate; the empty schema is self-explanatory. No parameter-level gaps to compensate for.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a precise verb+resource (return the count of Azure subscriptions) and adds a scope qualifier: only those visible to the authenticated identity. It also contrasts itself with a named sibling, search_official_guidance, so an agent can distinguish it without opening a schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides explicit when-to-use triggers via three example questions and an explicit when-not clause ('NÃO use search_official_guidance'). It does not mention the adjacent subscription-oriented siblings (e.g., get_subscription_direct_access_summary, list_resources) that an agent might otherwise pick, so the routing guidance is clear but not exhaustive.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_tenant_licensesB
Read-onlyIdempotent

Lista os SKUs licenciados do tenant, com unidades habilitadas, consumidas e ociosas, além dos planos de serviço ativos.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

B3.1/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, openWorldHint=true, idempotentHint=true, and destructiveHint=false, covering the safety and idempotency profile. The description adds the scope of returned data but nothing about freshness/caching, tenant scoping behavior implied by openWorldHint, or what 'ociosas' means operationally. Minimal added value beyond annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single compact sentence that front-loads the verb and resource. No wasted words, though it is somewhat under-elaborated given the absence of usage guidance.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The operation is simple (a read-only, no-param listing) and annotations cover its safety profile, so little is required. However, no output schema exists, so the description stands in for the return shape and only partially covers it (unit counts and service plans), and it omits any guidance distinguishing it from get_license_posture.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, which sets the baseline at 4. The description correctly implies a parameterless tenant-wide listing, but there is nothing more to document.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (Lista) and resource (SKUs licenciados do tenant) with the fields returned (habilitadas, consumidas, ociosas, planos de serviço). It is clearly distinct from the sibling get_license_posture only implicitly — it does not name it or explain the difference, so it falls short of the sibling-differentiation bar of 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No when-to-use or when-not-to-use information is given. The description does not reference the clearly related sibling get_license_posture, so the agent must guess which license tool to call. No exclusions or prerequisites are stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_user_authentication_methodsB
Read-onlyIdempotent

Retorna métodos de autenticação e status de MFA/passwordless de um usuário.

ParametersJSON Schema
NameRequiredDescriptionDefault
user_identifierYes

TDQS

B3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, destructiveHint=false and openWorldHint, so the safety profile is covered. The description adds the useful scope note that MFA/passwordless status is included, but says nothing about permissions, return shape, or pagination.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single efficient sentence with no wasted words, and the resource plus returned scope are front-loaded. Nothing extraneous is present.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple read-only single-user lookup with annotations covering safety, the description is adequate. However, with no output schema and a 0%-documented parameter, the identifier format and return contents remain unspecified.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0% and the single required parameter user_identifier has no description anywhere. The phrase 'de um usuário' implies a user reference but does not clarify whether the identifier is an email, UPN, or object ID, which is a real ambiguity for a lookup tool.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (retorna) and resource (métodos de autenticação), plus the scope of MFA/passwordless status for a single user. It is clear on its own, but it does not distinguish itself from siblings like get_authentication_methods_summary or get_authentication_strength_summary.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this per-user lookup versus the many authentication-related siblings (get_authentication_methods_summary, list_users_without_mfa, assess_privileged_mfa). The agent must infer usage from the name alone.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_user_effective_azure_accessB
Read-onlyIdempotent

Resolve acesso efetivo Azure RBAC de um usuário (direto + herdado via grupos transitive).

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNo
user_identifierYes

TDQS

B3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnly, idempotent, openWorld, and non-destructive, so the safety profile is covered. The description adds the meaningful behavioral fact that inheritance traverses transitive groups (not just direct assignments), which is beyond the annotations, but says nothing about pagination, result size, or how the tenant is scoped.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single sentence that is front-loaded with the verb and resource and has no filler. It is efficient, though its brevity means it also omits useful detail rather than being an example of economy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a nontrivial 'effective access' computation with no output schema, the description lets an agent know what is computed but not what comes back or how to page results. It is adequate to identify the tool but incomplete for calling it confidently.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must carry the param burden. It only implies a user argument ('de um usuário') via 'user_identifier' but never states the accepted formats (UPN vs object ID) and completely ignores the limit parameter, leaving both undocumented.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

It states a specific verb (resolve) and a precise resource (effective Azure RBAC access), and even scopes it to direct plus transitive-group-inherited access. It does not name or differentiate itself from siblings like list_users_with_direct_permissions or list_azure_role_assignments, so an agent must infer when this 'effective' view is preferred.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use or when-not-to-use guidance despite many closely related siblings (list_azure_role_assignments, list_users_with_direct_permissions, compute_identity_blast_radius). The description never says when resolving effective access beats listing raw assignments.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

graph_answer_identity_questionA
Read-onlyIdempotent

Interpreta uma pergunta de identidade em linguagem natural, escolhe a capability correta e executa a consulta de forma validada.

Use para perguntas abertas como:

  • "Quais usuários possuem Global Administrator?"

  • "Quais aplicações possuem Directory.ReadWrite.All?"

  • "Quem pode ativar Owner via PIM?"

Respeita a separação de fontes: Microsoft Graph para identidade/diretório e APIs Azure para Azure RBAC/recursos. Retorna erro explícito quando a capability não existe ou quando falta permissão. Nunca inventa chamadas Graph.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNo
questionYes

TDQS

A4.1/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond the annotations, which already declare read-only, open-world, idempotent, and non-destructive behavior, the description adds important operational traits: it respects source separation between Microsoft Graph and Azure APIs, returns explicit errors when a capability is missing or permission is lacking, and never fabricates Graph calls. This meaningfully reduces ambiguity about failure modes and trust boundaries.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is front-loaded with the core action, then supports it with examples and constraints. Every sentence earns its place: purpose, usage examples, source boundaries, error behavior, and the anti-hallucination guarantee. It is appropriately sized for a complex tool.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the complex natural-language query behavior, rich annotations, and absence of an output schema, the description covers purpose, usage, source boundaries, and error behavior well. It still leaves the 'limit' parameter unexplained and does not describe the general shape of returned results, which are minor but real gaps.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It explains the 'question' parameter as a natural-language identity question and provides examples, but it completely omits the 'limit' parameter and its default behavior. This is partial compensation only.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a clear verb and resource: it interprets a natural-language identity question, selects the correct capability, and executes a validated query. It also scopes the source to Microsoft Graph for identity/directory and Azure APIs for Azure RBAC/resources. However, it does not explicitly differentiate from sibling natural-language query tools such as iam_natural_language_query or pim_natural_language_query.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives clear usage context with concrete examples of open identity questions, such as Global Administrator holders and Directory.ReadWrite.All applications. It lacks explicit when-not-to-use guidance or named alternatives, but the examples provide strong context for selection.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

graph_assessmentB
Read-onlyIdempotent

Executa assessment de identidade baseado no capability registry, declarando explicitamente cobertura EVALUATED / PARTIAL / NOT_EVALUATED por domínio.

'scope' aceita: identity, directory, privileged, workload, azure. Nunca reporta NOT_EVALUATED como zero.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNo
scopeNoidentity

TDQS

B3.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, destructiveHint=false, and openWorldHint=true, so the safety profile is covered. The description adds a genuinely non-obvious output-semantics guarantee — coverage is declared explicitly per domain and NOT_EVALUATED is never collapsed to zero — which is a behavioral trait the annotations cannot express.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the action and its basis in the capability registry, followed by the scope enumeration and the NOT_EVALUATED caveat. Three short blocks, no filler; the only looseness is that the scope list could have been folded into the first sentence.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema the description carries some burden for return semantics, and it handles the coverage-labeling rule well. But it is silent on 'limit'/pagination behavior and offers no routing versus the overlapping run_*_assessment siblings, so an agent still has gaps before invoking it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It does so for 'scope' by enumerating the five accepted values (identity, directory, privileged, workload, azure), which the schema does not declare as an enum. However, 'limit' (default 200) receives no explanation at all, leaving half the parameters undocumented.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Executa assessment de identidade') and names the mechanism ('capability registry') plus what it produces (per-domain EVALUATED/PARTIAL/NOT_EVALUATED coverage). It stops short of distinguishing itself from close siblings such as run_iam_assessment, run_enterprise_identity_audit, and run_agent_assessment, which an agent would need help telling apart.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description only lists the accepted 'scope' values. It gives no when-to-use context, no prerequisites, and never names an alternative, even though the sibling set contains three other 'run_*_assessment' tools that overlap heavily in purpose.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

graph_directory_objectsC
Read-onlyIdempotent

Coleta objetos de diretório de um ou mais domínios e correlaciona a mesma identidade entre eles, preservando a fonte de cada evidência.

'domains' aceita lista separada por vírgula (ex.: "users,service_principals,directory_roles"). Se vazio, os domínios são inferidos da pergunta.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNo
domainsNo
questionNo

TDQS

C2.8/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations cover the safety profile (readOnlyHint, idempotentHint, destructiveHint=false, openWorldHint), so the description needn't restate safety. It adds a meaningful behavioral trait beyond annotations: it correlates the same identity across domains while preserving the source of each piece of evidence, which tells the agent the output merges identities and traces provenance. It does not mention limits, pagination, or what 'domains' values are valid beyond the single example.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short sentences, front-loaded with the core action before the parameter detail. No filler, though the quoted description includes leading/trailing whitespace artifacts. Efficient overall.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-required-parameter, cross-domain correlation tool with no output schema and 0% schema description coverage, the description should do more. It omits how 'limit' interacts with results, how 'question' drives inference, and what the correlated output looks like. It is not misleading but is incomplete for accurate invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate and it only partially does. It clarifies the 'domains' parameter format and the empty-value fallback, which is genuinely useful, but says nothing about 'limit' (a default of 100 with no semantics) or 'question' (how it drives domain inference). Two of three parameters remain undocumented.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a verb and resource ('Coleta objetos de diretório de um ou mais domínios') and adds a cross-domain correlation function, which is reasonably clear. However, it does not distinguish itself from the many sibling graph_* tools (graph_query, graph_list, graph_get, graph_relationship), leaving an agent unable to tell why this tool exists versus those. The purpose is legible but lacks sibling differentiation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explains the domains parameter behavior (comma-separated, inferred from the question if empty) but gives no explicit when-to-use guidance and no comparison to the numerous sibling tools that also retrieve directory objects. An agent is left to guess when this tool is the right choice over graph_query or list_users.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

graph_discover_capabilitiesA
Read-onlyIdempotent

Capability discovery de identidade.

Use ANTES de consultar quando não souber qual domínio/endpoint atende a pergunta.

  • Se 'question' for informado, retorna o roteamento sugerido (domínio, capability, fonte, permissões necessárias) SEM executar nada.

  • Se 'domain'/'source' forem informados, lista as capacidades catalogadas.

Fontes possíveis: microsoft_graph (identidade/diretório), azure_resource_graph / azure_management / azure_authorization (recursos e Azure RBAC).

ParametersJSON Schema
NameRequiredDescriptionDefault
domainNo
sourceNo
questionNo
include_gapsNo

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already establish read-only, idempotent, non-destructive, open-world behavior. The description adds useful context beyond that: a question returns routing and permissions without executing anything, and it enumerates possible sources for discovery.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with purpose, then usage, then mode behavior and possible sources in compact bullets. Every line contributes and there is no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a no-output-schema, optional-parameter routing/discovery tool, the description covers the main modes and source types adequately. The only notable gap is the unexplained include_gaps parameter, but an agent can still call the tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, so the description must compensate. It explains question, domain, and source well, but include_gaps is never mentioned, leaving one of four parameters undocumented.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States it is a capability-discovery tool for identity and explains the two modes: question → routing suggestion, domain/source → catalog listing. It is clear enough to distinguish from action-oriented siblings like graph_query or list_users, though no sibling is named explicitly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly says to use it BEFORE consulting when the domain/endpoint is unknown, and explains which input triggers which mode. No when-not condition or named alternative is provided, but the usage context is clear.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

graph_getC
Read-onlyIdempotent

Obtém um objeto específico de uma capability catalogada (operação 'get').

'id' deve ser um GUID ou identificador simples; valores fora desse padrão são rejeitados.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYes
selectNo
capability_idYes

TDQS

C2.7/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, destructiveHint=false, and openWorldHint, so the safety profile is covered. The description adds one genuine behavioral fact beyond them: 'id' must be a GUID or simple identifier and non-conforming values are rejected. It does not disclose error behavior for unknown capability_id or what the response contains.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences, front-loaded with the operation's purpose followed by the id constraint. No filler, though the parenthetical '(operação get)' is redundant with the tool name.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With 3 undocumented parameters, no output schema, and a vague 'capability catalogada' concept, the agent lacks the context to call this correctly: it cannot know what capability_id values are valid, what select does, or what shape of object comes back.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% across 3 parameters, so the description carries the burden. It clarifies only 'id' (accepts GUID or simple identifier, rejects other formats) and says nothing about the required 'capability_id' or the optional 'select' projection parameter.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a clear verb ('obtém') and resource ('um objeto específico de uma capability catalogada'), and the parenthetical '(operação get)' signals it is the read of a single item, implying a list sibling. However, 'capability catalogada' is never defined, so an agent cannot tell what kinds of objects this actually returns versus graph_list, graph_query, or graph_directory_objects.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this versus the many sibling retrieval tools (graph_list, graph_query, graph_relationship, graph_permissions, graph_role_assignments). The only constraint given is a format rule for 'id', which is a precondition rather than usage guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

graph_listB
Read-onlyIdempotent

Lista objetos de uma capability catalogada (operação 'list').

'capability_id' precisa existir no registry (use graph_discover_capabilities). 'filter' aceita apenas campos declarados como suportados pela capability. Endpoints arbitrários são rejeitados.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNo
filterNo
selectNo
capability_idYes

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, openWorldHint and destructiveHint=false, so the safety profile is covered. The description adds useful behavioral context (arbitrary endpoints rejected, filter limited to capability-declared fields), but says nothing about pagination, result size, or limit behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short, front-loaded sentences with no filler; the constraint on capability_id and filter is stated directly. Efficient for the information conveyed, though the parenthetical operation note is slightly redundant.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, 0% schema coverage and four parameters, the description should do more: it never explains return shape, pagination, or the semantics of 'limit'/'select'. Annotations cover the safety profile, but an agent still lacks guidance on two parameters and the response format.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It explains 'capability_id' (must exist in registry) and 'filter' (only capability-declared fields), but leaves 'limit' and 'select' completely unaddressed, covering only half the parameters.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('Lista') and resource ('objetos de uma capability catalogada') and notes it maps to the 'list' operation, which distinguishes it reasonably from graph_get and graph_query. It does not explicitly contrast with those siblings, so it stops short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides a prerequisite ('capability_id' must exist in the registry, use graph_discover_capabilities) and a constraint on 'filter', which is useful context for invocation. However, it never explicitly says when to prefer this over graph_query, graph_get, or the many domain-specific list_* siblings.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

graph_permissionsB
Read-onlyIdempotent

Retorna as permissões Microsoft Graph / Azure necessárias para uma consulta, incluindo suporte a delegated e application permission e requisito de licença.

Não executa consulta: serve para validar viabilidade antes de consultar.

ParametersJSON Schema
NameRequiredDescriptionDefault
questionNo
capability_idNo

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, openWorldHint and destructiveHint=false, so safety is covered structurally. The description usefully reinforces that it does not execute the query and states what the response itemizes (delegated/application permissions, license requirement), but adds no cost, rate-limit, or failure-mode context.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences, front-loaded with what it returns and followed by the key behavioral caveat. No padding, though the second sentence could be folded into the first.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema exists, so the description carries the burden of explaining returns; it does list the returned categories, which is helpful. But with 0% schema coverage on two parameters and no output schema, an agent still lacks enough to invoke it confidently with meaningful inputs.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% and neither of the two parameters (question, capability_id) is explained in the description. The phrase 'para uma consulta' hints at the question parameter but gives no guidance on capability_id, its format, or the relationship between the two.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a concrete verb+resource: returns the Microsoft Graph/Azure permissions required for a query, and enumerates the payload (delegated vs application permissions, license requirement). The second sentence distinguishes it from executing siblings like graph_query, though it does not name any sibling directly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

'Não executa consulta: serve para validar viabilidade antes de consultar' gives a clear condition for use — pre-flight validation before an actual query. No alternative tool is named and no exclusions are stated, so it falls short of a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

graph_queryB
Read-onlyIdempotent

Executa consulta parametrizada em fontes que aceitam query (ex.: Azure Resource Graph).

Somente leitura: operadores KQL de escrita ou de acesso externo são bloqueados. 'scope' é validado contra o formato de escopo Azure.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNo
queryNo
scopeNo
capability_idYes

TDQS

B3.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, and openWorldHint, but the description adds substantive behavioral constraints beyond them: KQL write operators and external-access operators are blocked, and 'scope' is validated against the Azure scope format. These are real operational facts an agent needs (query may be rejected on syntax grounds) and go past the annotation surface.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three tight sentences that front-load the core action, then the read-only constraint, then the scope-validation rule. No padding, and each sentence carries distinct information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 4-parameter query tool with 0% schema coverage and no output schema, the definition is thin: it never explains what a 'capability_id' is, what the query is run against, or what shape of result to expect. The scope-validation note helps but does not close the gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% across 4 parameters, so the description carries the full burden. It only clarifies 'scope' (validated against Azure scope format), while 'query', 'limit', and the required 'capability_id' are left entirely unexplained, including how capability_id relates to the query target.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource: 'Executa consulta parametrizada em fontes que aceitam query (ex.: Azure Resource Graph)'. An agent can tell this executes parameterized queries against query-capable sources. However, it does not distinguish itself from the many graph_* siblings (graph_list, graph_get, graph_relationship, graph_answer_identity_question), leaving the boundary fuzzy.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description names the source category ('fontes que aceitam query') but gives no when-to-use guidance versus alternatives such as graph_list, graph_get, or graph_answer_identity_question. No prerequisites, no mention of when this is preferable to a more targeted sibling.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

graph_relationshipB
Read-onlyIdempotent

Consulta relacionamentos de um objeto de diretório (membros, owners, membership transitiva, app role assignments de um principal).

Exemplos de capability_id:

  • graph.groups.members

  • graph.groups.transitive_membership

  • graph.applications.owners

  • graph.service_principals.owners

  • graph.app_role_assignments.assigned_to

ParametersJSON Schema
NameRequiredDescriptionDefault
idYes
limitNo
capability_idYes

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, openWorldHint and destructiveHint=false, so the safety profile is covered. The description adds the domain context of what relationships are exposed but says nothing about pagination via limit, result shape, or whether traversal is transitive by default.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Efficient structure: a single purpose sentence followed by a compact example list, with no filler. Slightly terse given the tool's breadth, but nothing is wasted.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema exists, so the description must carry the full burden of explaining what is returned, how many items, and how limit affects traversal — none of which is stated. For a graph traversal tool with an open-ended capability_id, this is adequate but thin.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% and no parameter is documented in the schema. The description partially compensates by giving valid capability_id example values, which is genuinely useful, but 'id' and 'limit' remain undefined, and the capability_id list is illustrative rather than authoritative.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource: it 'consulta relacionamentos de um objeto de diretório' and enumerates the relationship kinds (membros, owners, transitive membership, app role assignments). However, it does not distinguish itself from close siblings like graph_role_assignments, graph_permissions, or graph_query, which also surface role/relationship data.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The capability_id examples implicitly convey what can be queried and thus hint at when to use the tool, but there is no explicit when-to-use/when-not guidance and no routing to alternatives such as graph_role_assignments or graph_query. Usage must be inferred from the example list.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

graph_role_assignmentsA
Read-onlyIdempotent

Consulta atribuições de privilégio mantendo as fontes separadas:

  • Directory roles e PIM de diretório via Microsoft Graph

  • Azure RBAC via Azure Resource Graph

Não assume que Microsoft Graph cobre Azure RBAC.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNo
include_pimNo
include_azure_rbacNo
include_directory_rolesNo

TDQS

A3.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already establish readOnly, idempotent, openWorld, and non-destructive behavior. The description adds meaningful backend behavior: it separates Directory/PIM data from Azure RBAC data and explicitly warns not to assume Microsoft Graph covers Azure RBAC. It omits pagination, limit behavior, and required permissions, but is stronger than the annotation baseline.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Extremely concise and front-loaded: it states the core query, bullet-lists the source split, and ends with the key warning. Every sentence adds useful information with no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Covers the core source-separation concept and benefits from strong annotations, but with no output schema and 0% schema description coverage it leaves return shape, pagination, limit behavior, and parameter toggles unspecified. It is minimally adequate but incomplete for a four-parameter query tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% for four parameters. The description loosely maps to include_directory_roles, include_pim, and include_azure_rbac by naming the source categories, but it never mentions the limit parameter, defaults, or how the include flags toggle sources, so it fails to compensate for the schema gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Consulta atribuições de privilégio') and clearly scopes the tool to two source families: Directory roles/PIM via Microsoft Graph and Azure RBAC via Azure Resource Graph. The explicit warning that Microsoft Graph does not cover Azure RBAC distinguishes it from graph-only siblings and prevents a common source-assumption error.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Implies usage by describing source separation and the Graph-vs-Azure-RBAC assumption, but does not say when to choose this tool over siblings like list_azure_role_assignments or list_pim_role_states. No exclusions or alternative-selection guidance are provided.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

iam_natural_language_queryC
Read-onlyIdempotent

Interpreta uma pergunta de IAM em linguagem natural e responde com correlação Entra + Azure. Sempre retorna resumo e detalhes compreensíveis, sem depender de frase exata.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNo
questionYes

TDQS

C2.7/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, openWorldHint, and destructiveHint=false, covering the safety profile. The description adds that it always returns a summary and understandable details, and that it doesn't rely on exact phrasing, which is useful behavioral context beyond annotations. However, it doesn't disclose return format specifics, model/interpretation variability, or limitations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two concise sentences, front-loading the core action. No wasted words, though it sacrifices clarity for brevity given the missing details.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a natural-language query tool with no output schema, 0% schema coverage, and no annotations explaining the return structure beyond safety, the description is incomplete. It doesn't explain what 'correlação Entra + Azure' produces, the expected question format, or how results are structured.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, and the description provides no information about the 'question' or 'limit' parameters. The description mentions 'pergunta' implicitly but doesn't explain format, constraints, or what 'limit' controls. With 2 parameters at 0% coverage, the description fails to compensate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a verb+resource ('interpreta uma pergunta de IAM em linguagem natural e responde'), which conveys it's a natural-language query tool for IAM. However, it does not differentiate itself from siblings like timeline_natural_language_query, agent_natural_language_query, pim_natural_language_query, or graph_answer_identity_question, all of which are natural-language query interfaces. The purpose is clear at a high level but lacks sibling differentiation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool versus the many sibling natural-language query tools (timeline, agent, pim, graph). The phrase 'sem depender de frase exata' is a behavioral trait, not a usage condition. No when-to-use or when-not-to-use information is provided.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

identity_360A
Read-onlyIdempotent

Identity 360: visão completa de uma identidade — quem é, a que possui acesso (Entra + Azure, direto/grupo/PIM), como recebeu, ownership e risco (blast radius, findings). Aceita UPN, displayName ou objectId.

ParametersJSON Schema
NameRequiredDescriptionDefault
user_identifierYes

TDQS

A3.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare read-only, open-world, idempotent, and non-destructive behavior. The description adds scope context beyond those annotations: it specifies the aggregation across Entra and Azure, direct/group/PIM access paths, ownership, and risk signals such as blast radius and findings.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two tight sentences with no filler. It front-loads the tool's value proposition and then immediately clarifies accepted identifier formats.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

There is no output schema, but the description enumerates the key data domains returned (identity, access, provenance, ownership, risk). Combined with annotations that cover the safety profile, this is nearly complete, though it does not mention any return structure or limitations.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% for the single required parameter, so the description must compensate. It does so by stating that the identifier accepts UPN, displayName, or objectId, which gives clear semantics for the user_identifier field.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly identifies the tool as providing a complete identity view and enumerates the domains covered (who, access across Entra and Azure, ownership, risk). It is not a tautology, but it does not explicitly differentiate itself from siblings like get_identity_access_summary or compute_identity_blast_radius.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No when-to-use, when-not-to-use, or alternative guidance is provided. The description only states what the tool offers and what inputs it accepts, leaving the agent to infer when this aggregate view is preferable to more focused sibling tools.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_agent_identitiesB
Read-onlyIdempotent

Lista inventário de Agent Identities com owner, blueprint, permissões e risco.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNo
statusNoall

TDQS

B3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, destructiveHint=false and openWorldHint=true, so the safety profile is covered. The description adds only the payload fields (owner, blueprint, permissions, risk) and says nothing about pagination, ordering, or result volume, which for an open-world listing tool would be useful.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single short, front-loaded sentence with no filler; the resource and returned fields come first. It is efficient, though it could carry one more clause of routing or parameter guidance at little cost.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Annotations and the one-line scope cover the basics for a read-only listing tool with no output schema, but the two undocumented parameters and the absence of any when-to-use guidance leave real gaps for an agent choosing among ~50 sibling tools.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% for two parameters (limit, status), and the description mentions neither. It does not explain the default of 50, the pagination behavior, or the accepted values for 'status' (default 'all' implies an enum-like filter that is never documented).

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('Lista') and resource ('Agent Identities') and enumerates the fields the inventory carries (owner, blueprint, permissões, risco), which differentiates it from neighbors like get_agent_relationships or get_agents_by_owner. It stops short of explicitly contrasting itself with those siblings.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no indication of when to use this tool rather than the many adjacent agent/identity tools (get_agent_relationships, get_agents_by_owner, agent_natural_language_query). The agent must infer usage from the name alone.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_application_provenanceA
Read-onlyIdempotent

Diferencia app registrations criadas por usuários no seu tenant das aplicações nativas da Microsoft (first-party) e de apps de terceiros consentidos.

Classificação baseada em appOwnerOrganizationId (sinal autoritativo do diretório), não em heurística de nome.

'provenance' aceita: all, tenant (criadas no seu tenant), microsoft (nativas), thirdparty (terceiros), managedidentity.

Use para perguntas como:

  • "Quais aplicações foram criadas pelos usuários?"

  • "Qual a diferença entre app registrations próprias e nativas?"

  • "Quantas aplicações são nativas da Microsoft?"

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNo
provenanceNoall
include_managed_identitiesNo

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The annotations already declare readOnly/idempotent/non-destructive, so the safety profile is covered. The description adds real behavioral value by disclosing that classification is derived from appOwnerOrganizationId rather than name heuristics, which is a non-obvious implementation trait. It says nothing about result size, pagination, or how the limit default of 100 behaves.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the purpose, then the classification signal, then the parameter values, then example questions — a logical order with no filler. The example-question block is slightly verbose but directly aids tool selection, so it earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

There is no output schema, so the description should ideally sketch the return shape, and two of three parameters remain undocumented. It is adequate for a read-only listing tool because provenance semantics are explained, but the limit/include_managed_identities gap and absent return-format hint leave real holes.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description carries the burden. It fully documents the provenance parameter semantics (all, tenant, microsoft, thirdparty, managedidentity), but says nothing about limit or include_managed_identities, leaving two of three parameters unexplained. Partial compensation, not full.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description gives a specific verb and resource — listing app registrations classified by provenance — and names the authoritative signal used (appOwnerOrganizationId) as well as the four provenance classes. An agent can tell this apart from siblings like summarize_application_provenance or list_applications_without_owners without reading any schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Three concrete example questions make the intended use case unambiguous, and the enumeration of accepted provenance values tells the agent what scoping options exist. It stops short of naming an explicit alternative tool or a when-not-to-use condition, so it is clear context rather than full routing guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_application_secrets_expiringB
Read-onlyIdempotent

Lista secrets expirados ou próximos da expiração em aplicações do Entra ID.

ParametersJSON Schema
NameRequiredDescriptionDefault
daysNo
limitNo

TDQS

B3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, openWorldHint=true and destructiveHint=false, so the safety profile is covered. The description adds only the expiry-scope concept ('expirados ou próximos da expiração'); it does not explain what 'próximos' means, whether results are paginated, or rate-limit behavior. With annotations carrying the safety burden, this is an adequate 3.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with no wasted words; the verb and scope lead. It is efficient, though minimal enough that the missing parameter detail is a consequence of its brevity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple read-only list tool with no output schema, the description is minimally sufficient in stating what is listed. But with two undocumented parameters and no explanation of the expiry window or result capping, an agent lacks the detail needed to call it precisely.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% for both parameters. The description never mentions `days` or `limit`, so it does not explain that `days` defines the 'near expiration' window (default 30) or that `limit` caps results (default 100). With zero schema coverage, the description is expected to compensate and does not.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (Lista) and resource (secrets de aplicações do Entra ID) with a clear qualifier (expirados ou próximos da expiração). An agent can tell what it returns without opening the schema. However, it does not differentiate itself from any sibling or scope the result set, so it stops short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives no when-to-use context, no prerequisites, and no exclusions or alternative tools. Usage is only implied by the name. There is no sibling covering secrets expiry, so an agent has no guidance on routing or timing.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_applications_without_ownersC
Read-onlyIdempotent

Lista aplicações sem owner definido no Entra ID.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNo

TDQS

C2.8/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, destructiveHint=false, and openWorldHint=true, covering the entire safety profile. The description adds nothing beyond restating the purpose – no pagination behavior, no result-size expectations tied to the limit default of 50, no indication of how ownership is determined.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with no filler or redundancy. It is efficient, though the brevity borders on under-specification for the tool's scope.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple read-only listing tool with safety annotations already present, the description is minimally adequate. However, with no output schema and an undocumented limit parameter, it does not tell the agent what a result entry contains or how the list is bounded.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The only parameter, limit (default 50), has 0% schema description coverage and is not mentioned anywhere in the description. With one undocumented parameter the description fails to compensate for the coverage gap, leaving the agent to guess whether limit caps items, applications, or pages.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ("Lista") plus resource ("aplicações") and an explicit filter scope ("sem owner definido no Entra ID"), so an agent knows exactly what set is returned. It does not, however, differentiate itself from the near-identical sibling list_objects_without_owner, leaving ambiguity about which orphan-detection tool applies to applications versus general objects.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance, no prerequisites, and no mention of alternatives such as list_objects_without_owner, get_owned_objects, or list_application_provenance. The agent must infer the right context from the name alone.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_azure_role_assignmentsB
Read-onlyIdempotent

Lista role assignments Azure RBAC (User, Group, Service Principal e Managed Identity).

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNo

TDQS

B3.3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, destructiveHint=false, idempotentHint=true, and openWorldHint=true, so the safety and idempotency profile is covered. The description adds only the principal types enumerated, which is marginal context beyond annotations. No mention of pagination, result size, or subscription scoping.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

One short sentence, front-loaded with the verb and resource, with no filler. Nothing to trim.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a read-only list tool with annotations covering safety, the description is minimally adequate, but with 0% schema coverage on the parameter and no output schema, an agent lacks guidance on pagination, scoping, and result shape. Adequate but with clear gaps.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

One parameter (limit, default 50) with 0% schema description coverage. Baseline for low coverage kicks in, but with only a single parameter and the default value visible in the schema, the description adds nothing about what limit controls (page size vs. total results) — a gap, but small in scope.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (Lista) and resource (role assignments Azure RBAC) and enumerates the principal types covered (User, Group, Service Principal, Managed Identity). Sibling tools like list_privileged_azure_role_assignments and list_orphan_azure_role_assignments exist but the description doesn't differentiate from them, capping it at 4.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No indication of when to use this vs. the many sibling tools (list_privileged_azure_role_assignments, list_orphan_azure_role_assignments, list_deny_assignments, etc.). The parenthetical principal-type list implies scope but provides no routing guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_deny_assignmentsC
Read-onlyIdempotent

Lista deny assignments do Azure.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNo

TDQS

C2/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, destructiveHint=false and openWorldHint=true, covering the safety profile. The description adds nothing beyond that – no mention of pagination, scope of Azure subscriptions, or result volume.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

It is a single short sentence with no filler, though the mixed-language phrasing reduces clarity. Conciseness is fine; the issue is under-specification, not verbosity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema, no annotations explaining result scope, and a truncated description. For an Azure Graph API listing tool with a limit parameter, the description omits whether results are tenant-wide, subscription-scoped, or paginated.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0% for the single 'limit' parameter (default 50). The description does not mention pagination or any bound on results, leaving the agent without meaning for the one parameter it must pass.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description 'Lista deny assignments do Azure' mixes Portuguese and English and merely restates the tool name. It states a resource (deny assignments) but no scope or distinguishing detail versus many sibling list_* tools dealing with Azure role assignments.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines1/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No when-to-use, when-not-to-use, or alternative is named. Among dozens of list_* siblings including list_azure_role_assignments and list_privileged_azure_role_assignments, the agent gets no routing guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_disabled_users_with_active_rolesA
Read-onlyIdempotent

Lista usuários desabilitados no Entra ID que ainda possuem role assignments diretos ativos.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNo

TDQS

A3.7/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true and destructiveHint=false, so the safety profile is fully covered. The description adds scope value by restricting to Entra ID, disabled accounts, and direct (non-group-derived) assignments, but says nothing about result volume, pagination, or whether tenants are queried with elevated permissions.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

One front-loaded sentence with zero filler; the discriminating filter ('diretos ativos') is placed at the end where it does the most routing work.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a read-only listing tool with annotations covering safety and no output schema, the description covers what is enumerated but omits any hint of output shape, result size, or pagination behavior tied to the undocumented limit parameter.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% and the single 'limit' parameter (default 20) is never mentioned in the description, so no meaning is added beyond the schema. However, the parameter is a conventional result-cap integer, which keeps this near the baseline rather than a clear failure.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (lista), resource (usuários desabilitados no Entra ID) and a precise qualifying condition (que ainda possuem role assignments diretos ativos). The word 'diretos' distinguishes it from assignment-inheriting siblings such as list_users_with_direct_permissions and list_orphan_azure_role_assignments.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The condition it targets (disabled accounts holding live direct assignments) implies a dormant-privilege audit use case, but there is no explicit when-to-use statement, no exclusion of overlapping siblings like list_users_with_direct_permissions, and no prerequisite guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_entra_usersB
Read-onlyIdempotent

Lista usuários do Microsoft Entra ID com nome, UPN e e-mail.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNo

TDQS

B3.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, destructiveHint=false, so safety profile is covered. The description adds the specific returned attributes (nome, UPN, e-mail), which is useful behavioral detail beyond the annotations, but does not mention pagination or default limits.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single, front-loaded sentence with no waste. It states verb, resource, and returned fields efficiently.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple read-only list tool with annotations covering safety, the description is adequate but incomplete: no pagination/limit semantics and no differentiation from the many sibling user-listing tools leave context gaps.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% for the single 'limit' parameter. The description provides no information about the limit parameter, its default, or pagination behavior, so it fails to compensate for the schema gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

Clear verb 'Lista' and resource 'usuários do Microsoft Entra ID', with the returned fields (nome, UPN, e-mail) specified. It does not differentiate itself from the many sibling user-listing tools (list_users, list_azure_role_assignments, etc.), so it falls short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No when-to-use, when-not-to-use, or alternative tools are named. With many sibling list_* user tools, an agent gets no routing guidance at all.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_graph_critical_application_permissionsC
Read-onlyIdempotent

Lista aplicações/service principals com Microsoft Graph Application Permissions críticas.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNo

TDQS

C2.6/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare this as a safe, idempotent, read-only operation, so the safety profile is covered. Beyond that, the description adds nothing about what 'critical' means, whether results are paginated, how permissions are determined, or any rate limits. With annotations handling safety, the bar is lower, but the description still provides almost no behavioral context.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, front-loaded sentence with no filler. It is concise, though extremely terse by design; it could have used one more sentence for usage without becoming bloated.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a list tool with read-only annotations, no output schema, and a single parameter, the description is nearly minimal. It omits what 'critical' means, how results are ordered or paginated, and when to use this tool over its many siblings. The definition leaves substantial gaps for an agent to fill.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, but there is only one optional parameter (limit with a default of 100). The description does not mention the limit parameter or its units/meaning, which is a minor gap given the single parameter. Baseline 3 applies since 0-param/low-param tools get a baseline of 4, but the description could have added a sentence about limiting results.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a clear verb (Lista) and resource (aplicações/service principals with critical Microsoft Graph Application Permissions). However, it is a single sentence with no differentiation from siblings like list_application_provenance or graph_permissions, and the 'critical' qualifier is undefined, leaving the precise scope ambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool versus alternatives such as graph_permissions, list_application_provenance, or detect_toxic_combinations. The description offers no context, prerequisites, or exclusions, forcing the agent to infer usage from the name alone.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_objects_without_ownerB
Read-onlyIdempotent

Lista objetos sem owner (Groups, Applications, Service Principals, Agents, Blueprints).

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNo

TDQS

B3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, openWorldHint and destructiveHint=false, so the safety profile is fully covered. The description adds the useful detail of which object classes are scanned, but says nothing about return shape, ordering, or pagination behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with the enumeration in parentheses — no filler. It is tight, though the parenthetical list is the only substance beyond the name.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a read-only list tool with no output schema, the description identifies the scope well but leaves the agent without return-format, ordering, or pagination context, and the lone limit parameter is entirely unexplained.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% and there is one parameter (limit, default 200) that the description never mentions. With low coverage the description should compensate by explaining what limit controls and what happens when it is omitted or exceeded; it does neither.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

Names a specific verb (Lista) and resource (objetos sem owner) and enumerates the object types covered (Groups, Applications, Service Principals, Agents, Blueprints). This implicitly distinguishes it from the narrower sibling list_applications_without_owners, though the overlap is never called out explicitly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance, no prerequisite or scope note, and no mention of the closely overlapping sibling list_applications_without_owners that would let an agent choose between them. Usage is only inferable from the name and enumeration.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_orphan_azure_role_assignmentsC
Read-onlyIdempotent

Lista role assignments órfãos (principal não resolvido no tenant visível).

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNo

TDQS

C2.4/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, destructiveHint=false, and openWorldHint=true, covering the safety and side-effect profile. The description adds only the scoping condition for 'orphan', which is useful, but says nothing about pagination, tenant-visibility limits, or the fact that unresolved principals may be a transient/edge state. Against rich annotations, a 2 reflects minimal added behavioral context.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

One short parenthetical sentence that front-loads the resource and the defining filter. Nothing extraneous, though it is arguably too terse given the missing guidelines and parameter detail.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a list tool with one undocumented parameter, no output schema, and a dense sibling neighborhood of similar role-assignment tools, the definition is under-specified. The agent lacks guidance on selection, return shape, and limit behavior.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% and the single 'limit' parameter has no description in either schema or description. The description does not explain what 'limit' bounds, defaults, or whether it is a page size. With the only parameter undocumented, the description fails to compensate for the coverage gap, though a single param keeps the damage limited.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific resource ('role assignments órfãos') and qualifies it with 'principal não resolvido no tenant visível'. However, it is in Portuguese while tool names and siblings are in English, and it doesn't differentiate this tool from close siblings like list_azure_role_assignments or list_privileged_azure_role_assignments. Purpose is discernible but not sharp.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No when-to-use or when-not-to-use guidance. The agent must infer from the name which sibling to pick, and the description offers no routing hints despite ~8 closely related role-assignment tools in the sibling list.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_pim_role_statesB
Read-onlyIdempotent

Lista atribuições privilegiadas classificadas por estado: Active, Eligible e Permanent.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNo
include_permanentNo

TDQS

B3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, destructiveHint=false and openWorldHint=true, so the agent knows this is a safe, repeatable read. The description adds the state-classification grouping, which is mild context, but says nothing about pagination, result size, or auth requirements.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short lines that lead with the verb and resource, then enumerate the state categories. Nothing redundant, though the trailing colon fragment is slightly clipped.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a read-only list tool whose annotations already cover the safety profile, the description adequately conveys what is listed. However, with two undocumented parameters and no output schema, an agent still lacks pagination and filtering behavior needed to call it confidently.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It hints at the 'Permanent' category, which loosely maps to include_permanent, but the limit parameter is entirely unexplained and no formats or defaults are clarified. Compensation is only partial.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('Lista') and resource ('atribuições privilegiadas') and adds the classification dimension (Active, Eligible, Permanent). It is clear what the tool returns, though it never distinguishes itself from the closely-related sibling get_pim_state_summary.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool versus alternatives such as get_pim_state_summary or pim_natural_language_query. No prerequisites, exclusions, or context are given, leaving usage entirely to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_privileged_azure_role_assignmentsB
Read-onlyIdempotent

Lista role assignments privilegiados em Azure RBAC (Owner, User Access Administrator, Contributor).

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNo

TDQS

B3.1/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, openWorldHint and destructiveHint=false, so the safety profile is covered. The description adds useful domain context by defining exactly which roles count as 'privileged', but it says nothing about the default limit of 50 or the volume/scope of results returned, so it adds partial value beyond the structured fields.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence that states the action, the scope, and the defining role list with zero filler. Nothing is wasted.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description must carry more of the return-value burden, yet it does not describe the shape of results (assignment objects, principals, scope) or the effect of the limit parameter. The role enumeration helps, but the definition remains only minimally adequate for a security-focused listing tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% and the single 'limit' parameter carries only a default (50) with no explanation. The description does not mention the limit, its default, or what happens when it is omitted, so it fails to compensate for the schema gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('Lista') and resource ('role assignments privilegiados em Azure RBAC') and enumerates the exact privileged roles (Owner, User Access Administrator, Contributor), which cleanly distinguishes it from the sibling list_azure_role_assignments that presumably returns all assignments. It is clear and precise, though it does not explicitly name the sibling it narrows from.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use or when-not-to-use guidance. Given the dense sibling set (list_azure_role_assignments, list_orphan_azure_role_assignments, get_role_risk_score, list_pim_role_states), an agent gets no explicit signal about when this privileged-only view is preferred over the general role assignment listing.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_privilege_timeline_eventsC
Read-onlyIdempotent

Lista eventos de ganho/perda/ativação de privilégios no período informado.

ParametersJSON Schema
NameRequiredDescriptionDefault
daysNo
limitNo
actionNoall
providerNoall

TDQS

C2.9/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, destructiveHint=false and openWorldHint=true, so safety and repeatability are covered. The description usefully adds which event categories are returned (gain/loss/activation), but says nothing about ordering, pagination, or truncation behavior tied to the limit parameter.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with no filler; the scope constraint (period) is stated immediately. It is efficient, though too terse to carry the parameter burden the schema leaves open.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Four undocumented parameters, no output schema, no enums, and no sibling routing are all left unexplained. The one sentence covers the basic idea but is not sufficient to invoke the tool correctly with action/provider/limit combinations.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% for all four parameters (days, limit, action, provider). The description only indirectly hints at the time window ('período informado') and the event types that the action filter likely controls; it adds nothing about limit (default 200) or provider.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific verb and resource ('Lista eventos de ganho/perda/ativação de privilégios') and scopes it to a time period, so the purpose is unambiguous. It offers no differentiation from closely named siblings such as get_identity_privilege_timeline or summarize_privilege_timeline, and the Portuguese phrasing mismatches the English tool ecosystem.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no statement of when to use this tool versus the timeline siblings (get_identity_privilege_timeline, summarize_privilege_timeline, export_privilege_timeline_report). Usage must be inferred entirely from the name and the phrase 'no período informado'.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_public_ip_resourcesB
Read-onlyIdempotent

List Public IP resources. This does not prove that a workload is actually internet-exposed.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNo

TDQS

B3.1/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint/idempotent/openWorld/destructive=false, covering the safety profile. The description adds a genuinely useful interpretive caveat that a public IP alone does not prove internet exposure, but says nothing about pagination, result volume, or what the scan covers.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences, purpose front-loaded and the qualifier second. Every sentence carries weight and there is no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple, argument-light list tool with no output schema, the description covers the essential purpose. It still leaves gaps around what a returned 'Public IP resource' contains and how the caution should translate into follow-up queries.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% for the single 'limit' parameter, so the schema does not document it. The description provides no information about the parameter (default value, meaning, paging implications), failing to compensate for the coverage gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('List Public IP resources'), which is unambiguous and distinct from the identity/graph siblings in the list. It does not, however, explicitly differentiate itself from the other inventory-style tools (e.g. list_resources, get_management_group_inventory).

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no indication of when to reach for this tool versus related inventory tools like list_resources, nor any prerequisites or workflow context. The caveat sentence hints at interpretation but gives no selection guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_resource_groupsC
Read-onlyIdempotent

Lista Resource Groups com filtros read-only por nome.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNo
name_containsNo

TDQS

C2.6/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint and destructiveHint=false, so 'read-only' adds no new information. The description discloses nothing beyond the annotations: no pagination behavior, no default limit of 20, no statement of what a name filter matches (substring vs prefix).

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with no filler, which is structurally clean, but its brevity here reflects under-specification rather than economy given two undocumented parameters.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema exists, so the description should at least indicate what the returned resource groups carry and how paging works; instead it stops after purpose. For a listing tool with 0% parameter coverage, this is too thin for reliable invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description carries the whole burden, yet it only gestures at a name filter and says nothing about 'limit', its default of 20, or whether name_contains is a substring or exact match. It partially compensates for one of the two parameters and leaves the other undocumented.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Lista Resource Groups') plus the filter axis (by name), so the agent knows it enumerates resource groups rather than counting them. It does not, however, distinguish itself from the sibling get_resource_groups_count, which is the natural confusion point.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to use this versus the sibling get_resource_groups_count or list_resources, and no statement of prerequisites or scope limits. The agent must infer usage purely from the name.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_resourcesB
Read-onlyIdempotent

List Azure resources using controlled read-only filters. Use full Azure resource type when filtering, e.g. microsoft.compute/virtualmachines.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNo
name_containsNo
resource_typeNo

TDQS

B3.3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, openWorldHint, and destructiveHint=false, so the safety profile is covered. The description adds that filtering is controlled and read-only, but does not mention pagination, result volume, or scope. This is extra context but not rich behavioral detail.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two concise sentences, front-loaded with the action, and the example is directly useful. No wasted text.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a three-parameter listing tool with no output schema and no parameter descriptions in the schema, the description is incomplete: it does not cover result limits, name filtering semantics, or whether results are scoped to a subscription or tenant.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It explains resource_type format with an example ('microsoft.compute/virtualmachines'), which is valuable, but leaves 'limit' and 'name_contains' undefined and does not explain default behavior.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource: 'List Azure resources' with read-only filters. This is clear, but it does not differentiate from siblings like list_public_ip_resources or list_resource_groups, which also list Azure resource subsets.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to use this tool versus alternatives such as list_resource_groups. The example about using full Azure resource type is a parameter tip, not a usage guideline.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_top_blast_radiusC
Read-onlyIdempotent

Ranking de identidades por maior blast radius.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNo

TDQS

C2.9/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already specify readOnlyHint=true, idempotentHint=true, and destructiveHint=false, so the agent knows this is a safe, repeatable read operation. The description adds minimal behavioral context beyond what the annotations provide. No information is given about what the ranking is based on, the return format, or any limits. A baseline 3 is appropriate given the annotations cover safety, but the description does not enrich behavioral transparency.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, concise sentence with no wasted words, and the key information (ranking, metric) is front-loaded. It is appropriately sized for a simple tool.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the simplicity (1 optional param, no output schema) the description is not complete enough: it does not explain what 'blast radius' means in this context, how identities are ranked, or what the output contains. There is no output schema, so the description should clarify the return structure, but it does not. This is a significant gap for an agent trying to interpret results.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema has only one parameter ('limit') with no description (coverage 0%), so the description should compensate but does not. However, the parameter name is self-explanatory and has a default value. With only one simple parameter, the baseline is 4 per rules (0 params = 4), but here there is 1 param and 0% coverage, so a 3 reflects that the description adds no meaning and the schema is also silent.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific action (ranking identities) and a specific metric (blast radius), which is understandable. However, it does not distinguish this tool from its sibling 'compute_identity_blast_radius', leaving ambiguity about which tool to use for blast radius analysis. The purpose is clear but lacks sibling differentiation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is provided on when to use this tool versus the related 'compute_identity_blast_radius'. There is no indication of prerequisites, context, or exclusions. The agent is left to infer usage from the name alone.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_usersC
Read-onlyIdempotent

Lista usuários do Entra ID visíveis para a identidade autenticada.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNo
disabled_onlyNo
name_containsNo

TDQS

C2.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, destructiveHint=false, and openWorldHint=true, so the safety profile is covered. The description adds the scoping constraint that results are limited to what the authenticated identity can see, which is real behavioral context. It does not cover pagination, return shape, or rate limits.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

One short sentence with no waste and the scope constraint front-loaded. Efficient, though very minimal for a tool with undocumented parameters.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema and zero schema description coverage means the description must carry more weight. It omits filter semantics, whether any filters are mutually exclusive, pagination behavior, and how it differs from sibling user-listing tools. It is under-specified for the surrounding tool set.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% for all three parameters (limit, disabled_only, name_contains), and the description says nothing about them. With low coverage the description is supposed to compensate, but it does not mention filters, defaults, or pagination behavior at all.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (Lista) and resource (usuários do Entra ID), which is clear. However, it overlaps heavily with siblings list_entra_users and list_users_with_direct_permissions, and the description does not differentiate this tool from those alternatives.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides no when-to-use guidance, no alternatives, and no conditions that would select this tool over list_entra_users or the other user-listing siblings. The scope phrase 'visíveis para a identidade autenticada' is the only contextual hint, but it is not framed as usage guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_users_with_direct_permissionsC
Read-onlyIdempotent

Lista usuários com role assignments diretos no Azure RBAC.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNo
disabled_onlyNo

TDQS

C2.9/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, destructiveHint=false and openWorldHint=true, so the safety profile is fully covered by structured data. The description adds one genuine behavioral nuance – that only DIRECT (not inherited) assignments count – but says nothing about pagination, result size, or required permissions.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with no filler or repetition. It is well-structured, though its brevity contributes to the missing parameter and usage detail rather than being a flaw of the sentence itself.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 2-parameter tool with 0% schema description coverage, no output schema, and dozens of closely related siblings, a one-line description is insufficient. It omits the filter semantics (disabled_only), result limiting, and any tie-breaker against neighboring list tools.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% and both parameters (limit, disabled_only) are undocumented in the schema. The description never mentions either parameter, so it does not compensate for the coverage gap – an agent cannot tell that disabled_only filters to disabled accounts or how limit behaves from this text.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource ('Lista usuários com role assignments diretos no Azure RBAC'), and the 'diretos' qualifier meaningfully narrows it away from inherited assignments. It does not, however, explicitly name or differentiate itself from siblings like list_users, list_azure_role_assignments, or list_disabled_users_with_active_roles.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool versus the many adjacent siblings (list_users, list_azure_role_assignments, list_privileged_azure_role_assignments, get_user_effective_azure_access). The agent must infer the distinction from the tool name alone.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_users_without_mfaC
Read-onlyIdempotent

Lista usuários sem MFA registrado.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNo

TDQS

C2.8/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint and destructiveHint=false, so the safety profile is covered. The description adds nothing beyond that: no pagination behaviour, no note on how 'sem MFA registrado' is determined, no scale/rate context.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with zero filler, which is appropriate for the scope. It is concise rather than wasteful, though it verges on under-specification.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple read-only, zero-required-parameter listing tool with annotations and no output schema, the minimum is arguably met. Still missing is any hint about result size limits or how to narrow the list.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% and the single 'limit' parameter (default 100) is never mentioned in the description. Since the schema does not document it either, the description fails to compensate for the coverage gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a clear verb-plus-resource ('Lista usuários sem MFA registrado') that precisely defines the result set. It does not, however, distinguish this from near-neighbours such as list_users_with_weak_authentication or get_user_authentication_methods.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no statement of when to use this tool, no prerequisites, and no mention of the many sibling tools that also cover authentication posture. The agent must infer applicability from the name alone.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_users_with_passkeyC
Read-onlyIdempotent

Lista usuários com passkey/FIDO2 registrado.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNo

TDQS

C2.6/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true and destructiveHint=false, so the safety profile is covered. The description adds no behavioral context beyond that: no mention of result size, pagination, truncation at the limit, or whether the read is tenant-scoped. It is neither contradictory nor additive.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

One short, front-loaded sentence with no filler or redundancy. It is tight and readable, though the sentence is so sparse that it barely carries information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple read-only list tool with no output schema, an agent mostly needs the filter condition and the result-shaping parameter behavior. The filter condition is communicated, but with 0% schema coverage and no output schema, the meaning of 'limit' and any pagination semantics are left entirely unexplained.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% for the single parameter 'limit' (integer, default 100). The description says nothing about the limit, pagination, or whether 100 is a ceiling, so it fails to compensate for the undocumented parameter.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The sentence names a specific verb and resource (lists users) with a filter qualifier (registered passkey/FIDO2), so the purpose is understandable. However, it is essentially a restatement of the tool name plus the FIDO2 synonym, adding no scope detail (tenant-wide? paginated? does it include disabled users?) and no differentiation from siblings such as list_users or list_users_without_mfa.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance and no alternatives named. Given the dense sibling set of user-listing tools (list_users, list_users_with_direct_permissions, list_users_without_mfa, list_users_with_weak_authentication), the agent must infer that this is the passkey-registration view; the description does not state that or draw any boundary.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_users_with_weak_authenticationC
Read-onlyIdempotent

Lista usuários com métodos fracos registrados (sms/voice/email/password), com evidência técnica.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNo
include_mfa_registeredNo

TDQS

C2.9/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, destructiveHint=false and openWorldHint, so the safety profile is covered. The description adds only the vague phrase 'com evidência técnica', which hints that results carry supporting evidence but never says what that evidence is or how it is shaped.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single compact sentence with the scope qualifier front-loaded via the parenthetical enumeration. Efficient, though the language differs from the English sibling vocabulary, which slightly weakens matching.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a filtered-list tool with two undocumented parameters and no output schema, the description should explain the parameter effects and roughly what comes back. It instead stops at scope, leaving the agent to guess how limiting or MFA inclusion behave.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% for both parameters, so the description carries the burden and fails it. It never explains what 'limit' controls or what 'include_mfa_registered' toggles, even though that boolean materially changes the result set.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Lista usuários') and defines the scope precisely by enumerating which authentication methods count as weak (sms/voice/email/password). It is distinguishable from adjacent siblings like list_users_without_mfa, though it never names them or the contrast explicitly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No when-to-use guidance, no prerequisites, and no alternatives named. An agent must infer from the name alone whether this or list_users_without_mfa / assess_privileged_mfa is the right call for a given authentication-posture question.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

pim_natural_language_queryC
Read-onlyIdempotent

Interpreta perguntas de PIM em linguagem natural e retorna resposta com resumo e detalhes.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNo
questionYes

TDQS

C2.9/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, destructiveHint=false, and openWorldHint=true, so the safety profile is covered by structured data. The description adds one behavioral fact beyond that: the return payload contains both a summary and details, which is useful given there is no output schema. It says nothing about the open-world nature of the answers, latency, or how the 'limit' affects returned detail.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with no filler; the purpose is delivered immediately. It is efficient, though the terseness leaves no room for the usage and parameter detail the definition otherwise lacks.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

This is an open-world natural-language query tool with no output schema and two undocumented parameters, and it sits among several near-identically named siblings. The description covers only the bare intent plus a hint at the return shape, leaving routing between the natural-language query family and the meaning of 'limit' unresolved.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% and neither parameter is described in the schema. The description only implies that 'question' is free-form natural language; it never explains what 'limit' bounds (rows, entities, detail items) or what the default of 20 means in practice. For a zero-coverage schema, the description does not compensate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: it interprets PIM questions in natural language and returns a response with summary and details. The 'PIM' domain qualifier implicitly distinguishes it from the sibling natural-language-query tools (iam_natural_language_query, timeline_natural_language_query, agent_natural_language_query), though the differentiation is never made explicit.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool versus the four sibling natural-language-query tools, nor any stated prerequisites, scope limits, or examples of acceptable questions. The agent must infer usage entirely from the tool name and the 'PIM' qualifier.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

run_agent_assessmentC
Read-onlyIdempotent

Executa assessment específico de Agent Identities e retorna ranking de risco.

ParametersJSON Schema
NameRequiredDescriptionDefault
top_risksNo

TDQS

C2.9/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, destructiveHint=false, so the safety profile is covered. The description adds only that the result is a 'ranking de risco'; it says nothing about scoring method, cost, latency, or whether the openWorldHint implies external calls. Some added value, but thin.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with no filler; the verb, scope and return artifact come in order. It is efficient, though terse to the point of leaving capability undefined.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For an assessment tool with no output schema, one undocumented parameter, and dozens of near-neighbor siblings, the description should explain what the risk ranking contains and how it relates to the adjacent graph/IAM assessment tools. As written, an agent lacks enough to call it confidently or interpret results.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The single parameter top_risks has a default of 10 but 0% schema description coverage, and the description never mentions it. The name is largely self-explanatory (number of top-ranked risks), which prevents a 1, but the description fails to compensate for the coverage gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and a narrow resource ('assessment específico de Agent Identities') plus the output artifact ('ranking de risco'), which separates it from broader siblings like run_iam_assessment and run_enterprise_identity_audit. It does not name those siblings explicitly, but the 'Agent Identities' scope is distinctive enough for selection.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No when-to-use, when-not-to-use, or prerequisites are given. With ~40 siblings including list_agent_identities, graph_assessment, run_iam_assessment and run_enterprise_identity_audit, the agent gets no routing guidance about when this assessment is the right call versus those alternatives.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

run_enterprise_identity_auditB
Read-onlyIdempotent

Executa auditoria IAM enterprise consolidada (Entra + Azure + subscriptions + management groups).

ParametersJSON Schema
NameRequiredDescriptionDefault
top_risksNo

TDQS

B3.1/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, destructiveHint=false, idempotentHint=true and openWorldHint=true, so the safety and idempotency profile is covered. The description adds the scope of the audit (Entra, Azure, subscriptions, management groups), which is meaningful context, but does not disclose cost, latency, rate limits, or authentication requirements for what is likely an expensive and broad read operation across the tenant.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single, front-loaded sentence with no filler. It is appropriately sized for a description that, however, under-specifies behavior; conciseness itself is good here since there is no wasted text.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a broad, open-world consolidated IAM audit with no output schema, one undocumented parameter, and no annotation-independent behavioral context, the description is too thin. It does not explain what the audit returns, what the top_risks parameter does, or when to prefer it over overlapping siblings such as run_iam_assessment and identity_360.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

There is one parameter, top_risks, with 0% schema description coverage and a default of 15. The description says nothing about this parameter, so the agent cannot tell what top_risks controls (number of surfaced risks, ranking cutoff, etc.) or what the default does. With low coverage and an undocumented parameter, the description should have compensated but does not.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb (Executa auditoria) and a specific resource/scenario (IAM enterprise consolidada across Entra, Azure, subscriptions, management groups). This is clearly differentiated from the many sibling tools that audit narrow slices (e.g., list_users, list_azure_role_assignments, get_environment_summary) because it declares itself the consolidated enterprise-wide audit. It is clear but slightly less precise than a 5 because it does not name the specific outputs or distinguish exactly when to prefer it over run_iam_assessment.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is implied by the scope (an enterprise-wide consolidated audit). However there is no explicit when-to-use guidance and no named alternatives among siblings, notably run_iam_assessment, iam_natural_language_query, and identity_360, which appear to overlap. The agent must infer that this tool is the broad enterprise audit rather than the narrower sibling assessments.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

run_iam_assessmentB
Read-onlyIdempotent

Executa um IAM Assessment no ambiente e retorna riscos principais e plano de remediação.

ParametersJSON Schema
NameRequiredDescriptionDefault
top_risksNo

TDQS

B3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, openWorldHint=true and destructiveHint=false, so the safety profile is covered and the description does not contradict it. The description adds the useful fact that the result includes both risks and a remediation plan, but says nothing about scan duration, breadth of the environment scanned, or cost of an open-world sweep.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

One compact sentence that front-loads the action and then the outputs; nothing is wasted. It is arguably too terse for a broad assessment tool, but there is no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description usefully names the two return components (top risks, remediation plan), which partially covers return-value expectations. However it leaves the only input parameter undocumented and gives no sense of scope or runtime for an open-world assessment.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% and the single parameter 'top_risks' (default 10) is never mentioned in the description, so the agent gets no semantic guidance on what the number controls. The description does not compensate for the coverage gap at all.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Executa um IAM Assessment no ambiente') plus the outputs it produces (top risks and a remediation plan), so the agent knows what it gets. It does not, however, distinguish itself from near-neighbour siblings such as run_agent_assessment, run_enterprise_identity_audit or graph_assessment, which the agent must disambiguate on its own.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no statement of when to run this assessment versus the several sibling assessment/audit tools (run_agent_assessment, run_enterprise_identity_audit, graph_assessment), nor any prerequisites or scope conditions. Usage must be fully inferred from the tool name alone.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

search_official_guidanceB
Read-onlyIdempotent

Return curated official Microsoft Learn references relevant to an Azure topic or best-practice question.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNo
topicYes

TDQS

B3.3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare readOnlyHint=true, idempotentHint=true, openWorldHint=true, and destructiveHint=false, so the safety profile is clear. The description adds that results are 'curated' and 'official', which is useful behavioral context, but it doesn't disclose rate limits, result freshness, or response structure. With annotations covering safety, a 3 is appropriate.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, well-structured sentence that is front-loaded with the main action and resource. Every word earns its place, with no redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple search tool with no output schema and minimal parameters, the description covers the core purpose but lacks details on parameter usage, result format, and when to prefer this tool over others. It is adequate but incomplete for an agent that must decide between this and sibling tools.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the schema only provides types and defaults (e.g., limit defaults to 4). The description mentions 'topic or best-practice question' but doesn't explain the expected format, constraints, or the role of the 'limit' parameter. The description fails to compensate for the low schema coverage, leaving both parameters ambiguous.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('Return') and resource ('curated official Microsoft Learn references'), and scopes it to an 'Azure topic or best-practice question'. It is clearly distinguishable from the sibling tools, which focus on identity, access, and resource data rather than documentation. The only gap is that it doesn't clarify the format of the references or the exact nature of the 'curated' set.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage for topics or best-practice questions but does not explicitly state when to use this tool versus alternative search or knowledge tools. It gives context ('official Microsoft Learn') but no exclusions or prerequisites, leaving the agent to infer.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

summarize_application_provenanceA
Read-onlyIdempotent

Resumo executivo da procedência das aplicações do tenant: quantas estão sob sua governança (criadas no tenant) versus pré-provisionadas pela Microsoft, com percentuais e foco de governança.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A3.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, openWorldHint and non-destructive, so the safety profile is fully covered. The description adds that output is count/percentage oriented with a governance focus, which is useful context beyond the annotations, but omits anything on data freshness or the scope of the count.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single, front-loaded sentence that names the subject first and then the comparison axis. The parenthetical 'criadas no tenant' clarifies the jargon immediately; slightly dense but no wasted clauses.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-argument, read-only aggregation with no output schema, the description supplies enough for an agent to know what the summary contains (governance vs pre-provisioned split with percentages). Only the relationship to the sibling list tool is left unstated.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters and the schema is empty at 100% coverage, so there is nothing for the description to disambiguate. Baseline 4 applies with no parameter gaps to compensate for.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific verb and resource ('Resumo executivo da procedência das aplicações do tenant') and adds the exact axis of comparison: tenant-governed vs Microsoft pre-provisioned, with percentages. It does not, however, explicitly distinguish itself from the close sibling list_application_provenance, so the agent must infer that this is the aggregated/summary variant.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The 'executive summary' framing implies a reporting use case, but the description never states when to pick this over list_application_provenance (the raw list) or other assessment tools. Usage is inferable from the noun 'resumo', not stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

summarize_privilege_timelineB
Read-onlyIdempotent

Resumo da timeline de privilégios (ganhos, revogações, ativações e anomalias).

ParametersJSON Schema
NameRequiredDescriptionDefault
daysNo

TDQS

B3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true and destructiveHint=false, so the safety profile is fully covered externally. The description adds a useful statement of what the summary aggregates (gains, revocations, activations, anomalies), but does not disclose the default lookback window, aggregation granularity, or whether anomalies are inferred.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with a parenthetical enumeration; no filler. It is efficient, though the parenthetical list could be read as slightly redundant with the resource name.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description carries the burden of conveying the return shape; the enumeration of included categories partially does this. However, it omits the time window (the only parameter) and the level of aggregation, leaving meaningful gaps for a reporting tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% and the sole parameter ('days', default 30) is never referenced in the description. Because the schema does not document the lookback window, the description needed to compensate and does not, leaving the time scope of the summary unstated.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a clear verb ('Resumo' = summary) and resource ('timeline de privilégios'), and enumerates the summarized content (gains, revocations, activations, anomalies). It distinguishes itself implicitly from the listing siblings (list_privilege_timeline_events, get_identity_privilege_timeline) by being an aggregation rather than a raw event list, but never names those alternatives.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance, no mention of prerequisites, and no routing to the closely related siblings list_privilege_timeline_events or get_identity_privilege_timeline. The agent must infer that a summary is preferred over a raw listing whenever an overview is wanted.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

timeline_natural_language_queryC
Read-onlyIdempotent

Interpreta perguntas de auditoria temporal de privilégios em linguagem natural.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNo
questionYes

TDQS

C2.6/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnly, idempotent, openWorld and non-destructive, so the safety profile is covered structurally. The description adds nothing beyond that: no mention of how questions are interpreted, whether results are ranked/paginated, or any limits on query scope.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with no wasted words, but it is so terse that conciseness comes at the cost of substance. Adequate minimum-viable brevity rather than efficient completeness.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a natural-language query tool with no output schema, 0% parameter coverage, and several closely related siblings, the description is too thin. It omits what the returned timeline contains, how to phrase questions, and how it differs from the other *_natural_language_query tools.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the schema documents neither 'question' nor 'limit'. The description only implies that 'question' takes natural language; it never explains the 'limit' parameter (default 20, presumably result count) or the expected question format.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('Interpreta') and resource ('perguntas de auditoria temporal de privilégios em linguagem natural'), which scopes it to the timeline/privilege-audit domain rather than generic querying. However, it does not explicitly distinguish itself from the four sibling natural-language query tools (pim_, iam_, agent_, graph_answer_), leaving domain boundaries partly to inference.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No when-to-use guidance, no prerequisites, and no mention of alternatives such as iam_natural_language_query or pim_natural_language_query. The agent must guess which NL-query tool fits a given question from the name alone.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 66 tool updatesv0.1.0
    • First observedagent_natural_language_query
    • First observedassess_privileged_mfa
    • First observedcompute_identity_blast_radius
    • First observeddetect_toxic_combinations
    • First observedexport_privilege_timeline_report
    • First observedget_agent_relationships
    • First observedget_agents_by_owner
    • First observedget_authentication_methods_summary
    • First observedget_authentication_strength_summary
    • First observedget_environment_summary
    • First observedget_identity_access_summary
    • First observedget_identity_privilege_timeline
    • First observedget_license_posture
    • First observedget_management_group_inventory
    • First observedget_owned_objects
    • First observedget_pim_state_summary
    • First observedget_resource_groups_count
    • First observedget_role_risk_score
    • First observedget_subscription_direct_access_summary
    • First observedget_subscriptions_count
    • First observedget_tenant_licenses
    • First observedget_user_authentication_methods
    • First observedget_user_effective_azure_access
    • First observedgraph_answer_identity_question
    • First observedgraph_assessment
    • First observedgraph_directory_objects
    • First observedgraph_discover_capabilities
    • First observedgraph_get
    • First observedgraph_list
    • First observedgraph_permissions
    • First observedgraph_query
    • First observedgraph_relationship
    • First observedgraph_role_assignments
    • First observediam_natural_language_query
    • First observedidentity_360
    • First observedlist_agent_identities
    • First observedlist_application_provenance
    • First observedlist_application_secrets_expiring
    • First observedlist_applications_without_owners
    • First observedlist_azure_role_assignments
    • First observedlist_deny_assignments
    • First observedlist_disabled_users_with_active_roles
    • First observedlist_entra_users
    • First observedlist_graph_critical_application_permissions
    • First observedlist_objects_without_owner
    • First observedlist_orphan_azure_role_assignments
    • First observedlist_pim_role_states
    • First observedlist_privilege_timeline_events
    • First observedlist_privileged_azure_role_assignments
    • First observedlist_public_ip_resources
    • First observedlist_resource_groups
    • First observedlist_resources
    • First observedlist_top_blast_radius
    • First observedlist_users
    • First observedlist_users_with_direct_permissions
    • First observedlist_users_with_passkey
    • First observedlist_users_with_weak_authentication
    • First observedlist_users_without_mfa
    • First observedpim_natural_language_query
    • First observedrun_agent_assessment
    • First observedrun_enterprise_identity_audit
    • First observedrun_iam_assessment
    • First observedsearch_official_guidance
    • First observedsummarize_application_provenance
    • First observedsummarize_privilege_timeline
    • First observedtimeline_natural_language_query

TDQS

C2.7/5.0

Scored across 66 tools

Disambiguation2/5

Many tools overlap heavily: list_users and list_entra_users both list Entra users; the generic graph_* tools (graph_list, graph_get, graph_query, graph_relationship) duplicate specific list_*/get_* tools; and multiple assessment/NL-query tools (run_iam_assessment, run_enterprise_identity_audit, graph_assessment, iam_natural_language_query, graph_answer_identity_question) have unclear boundaries. An agent would struggle to pick the right tool without opening descriptions.

Naming Consistency4/5

Nearly all tool names use lower snake_case with a verb_noun pattern (list_users, get_role_risk_score, run_iam_assessment, graph_list), so the set is predictable overall. Minor deviations like identity_360, graph_directory_objects, and the *_natural_language_query suffix are readable but break the strict verb_noun convention.

Tool Count1/5

66 tools is extreme for an MCP server and far exceeds the 15-tool guideline. The surface contains many redundant wrappers and overlapping assessment/query tools, so the count is not justified by distinct capabilities.

Completeness4/5

For a read-only Azure identity and security assessment server, the surface is broad: it covers users, RBAC, PIM, MFA, applications, agents, licenses, resources, and ownership. Minor gaps exist around remediation or write operations, but those may be intentionally out of scope.

Maintenance

ActivityNo data
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • -
    license
    Not graded
    quality
    Not graded
    maintenance
    Provides secure access to Microsoft Entra ID (Azure AD) resources including users, devices, and applications through Microsoft Graph API. Enables querying organizational data with comprehensive audit logging to Azure Blob Storage.
    -
  • F
    license
    Not graded
    quality
    C
    maintenance
    Enables AI assistants to inspect and audit Azure Landing Zones by inventorying resources, auditing tagging, evaluating policy compliance, and detecting infrastructure drift, all in read-only mode.
    -