Skip to main content
Glama
shigechika

entraadm-mcp

by shigechika

signin_failure_stats

Aggregate tenant-wide Entra ID sign-in failures to identify top error codes, users, apps, source IPs, and password-spray suspects.

Instructions

Tenant-wide sign-in failure aggregation -- the Entra ID counterpart to the RADIUS failure patrol.

Time-bounded (ENTRAADM_DEADLINE, default 45 s): on a wide window or a busy day the scan stops early and capped=true marks the counts as a lower bound; narrow hours for a full count.

Aggregates failed sign-ins across the whole tenant into four views: top AADSTS error codes (with the same meaning annotations as signin_logs), top failing users, top applications, and top source IPs. spray_suspects flags any IP with failed sign-ins against 5 or more distinct users -- Entra's smart lockout is per-account, so a low-and-slow password spray from one IP across many accounts does not trip it the way a brute force against one account does; this is the observation a per-account view cannot make on its own. This mirrors the KeyCloak-side spray detection this fleet already relies on; neither the official Microsoft MCP Server for Enterprise nor Graph itself offers this aggregation.

Read-only (AuditLog.Read.All application permission, or -- for azure-cli auth -- the Reports Reader directory role). Graph cannot filter sign-ins on status/errorCode server-side, so this walks up to max_pages of the full sign-in log for the window and aggregates client-side -- capped=true means the page budget ran out before the window was fully scanned, so the counts below are a sample of the window, not a census of it.

Args: hours: How far back to look, clamped to [1, 720] (30 days). max_pages: Page budget (default: ENTRAADM_MAX_PAGES_DEFAULT).

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
hoursNo
max_pagesNo

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observedv0.1.0

TDQS

A4.5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does so thoroughly: read-only, required permission (AuditLog.Read.All) or azure-cli role (Reports Reader), a 45 s ENTRAADM_DEADLINE budget, and the crucial caveat that Graph can't filter server-side so results are a client-side sample, with `capped=true` marking lower-bound counts. This is exactly the behavioral disclosure an agent needs.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the core purpose well, but the middle paragraphs get verbose and partly promotional ('This mirrors the KeyCloak-side spray detection this fleet already relies on; neither the official Microsoft MCP Server for Enterprise nor Graph itself offers this aggregation'), which does not help an agent decide or call correctly.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Despite having no output schema, the description enumerates the return shape (four views plus `spray_suspects`), explains the `capped` flag semantics, and covers auth/permission requirements and time-budget behavior. An agent has everything needed to invoke and interpret it.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, so the description must compensate, and it does: `hours` is documented as clamped to [1,720] and `max_pages` as a page budget tied to capping. It gives the default of max_pages only by env-var reference rather than a concrete number, a minor gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a precise verb and resource — tenant-wide aggregation of failed sign-ins into four named views — and explicitly contrasts itself with the per-event `signin_logs` it borrows annotations from. The scope (tenant-wide, failure-only, grouped) is unmistakable from the sibling set.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explains the core use case (detecting low-and-slow spray that per-account lockout misses) and gives actionable tuning guidance: 'narrow hours for a full count.' It references `signin_logs` for per-event detail, but stops short of an explicit 'use this instead of X when Y' routing statement for the other siblings.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.