Skip to main content
Glama
shigechika

entraadm-mcp

by shigechika

entraadm-mcp

English | 日本語

MCP server for Microsoft Entra ID sign-in and audit-log triage. Read-only.

Why this instead of the official Microsoft MCP Server for Enterprise

Microsoft ships an official MCP Server for Enterprise for Entra ID data. It is a good fit for an interactive admin at a keyboard, and is not a fit for an unattended triage bot:

  • Delegated auth only. The official server does not support app-only (client credentials) auth, so it cannot run headless behind a service account. entraadm-mcp is built for that case: app-only in production, with a delegated (az login) fallback for local development.

  • A general-purpose Graph query tool, not a fixed tool set. The official server exposes one tool that lets the model construct arbitrary GET/schema-discovery calls against Microsoft Graph. That is flexible for a human, and awkward to put behind an allow-list for an automated triage profile. entraadm-mcp exposes nine fixed, read-only tools instead.

  • No AADSTS translation. Sign-in failures come back as raw error codes; triage still needs a lookup table. entraadm-mcp annotates every sign-in failure with what the code actually means.

  • No cross-request aggregation. Microsoft Graph itself cannot filter sign-ins on status/errorCode server-side, and has no built-in password-spray view. signin_failure_stats aggregates client-side and flags IPs with failed sign-ins against many distinct users — the pattern Entra's per-account smart lockout does not catch on its own.

Related MCP server: Microsoft Graph MCP Server

Tools

Tool

What it answers

health_check

Is Graph reachable, and can this credential read sign-in logs?

get_user

Is this account enabled, synced from on-prem, and what are its licenses?

signin_logs

Why did this user's sign-in fail (or succeed), with the AADSTS code translated?

signin_failure_stats

Tenant-wide failure aggregation: top error codes, users, apps, source IPs, and password-spray suspects

signin_success_stats

Tenant-wide success aggregation by source IP: IPs shared by several accounts, legacy-auth (SMTP/IMAP) successes — the view that finds a breach

signin_by_ip

Every sign-in from one source IP: who got in from it, who was tried, and when

directory_audits

Who changed what in the directory (block/unblock, attribute edits), and when?

get_user_auth_methods

Is MFA actually registered for this account?

daily_brief

One-call summary combining signin_failure_stats and directory_audits

Every tool is read-only. Write operations (unblocking an account, resetting a password, revoking a session) are out of scope for this server.

Auth model

Two auth modes, selected by which environment variables are set:

Mode

When

Env vars

app-only

All three set

ENTRAADM_TENANT_ID, ENTRAADM_CLIENT_ID, ENTRAADM_CLIENT_SECRET

azure-cli

None set

(uses the current az login session)

Setting one or two of the three app-only variables is a configuration error and the server refuses to start, rather than silently falling back to a different auth mode than intended.

Required Graph permissions

Tool(s)

Permission

Notes

get_user (base fields)

User.Read.All

signin_logs, signin_failure_stats, signin_success_stats, signin_by_ip, directory_audits, get_user's sign_in_activity field

AuditLog.Read.All (app-only) or the Reports Reader directory role (delegated)

get_user_auth_methods

UserAuthenticationMethod.Read.All

App-only only; not available under delegated (az login) auth in a typical tenant role assignment

A missing permission never crashes a tool. It degrades that tool (or that one field) to {"error": "...", "missing_permission": "..."} with a human-readable explanation of what role or permission is needed, so health_check and every other tool stay usable even before full permissions are granted.

Setup

uv tool install entraadm-mcp
# or
pip install entraadm-mcp

Configuration

Set the three app-only variables for production/unattended use:

export ENTRAADM_TENANT_ID=00000000-0000-0000-0000-000000000000
export ENTRAADM_CLIENT_ID=00000000-0000-0000-0000-000000000000
export ENTRAADM_CLIENT_SECRET=your-client-secret

Or leave all three unset and run az login first for local development.

Optional:

# Default page cap for the log-scanning tools (1-50, default 50).
export ENTRAADM_MAX_PAGES_DEFAULT=50
# Wall-clock budget per tool call in seconds (default 45; 0 disables). A hosted MCP
# client cuts a call off at about 60 s, so a scan stops at the budget and returns
# what it has with capped=true.
export ENTRAADM_DEADLINE=45

Usage

Claude Code (plugin)

/plugin marketplace add shigechika/entraadm-mcp
/plugin install entraadm-mcp@entraadm-mcp

Claude Code (manual)

Add to .mcp.json:

{
  "mcpServers": {
    "entraadm-mcp": {
      "type": "stdio",
      "command": "uvx",
      "args": ["entraadm-mcp"],
      "env": {
        "ENTRAADM_TENANT_ID": "${ENTRAADM_TENANT_ID:-}",
        "ENTRAADM_CLIENT_ID": "${ENTRAADM_CLIENT_ID:-}",
        "ENTRAADM_CLIENT_SECRET": "${ENTRAADM_CLIENT_SECRET:-}"
      }
    }
  }
}

Direct execution

entraadm-mcp

CLI options

Option

Effect

--version

Print the version and exit

--check

Resolve auth, probe Graph reachability and sign-in log access, print a report, exit 0 (or 1 on config error)

Notes

  • Coverage contract. Every result that walks a paged Graph collection carries a capped boolean when its window was not fully scanned — a partial scan is never reported as if it were exhaustive.

  • found: false is not an error. get_user and get_user_auth_methods answer a nonexistent account with {"found": false, ...}, not an error key — a typo'd userPrincipalName should never look like this server being broken.

  • Retention. Entra ID P1 retains sign-in and directory audit logs for 30 days. A window beyond that returns an empty result, not an error.

Development

uv sync --dev
uv run pytest -v
uv run ruff check .
uv run ruff format --check .

Live smoke test

uv run python scripts/smoke_test.py

Read-only, no payloads printed (tool names/statuses/row counts only), and bounded (small explicit windows/page caps) — nothing here writes to the tenant or scans more than a day of logs.

Releasing

This repository uses release-please driven by Conventional Commits. Merge a feat:/fix: PR to main, and release-please opens (or updates) a release PR; merging that PR tags a release and triggers the publish pipeline (PyPI, MCP Registry).

License

MIT

Available Tools

9 tools
daily_briefA

One-call morning-patrol summary: sign-in failures, spray suspects, and admin actions.

Combines signin_failure_stats and directory_audits into one result with a compact summary on top, matching the shape of this fleet's other daily_brief tools. A permission failure in one section degrades only that section's contribution to summary -- the other section still returns in full.

Runs both sections synchronously in one tool call, unlike the sibling gwsadm-mcp's job+poll daily_brief. If this proves too slow for a tenant's sign-in volume against the client's tool-call timeout, port that job+poll pattern here (tracked in this repo's CLAUDE.md Roadmap).

Args: hours: How far back to look, clamped to [1, 720] (30 days). max_pages: Page budget passed to both sections (default: ENTRAADM_MAX_PAGES_DEFAULT). samples: Reserved for a future drill-down sample size; currently unused.

ParametersJSON Schema
NameRequiredDescriptionDefault
hoursNo
samplesNo
max_pagesNo

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description fully compensates for missing annotations by disclosing: synchronous execution, per-section permission degradation behavior, page budget handling, hour clamping, and the unused samples parameter. It even notes the performance consideration and potential porting path. No explicit read-only statement, but a summary tool is implicitly non-destructive and the behavior is well specified.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is well-structured: a crisp one-line summary up front, then behavior details, then a clean Args section. Every sentence adds value, though a few extra details (e.g., roadmap reference) could be trimmed. It is appropriately sized for the tool's complexity without fluff.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the moderate complexity and missing output schema, the description is largely complete: it covers behavior, parameter semantics, and error degradation. It references the return shape via 'matching the shape of this fleet's other daily_brief tools' but does not explicitly enumerate return fields, which is a minor gap. However, it clearly explains the combined nature and the summary on top, sufficient for an agent to call it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With 0% schema description coverage, the description carries the full burden and excels: 'hours' is explained with clamp range, 'max_pages' with default and purpose, and 'samples' is explicitly flagged as reserved/unused. Each parameter gets meaningful semantics beyond the raw type.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb-resource pair ('combines signin_failure_stats and directory_audits into one result') and names the exact content (sign-in failures, spray suspects, admin actions). It clearly distinguishes itself from siblings by explicitly naming the two source tools and contrasting with the job+poll pattern.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description clearly implies when to use it (morning patrol, one-call summary) and differentiates from the sibling gwsadm-mcp's job+poll daily_brief. It does not explicitly list exclusions or alternatives beyond that sibling, but the purpose statement makes it obvious that this is the combined-synchronous variant.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

directory_auditsA

Who did what to the directory, and when -- the operator-side counterpart to signin_logs.

Every admin action against a user object (block/unblock, password reset, role assignment, attribute edits) appears here, naming the actor (initiated_by) and the affected object(s) (target_resources). This is the record a manual "unblock and reset" intervention -- like the one that closed the 2026-08-21 case this server exists to shorten -- leaves behind; it is how a later triage can tell "already handled by a human" from "still open".

user, when given, matches audits where that account is either the initiator or a target resource. Graph's directoryAudits endpoint only supports server-side $filter on the initiator (initiatedBy/user/userPrincipalName), not on targetResources, so this fetches the full time window and matches both sides client-side -- a window with many unrelated admin actions can need a larger max_pages budget than signin_logs/signin_failure_stats to find one specific user's audits; capped=true warns when that budget ran out before the window was fully scanned.

Read-only (AuditLog.Read.All application permission, or -- for azure-cli auth -- the Reports Reader directory role). Entra ID retains directory audit logs for 30 days, same as sign-in logs.

Time-bounded: the scan stops after ENTRAADM_DEADLINE seconds (default 45) and returns what it has with capped=true. Over several days the log is dominated by device-registration noise ("Update device"), so pass category (for example UserManagement, RoleManagement, GroupManagement, ApplicationManagement) to have Graph filter server-side; that keeps a multi-day window inside the budget.

Args: category: Only this Graph audit category (letters only; default: all). user: Restrict to audits naming this userPrincipalName as actor or target (default: all). hours: How far back to look, clamped to [1, 720] (30 days). top: Maximum records to return, clamped to [1, 500]. max_pages: Page budget (default: ENTRAADM_MAX_PAGES_DEFAULT).

ParametersJSON Schema
NameRequiredDescriptionDefault
topNo
userNo
hoursNo
categoryNo
max_pagesNo

TDQS

A4.9/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations present, the description carries the full behavioral burden and does so: read-only, required AuditLog.Read.All application permission or Reports Reader role for azure-cli auth, 30-day Entra ID retention, a 45s ENTRAADM_DEADLINE scan cap returning partial results with capped=true, and the important client-side vs. server-side filter distinction (only initiatedBy supports $filter). This is well beyond what any structured field supplies.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core 'who did what to the directory, and when' framing, and the Args block is dense and useful. It is longer than strictly necessary, though -- the aside about the 2026-08-21 case this server exists to shorten adds narrative color rather than invocation-relevant information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema, no annotations, and a scan-heavy tool with a page budget and timeout, so the description has to be self-sufficient -- and it is. It covers return-relevant fields (initiated_by, target_resources, capped), permissions, retention window, and the operational tradeoff of narrowing by category, leaving no gap an agent would need filled before calling it.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate fully and it does: every one of the 5 args is documented with meaning plus constraints (hours clamped to [1,720], top clamped to [1,500], category letters-only, max_pages default reference, user matches either actor or target). The user param's dual actor/target semantics in particular is only knowable from the prose.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Names a specific resource (directory audits) and states exactly what it captures: admin actions against user objects (block/unblock, password reset, role assignment, attribute edits), with actor and target fields. It explicitly positions itself as the 'operator-side counterpart to signin_logs,' so an agent can separate it from the sibling sign-in tools without opening a schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives explicit routing guidance: use this for admin/intervention history rather than sign-in events, and pass category (e.g. UserManagement) to keep multi-day windows within budget. It also names the sibling tools it shares a page budget model with and states the condition under which a larger max_pages is needed (many unrelated admin actions in-window).

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_userA

One account's identity/lifecycle state -- the first thing to check on any triage report.

account_enabled=false means the account itself is the whole story; stop there. A stale last_password_change_date_time alongside a fresh "wrong password" complaint (AADSTS50126 in signin_logs) is the most common on-the-ground pattern: the password changed or expired somewhere, and a cached credential on one device is now stale. on_premises_sync_enabled=true means this account is synced from an on-premises directory (Entra Connect) -- Entra is a downstream copy of its password via Password Hash Sync, not the source of truth. licenses names are resolved from the tenant's SKU catalog (/subscribedSkus, page budget from ENTRAADM_MAX_PAGES_DEFAULT); licenses_capped: true appears only when that scan was cut short before resolving one of this account's own licenses -- when present, one or more licenses entries is a raw skuId rather than a friendly name.

sign_in_activity needs an additional Graph read (AuditLog.Read.All application permission, or -- for azure-cli auth -- the Reports Reader directory role) beyond what the rest of this tool needs. If that permission is missing, every other field above still returns and sign_in_activity alone degrades to {"error": ..., "missing_permission": "AuditLog.Read.All"}.

A nonexistent account is a normal answer, not a tool failure: the result is {"found": false, "user_principal_name": upn} rather than an error key, so a typo'd UPN in a triage report cannot be mistaken for this tool being broken.

Read-only (User.Read.All application permission, or an equivalent delegated read). Requires an exact userPrincipalName, not a display name or partial match.

Args: upn: The account's userPrincipalName, e.g. "user@example.edu".

ParametersJSON Schema
NameRequiredDescriptionDefault
upnYes

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and excels. It discloses the meaning of key fields (account_enabled, on_premises_sync_enabled, licenses_capped), the graceful degradation of sign_in_activity without the required permission, and that a nonexistent account returns {'found': false} rather than an error. It also states the read-only nature and exact UPN requirement. No contradictions.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Though long, the description is densely packed with prerequisite triage knowledge—field semantics, permission nuances, and edge cases—and every sentence adds value. It is structured with clear paragraphs and code-formatting for fields, making it scannable. The front-loaded purpose statement ensures the agent quickly knows what the tool does.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with no output schema, the description covers the essential return fields (account_enabled, last_password_change_date_time, on_premises_sync_enabled, licenses, licenses_capped, sign_in_activity), their meanings, and expected edge behavior. It also describes permission requirements and the graceful degradation path. Nothing an agent needs to call it and interpret results is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage for the only parameter, upn, is 0%, but the description fully compensates: 'Args: upn: The account's userPrincipalName, e.g. "user@example.edu".' It also stresses the need for an exact match, not display name. This gives the agent everything needed to format the argument correctly.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with 'One account's identity/lifecycle state -- the first thing to check on any triage report.' This is a specific verb+resource (get identity/lifecycle state) and explicitly positions it as the initial step, distinguishing it from sibling tools that focus on sign-in logs, auth methods, or audits. The purpose is unmistakable and actionable.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives clear context: 'first thing to check' and instructs to 'stop there' if account_enabled=false. It also explains a common pattern for stale passwords. However, it never explicitly names alternative tools or states when not to use this tool (e.g., 'for auth methods, use get_user_auth_methods'). The order is implied but not contrasted with siblings.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_user_auth_methodsA

Registered authentication methods for one account -- is MFA actually set up?

mfa_registered answers "would this account survive a password-spray hit": True iff at least one non-password method is registered (Authenticator app, phone, FIDO2 security key, Windows Hello, a temporary access pass, software OATH token, or a platform credential/passkey). password itself is excluded from that count -- every account has one, so its presence alone says nothing about MFA coverage.

A nonexistent account is a normal answer, not a tool failure: the result is {"found": false, "user_principal_name": upn} rather than an error key, matching get_user's contract.

Read-only (UserAuthenticationMethod.Read.All application permission). This endpoint is app-only only: it is not exposed to delegated (azure-cli) auth under this tenant's current role assignment, so it degrades to a permission error under azure-cli auth even when other tools work.

Args: upn: The account's userPrincipalName.

ParametersJSON Schema
NameRequiredDescriptionDefault
upnYes

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure. It explicitly states the read-only permission, the app-only limitation that leads to permission errors under azure-cli, the behavior for nonexistent accounts (returns a specific structure instead of an error), and the precise semantics of mfa_registered. It does not detail the full response structure for existing accounts, which is a minor gap, but the key behaviors are well covered.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is well-structured with a clear opening statement, a code-formatted explanation of the key field, a note on nonexistent accounts, and a security/permission caveat. It is informative but not bloated; each paragraph serves a distinct purpose. It could be slightly more concise, but the structure aids readability.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description fully covers the MFA-checking purpose and the nonexistent-account case, but it does not specify the full response format for an existing account beyond the mfa_registered boolean. Since there is no output schema, an agent may not know if the response includes a list of methods, their types, or other fields. This is a notable gap for a tool that returns a data structure, though it may be sufficient for the primary use case of checking MFA.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema provides only the parameter name 'upn' with no description (0% coverage). The description compensates by stating 'Args: upn: The account's userPrincipalName.' This adds the meaning and domain context, making it clear what value to provide. While it doesn't elaborate on format or constraints, the parameter is simple and the explanation is sufficient.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a clear, specific purpose: retrieving registered authentication methods for an account and determining whether MFA is actually set up. It explains the key field (mfa_registered) and distinguishes itself from sibling tools that deal with sign-in logs or audits. The verb and resource are unambiguous, and the contrast with other tools is implicit but effective.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides clear context on when to use this tool (to check MFA coverage) and includes a practical caveat about app-only auth under azure-cli. It references get_user's contract for nonexistent accounts, offering a form of alternative comparison. However, it does not explicitly name siblings like signin_logs or signin_failure_stats as alternatives, leaving some usage inference to the agent.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

health_checkA

Fleet-standard health probe: service/version/status plus two independent Graph probes.

graph confirms Microsoft Graph is reachable at all (GET /users with $top=1 -- needs only User.Read.All, the minimum permission every deployment of this server needs anyway). signin_probe additionally confirms the current credential can read sign-in logs -- the permission every other tool here except get_user depends on. Both probes always run, independently of each other: a tenant that has AuditLog.Read.All but not (yet) the baseline User.Read.All would otherwise have this report "Graph unreachable" -- a fabricated diagnosis, since Graph plainly is reachable if the other probe succeeds. status is derived from the two outcomes: healthy when both succeed, degraded when exactly one does (Graph is reachable but some permission is missing), error only when neither does.

Read-only. Always returns the same keys regardless of outcome (detail is null on success, a translated message on failure), so a caller never has to branch on which keys are present.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4.5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description fully carries the behavioral burden — and it succeeds admirably. It discloses read-only status, that both probes always run independently, and the exact status derivation (healthy/degraded/error), plus the output contract (always same keys, detail null on success, translated message on failure). This is thorough, non-obvious behavioral disclosure.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is long (~130 words) but every sentence earns its place — the independence rationale, the derived-status logic, and the output contract are all essential behavioral details that prevent misdiagnosis. It is front-loaded with the core purpose and only moderately verbose, justified by the nuance it conveys.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no parameters and no output schema, the description carries a heavy completeness burden and mostly meets it: it explains all three status values, the constant key contract, and the translated-message behavior. The only small gap is that it references the return keys ('detail', 'status', 'graph', 'signin_probe') without enumerating the full key set a caller should expect.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has zero parameters with an empty schema, so there is nothing for the description to compensate for. Per the baseline for 0-param tools, a 4 is appropriate; there is no parameter documentation gap to fill.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description gives a specific verb and resource ('Fleet-standard health probe: service/version/status plus two independent Graph probes') and then defines each component precisely. It distinguishes itself from the sibling data tools (get_user, signin_logs, etc.) by being the diagnostic probe rather than a data-retrieval tool.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explains that signin_probe validates 'the permission every other tool here except get_user depends on', which implicitly tells the agent this tool verifies prerequisites before relying on sibling tools. However, it never explicitly states 'use this when you need to check if deployments are healthy' or gives an explicit when-not-to-use directive, so it falls short of a full 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

signin_by_ipA

Every sign-in from one source IP: who got in from it, who was tried, and when.

The follow-up to a spray_suspects or shared_ips hit: Graph can filter sign-ins on ipAddress server-side, so this is one cheap query, not a log walk. users summarises the IP per account (successes, failures, first/last seen, up to 50 accounts) over every row fetched; events lists the newest top entries that match result ("all" by default, or "success" / "failure"), each carrying the account name and the same AADSTS annotation as signin_logs. capped=true means the page budget or the deadline ran out before the window was fully read; events_truncated=true means more matching rows were read than top returns (the users summary still counts them).

Read-only (AuditLog.Read.All application permission, or -- for azure-cli auth -- the Reports Reader directory role).

Args: ip: The source IPv4 or IPv6 address, exactly as the sign-in log shows it. hours: How far back to look, clamped to [1, 720] (30 days). result: "all" (default), "success", or "failure" -- which events to list. top: Maximum events to return, clamped to [1, 500]. max_pages: Page budget (default: ENTRAADM_MAX_PAGES_DEFAULT).

ParametersJSON Schema
NameRequiredDescriptionDefault
ipYes
topNo
hoursNo
resultNoall
max_pagesNo

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does so thoroughly: it declares read-only status, states the exact permission/role required (AuditLog.Read.All application permission or Reports Reader role), and explains operational signals (capped, events_truncated) beyond anything in structured fields.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with purpose and behavior before the Args block. The prose is dense and slightly wordy in places (e.g., the elaborated users/events description), but each sentence conveys usable information rather than filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no annotations and no output schema, the description still explains return structure (users summary with up to 50 accounts, events with AADSTS annotations, capped/events_truncated flags), permissions, and all parameter semantics. Nothing an agent needs to invoke it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate, and it documents all five parameters: ip format, hours clamp [1,720], result enum values, top clamp [1,500], and max_pages default. This is exactly the compensation required when the schema is silent.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The opening sentence states a specific verb+resource+scope: 'Every sign-in from one source IP: who got in from it, who was tried, and when.' It also implicitly separates itself from signin_logs by framing the results as IP-scoped, and names upstream triggers (spray_suspects, shared_ips).

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Clearly positions itself as 'the follow-up to a spray_suspects or shared_ips hit' and explains that Graph filters on ipAddress server-side, making it a cheap query rather than a log walk. It does not explicitly name signin_logs as the alternative to use for non-IP-scoped queries, so the when-not is only implied.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

signin_failure_statsA

Tenant-wide sign-in failure aggregation -- the Entra ID counterpart to the RADIUS failure patrol.

Time-bounded (ENTRAADM_DEADLINE, default 45 s): on a wide window or a busy day the scan stops early and capped=true marks the counts as a lower bound; narrow hours for a full count.

Aggregates failed sign-ins across the whole tenant into four views: top AADSTS error codes (with the same meaning annotations as signin_logs), top failing users, top applications, and top source IPs. spray_suspects flags any IP with failed sign-ins against 5 or more distinct users -- Entra's smart lockout is per-account, so a low-and-slow password spray from one IP across many accounts does not trip it the way a brute force against one account does; this is the observation a per-account view cannot make on its own. This mirrors the KeyCloak-side spray detection this fleet already relies on; neither the official Microsoft MCP Server for Enterprise nor Graph itself offers this aggregation.

Read-only (AuditLog.Read.All application permission, or -- for azure-cli auth -- the Reports Reader directory role). Graph cannot filter sign-ins on status/errorCode server-side, so this walks up to max_pages of the full sign-in log for the window and aggregates client-side -- capped=true means the page budget ran out before the window was fully scanned, so the counts below are a sample of the window, not a census of it.

Args: hours: How far back to look, clamped to [1, 720] (30 days). max_pages: Page budget (default: ENTRAADM_MAX_PAGES_DEFAULT).

ParametersJSON Schema
NameRequiredDescriptionDefault
hoursNo
max_pagesNo

TDQS

A4.5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does so thoroughly: read-only, required permission (AuditLog.Read.All) or azure-cli role (Reports Reader), a 45 s ENTRAADM_DEADLINE budget, and the crucial caveat that Graph can't filter server-side so results are a client-side sample, with `capped=true` marking lower-bound counts. This is exactly the behavioral disclosure an agent needs.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the core purpose well, but the middle paragraphs get verbose and partly promotional ('This mirrors the KeyCloak-side spray detection this fleet already relies on; neither the official Microsoft MCP Server for Enterprise nor Graph itself offers this aggregation'), which does not help an agent decide or call correctly.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Despite having no output schema, the description enumerates the return shape (four views plus `spray_suspects`), explains the `capped` flag semantics, and covers auth/permission requirements and time-budget behavior. An agent has everything needed to invoke and interpret it.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, so the description must compensate, and it does: `hours` is documented as clamped to [1,720] and `max_pages` as a page budget tied to capping. It gives the default of max_pages only by env-var reference rather than a concrete number, a minor gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a precise verb and resource — tenant-wide aggregation of failed sign-ins into four named views — and explicitly contrasts itself with the per-event `signin_logs` it borrows annotations from. The scope (tenant-wide, failure-only, grouped) is unmistakable from the sibling set.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explains the core use case (detecting low-and-slow spray that per-account lockout misses) and gives actionable tuning guidance: 'narrow hours for a full count.' It references `signin_logs` for per-event detail, but stops short of an explicit 'use this instead of X when Y' routing statement for the other siblings.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

signin_logsA

One user's recent sign-in events, AADSTS-annotated.

The most direct answer to "why can't this person log in": each entry's error_code_meaning translates the raw AADSTS code (e.g. 50126 -> "invalid credentials (wrong password)") so triage rarely needs a second lookup. result filters client-side after the Graph fetch (Graph cannot filter sign-ins on status/errorCode server-side): "failure" (the default) keeps only failed attempts, "success" keeps only clean ones, "all" keeps everything.

Because the filter is client-side, this walks pages until it has collected top matching entries or exhausts max_pages -- a mostly- successful user can otherwise mean paging through hundreds of rows to find a handful of failures. capped=true means the page budget ran out (or top was reached) before the whole window was scanned; a low match count alongside capped=true is evidence of "no more found within the budget", not "no more exist".

Read-only (AuditLog.Read.All application permission, or -- for azure-cli auth -- the Reports Reader directory role). Entra ID P1 retains sign-in logs for 30 days; hours beyond that returns an empty result, not an error.

Args: user: The account's userPrincipalName. hours: How far back to look, clamped to [1, 720] (30 days). result: "failure" (default), "success", or "all". top: Maximum matching entries to return, clamped to [1, 500]. max_pages: Page budget for the client-side filter walk (default: ENTRAADM_MAX_PAGES_DEFAULT).

ParametersJSON Schema
NameRequiredDescriptionDefault
topNo
userYes
hoursNo
resultNofailure
max_pagesNo

TDQS

A5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden and excels. It discloses read-only permissions (AuditLog.Read.All or Reports Reader), retention limits (30 days), clamping behavior, client-side filtering, page-walking mechanics, and the meaning of capped=true. It even clarifies that low match counts with capped=true indicate a budget limit, not absence of data. This is a model of transparency.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is well-structured, front-loading the core purpose, then providing necessary nuance about filtering and paging, then permissions, then a clean parameter list. Every sentence contributes to correct usage; there is no filler or redundancy. It is long but appropriately so for a tool with complex client-side behavior.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the complexity (client-side filtering, pagination, permission requirements, retention window, edge cases) and the absence of an output schema, the description is remarkably complete. It covers all scenarios an agent needs to invoke the tool correctly and interpret results, including the capped flag and empty results. Nothing critical is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, but the description compensates thoroughly. It explains each parameter in detail: user's userPrincipalName, hours clamped to [1,720], result options with default 'failure', top clamped to [1,500], and max_pages with default. It also explains the functional meaning of result (client-side filter) and the cap semantics, providing far more than the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description explicitly states the tool returns 'one user's recent sign-in events' and frames it as 'the most direct answer to why can't this person log in'. It identifies the AADSTS annotation as a key value-add, distinguishing it from sibling tools like signin_failure_stats that likely provide aggregated statistics. The purpose is specific and immediately actionable.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives explicit usage context: it is positioned as the go-to for individual login troubleshooting, and it explains important behavioral caveats like client-side filtering that require adjusting expectations. It implies when not to use it (e.g., for stats, use signin_failure_stats) and details what happens when the page budget is exhausted (capped=true). Alternatives are not explicitly named but clearly indicated by context.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

signin_success_statsA

Tenant-wide successful sign-in aggregation by source IP -- the view that finds a breach.

signin_failure_stats shows who is being attacked; this shows whether anyone got in. The breach signature is one source IP signing in successfully as several different accounts, most often over a legacy protocol (clientAppUsed such as "Authenticated SMTP" or "IMAP4", which carry no MFA). shared_ips lists the IPs with successes for min_distinct_users or more distinct accounts, most-shared first (up to 50 IPs, shared_ips_capped when more qualified; account names up to 25 per IP), with the client apps and the countries seen. legacy_auth_users lists the accounts that succeeded over a legacy protocol at all, with how many IPs and countries they came from.

A campus NAT, a VDI farm or a shared proxy also puts many accounts behind one IP, so a shared IP is a lead, not a verdict: the caller excludes its own egress ranges and reads the client apps and countries before calling anything a breach. Graph cannot filter sign-ins on status/errorCode server-side, so like signin_failure_stats this walks the sign-in log for the window and aggregates client-side. The walk covers interactive sign-ins only (Graph's default listing): every legacy-protocol authentication is logged as interactive, so none is missed, but non-interactive token refreshes are not counted; capped=true means the page budget or the deadline (ENTRAADM_DEADLINE, default 45 s) ran out first and the counts are a lower bound -- narrow hours for a full count.

Read-only (AuditLog.Read.All application permission, or -- for azure-cli auth -- the Reports Reader directory role).

Args: hours: How far back to look, clamped to [1, 720] (30 days). max_pages: Page budget (default: ENTRAADM_MAX_PAGES_DEFAULT). min_distinct_users: Distinct accounts an IP needs to appear in shared_ips (default 2, clamped to >= 2).

ParametersJSON Schema
NameRequiredDescriptionDefault
hoursNo
max_pagesNo
min_distinct_usersNo

TDQS

A4.9/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden and does so: read-only, required permission (AuditLog.Read.All / Reports Reader role), client-side aggregation because Graph cannot filter on status/errorCode server-side, interactive-sign-ins-only scope with the reasoning that legacy auth is logged as interactive, and lower-bound semantics when the page budget or ENTRAADM_DEADLINE expires. This is exactly the context annotations would otherwise supply.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Purpose and the failure/success contrast are front-loaded, and the dense middle does real work. It is longer than strictly necessary — the breach-signature narrative and the NAT/VDI caveat could be tightened — but almost every sentence carries operational information rather than padding.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema exists, so the description must explain return values directly, and it does: shared_ips (up to 50, ordering, per-IP account/client-app/country detail, shared_ips_capped) and legacy_auth_users with IP/country counts. Combined with the permission, scope, and cap caveats, an agent has everything needed to call and interpret this tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, so the description must compensate and it does for all three params: hours clamped to [1,720], max_pages defaulting to ENTRAADM_MAX_PAGES_DEFAULT (page budget), and min_distinct_users defaulting to 2 with a clamp of >=2. No parameter is left to inference.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a precise verb+resource+scope (tenant-wide successful sign-in aggregation by source IP) and explicitly contrasts itself with the sibling signin_failure_stats ('shows who is being attacked; this shows whether anyone got in'). An agent can distinguish it from every sibling without opening a schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicit routing to signin_failure_stats, plus genuine interpretation guidance: shared IPs are 'a lead, not a verdict' because NAT/VDI/proxy produce the same signature, and it tells the caller to exclude its own egress ranges before concluding a breach. It also names the condition that invalidates results (capped=true) and the remedy (narrow hours).

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 3 tool updatesv1.1.0
    • Changeddirectory_audits1 field changed
      • addedInput schema / properties / category
        Added value: +{
        +  "anyOf": [
        +    {
        +      "type": "string"
        +    },
        +    {
        +      "type": "null"
        +    }
        +  ],
        +  "default": null,
        +  "title": "Category"
        +}
    • Addedsignin_by_ip
    • Addedsignin_success_stats
  2. 7 tool updatesv0.1.0
    • First observeddaily_brief
    • First observeddirectory_audits
    • First observedget_user
    • First observedget_user_auth_methods
    • First observedhealth_check
    • First observedsignin_failure_stats
    • First observedsignin_logs

TDQS

A4.4/5.0

Scored across 9 tools

Disambiguation4/5

The tools target distinct axes: per-user (get_user, get_user_auth_methods, signin_logs), per-IP (signin_by_ip), tenant-wide aggregates (signin_failure_stats, signin_success_stats), directory audits, and health. Failure vs success stats and by-IP vs by-user sign-ins are clearly delineated. The only mild overlap is daily_brief, which re-bundles signin_failure_stats and directory_audits, though it is explicitly framed as a convenience summary.

Naming Consistency4/5

All names are snake_case and readable, giving a coherent family (signin_*, get_user*, *_stats, *_audits). The only wobble is verb usage: get_ is applied to the user tools but not elsewhere (signin_logs, directory_audits, daily_brief are action-less nouns), so the prefix convention isn't uniform. Still predictable enough to navigate.

Tool Count5/5

Nine tools is well-scoped for an Entra ID triage server, with each tool earning its place across the user/IP/tenant/permission axes. No redundant or filler tools, and nothing is so fine-grained that it fragments a single workflow.

Completeness4/5

For its stated read-only diagnostic purpose the surface is broad: identity, MFA registration, per-user and per-IP sign-ins, tenant-wide success/failure aggregation, directory audits, a daily brief, and a health probe. The obvious omissions (no account disable/revoke-session remediation, no group/role listing) are consistent with the deliberate read-only scope, so gaps are minor rather than blocking.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • -
    license
    Not graded
    quality
    Not graded
    maintenance
    Provides secure access to Microsoft Entra ID (Azure AD) resources including users, devices, and applications through Microsoft Graph API. Enables querying organizational data with comprehensive audit logging to Azure Blob Storage.
    -