Enterprise Microsoft 365 MCP Server
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@Enterprise Microsoft 365 MCP ServerSet an out-of-office auto-reply for jane@contoso.com starting Monday"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
โก Enterprise Microsoft 365 MCP Server
The definitive Model Context Protocol (MCP) server for enterprise Microsoft 365, Exchange Online, Entra ID, and Microsoft Graph administration.
๐ Overview
The Enterprise Microsoft 365 MCP Server connects Large Language Models (LLMs) and AI agents (such as Claude Desktop, Cursor, Antigravity IDE, Cline, and GitHub Copilot) directly to enterprise Microsoft 365 environments.
Equipped with 34 enterprise-grade tools, it bridges Microsoft Graph REST API (for lightning-fast directory and messaging operations) with Exchange Online Management v3 (for deep administrative transport rules, litigation holds, quarantine management, and mail flow forensics).
Related MCP server: OWA Exchange MCP Server
๐ Key Capabilities
โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ
โ ENTERPRISE CAPABILITIES โ
โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโฌโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโฌโโโโโโโโโโโโโโโโค
โ ๐
Vacation & OOF Automation โ ๐ก๏ธ EOP Quarantine & Defender โ ๐ฅ Entra ID โ
โ โข Scheduled HTML auto-repliesโ โข Query quarantined emails โ โข GAL Search โ
โ โข Temporal redirect rules โ โข Tenant Allow/Block Lists โ โข Org Sync โ
โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโผโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโผโโโโโโโโโโโโโโโโค
โ ๐ฌ Mailbox & Delegation Ops โ โ๏ธ Compliance & Governance โ ๐ Analytics โ
โ โข SharedMailbox conversion โ โข Legal & Litigation Holds โ โข License SKUsโ
โ โข FullAccess & SendAs rights โ โข Retention policies & tags โ โข Daily Digestโ
โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโผโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโผโโโโโโโโโโโโโโโโค
โ โก Microsoft Graph Engine โ ๐ Forensics & Mail Flow โ ๐ Transport โ
โ โข High-speed HTML email send โ โข Message delivery traces โ โข Tenant rulesโ
โ โข Teams Adaptive Cards โ โข DKIM status & connectors โ โข Inbox rules โ
โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโดโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโดโโโโโโโโโโโโโโโโ๐๏ธ Architecture
The server employs a Dual-Engine Architecture balancing high throughput with deep administrative capability:
flowchart TD
subgraph AI Client Layer
C1[Claude Desktop]
C2[Cursor IDE]
C3[Antigravity IDE]
C4[Custom Autonomous Agent]
end
subgraph MCP Server Gateway [Stdio Transport]
PROTO[FastMCP JSON-RPC Protocol Dispatcher]
GUARD[Security Boundary & Environment Sanitizer]
DISPATCH[Tool Router - 34 Enterprise Tools]
end
subgraph Dual Engines
subgraph Graph Engine [Graph REST API v1.0]
MSAL[MSAL Client Credentials OAuth2]
G_HTTP[Async HTTP Graph Session]
end
subgraph Exchange Engine [Exchange Management v3]
PS_BRIDGE[PowerShell Isolated Script Bridge]
EXO_CMD[Exchange Online Cmdlets]
end
end
subgraph Microsoft Cloud Fabric
M365_GRAPH[(Microsoft Graph API)]
EXO_SVC[(Exchange Online Fabric)]
ENTRA_DIR[(Microsoft Entra ID)]
end
C1 & C2 & C3 & C4 -->|STDIO JSON-RPC| PROTO
PROTO --> GUARD --> DISPATCH
DISPATCH -->|Direct Graph Calls| MSAL --> G_HTTP --> M365_GRAPH & ENTRA_DIR
DISPATCH -->|Administrative Cmdlets| PS_BRIDGE --> EXO_CMD --> EXO_SVCFor complete details on thread safety, token lifecycle, and error budgets, see the Architecture Blueprint.
๐ ๏ธ Tool Catalog (34 Tools)
Category | Tool Identifier | Engine | Purpose |
Out of Office |
| Exchange | Inspect active auto-reply window and redirect rules |
| Exchange | Schedule HTML auto-reply + temporal team forwarding rule | |
| Exchange | Instantly deactivate OOF and clear vacation rules | |
Offboarding |
| Exchange | Audit disabled accounts to avoid NDR 550 5.1.10 errors |
| Exchange | Convert to free SharedMailbox, set forward with copy & hide | |
Mailbox Ops |
| Exchange | Retrieve quota, protocols, UPN, and forwarding config |
| Exchange | Filter mailboxes by type (Shared, User, Room) | |
| Exchange | Configure FullAccess, SendAs, and sent items retention | |
Rules & Flow |
| Exchange | List client-side and server-side rules on a mailbox |
| Exchange | Inspect tenant-wide mail flow and transport rules | |
| Exchange | Delete a specific inbox rule by name | |
Graph API |
| Graph API | Send corporate HTML emails with zero client latency |
| Graph API | Retrieve recent mailbox messages and read status | |
| HTTP / Webhook | Broadcast rich formatted cards to Microsoft Teams | |
Directory (GAL) |
| Graph API | Search Entra ID users by name, email, or department |
| Exchange | Audit external MailContacts filtered by domain | |
Compliance |
| Exchange | Audit tenant disclaimer and HTML signature rules |
| Exchange | Audit retention policies and MRM tags (GDPR/LGPD) | |
| Exchange | Audit mailboxes with Litigation/Retention Hold active | |
| Exchange | Audit Recoverable Items quota & folder size hygiene | |
Diagnostics |
| Exchange | Trace message delivery events, failures, and routing hops |
| Exchange | Verify cryptographic DKIM signing keys and status | |
| Exchange | List inbound and outbound email connectors | |
Defender & EOP |
| Exchange | Query quarantined messages in Microsoft Defender |
| Exchange | List allowed senders, domains, and URLs | |
| Exchange | Whitelist partner domains or senders with expiration | |
Distribution |
| Exchange | List distribution lists and external delivery policies |
| Exchange | Enumerate distribution group members | |
| Exchange | Add or remove members from distribution groups | |
| Exchange | Control unauthenticated external email permissions | |
Licensing |
| Graph API | Query tenant SKUs, consumed seats, and free licenses |
Profile Sync |
| Graph API | Fetch department, title, phone, manager, and office |
| Graph API | Update organizational attributes in Entra ID | |
Intelligence |
| Graph API | Categorize recent messages into Urgent, Support, Financial |
Detailed parameter schemas and return types are documented in the Tools Reference Manual.
โก Quickstart
Prerequisites
Python 3.10+
Windows PowerShell 5.1 or PowerShell 7+ with the
ExchangeOnlineManagementmodule:Install-Module -Name ExchangeOnlineManagement -Scope CurrentUser -ForceAn Azure Entra ID App Registration (Follow the Entra ID Setup Guide).
1. Clone & Install
git clone https://github.com/ManoAlee/MCP-M365-MICROSOFT.git
cd MCP-M365-MICROSOFT
# Create virtual environment
python -m venv .venv
# On Windows:
.venv\Scripts\activate
# On Linux/macOS:
source .venv/bin/activate
# Install dependencies
pip install -e .2. Configure Environment
Copy the example environment file:
cp .env.example .envEdit .env with your Azure Entra ID credentials:
M365_TENANT_ID=00000000-0000-0000-0000-000000000000
M365_CLIENT_ID=11111111-1111-1111-1111-111111111111
M365_CLIENT_SECRET=your_client_secret_here
M365_PRIMARY_ADMIN=admin@yourtenant.onmicrosoft.com
MCP_LOG_LEVEL=INFO3. Run Self-Tests
python -m unittest discover tests๐ป Client Configuration
Add this server to your preferred MCP client configuration:
๐ท Claude Desktop
Edit %APPDATA%\Claude\claude_desktop_config.json (Windows) or ~/Library/Application Support/Claude/claude_desktop_config.json (macOS):
{
"mcpServers": {
"m365-microsoft": {
"command": "python",
"args": ["-m", "mcp_m365"],
"cwd": "C:\\path\\to\\MCP-M365-MICROSOFT",
"env": {
"M365_TENANT_ID": "your-tenant-id",
"M365_CLIENT_ID": "your-client-id",
"M365_CLIENT_SECRET": "your-client-secret",
"M365_PRIMARY_ADMIN": "admin@yourdomain.com"
}
}
}
}โก Cursor IDE
Add to .cursor/mcp.json or Global Cursor Settings:
{
"mcpServers": {
"m365-microsoft": {
"command": "python",
"args": ["-m", "mcp_m365"],
"cwd": "/path/to/MCP-M365-MICROSOFT"
}
}
}๐ช Antigravity IDE / Cline
Add to ~/.gemini/config/mcp_config.json:
{
"mcpServers": {
"m365-governance": {
"command": "python",
"args": ["-m", "mcp_m365"],
"cwd": "C:\\path\\to\\MCP-M365-MICROSOFT"
}
}
}๐ Security & Compliance
This repository enforces strict enterprise security practices:
Zero Hardcoded Secrets: No tenant IDs, credentials, or client secrets are committed. Everything is dynamically populated through environment variables or local ignored files.
Strict
.gitignore: Excludesconfig.json,.env,*.log,__pycache__, and temporary test tokens.Principle of Least Privilege: Read-only operations are favored where possible, and administrative PowerShell cmdlets require authorized service accounts.
Full Traceability: Every tool invocation logs execution time and outcome for auditability.
๐ค Contributing
Contributions are welcome! Please check out CONTRIBUTING.md and adhere to the project's code style and security guardrails.
Fork the Project
Create your Feature Branch (
git checkout -b feature/AmazingFeature)Commit your Changes (
git commit -m 'feat: Add new M365 compliance tool')Push to the Branch (
git push origin feature/AmazingFeature)Open a Pull Request
๐ License
Distributed under the MIT License. See LICENSE for more information.
Available Tools
34 toolscompliance_audit_domain_signaturesB
Audits tenant-wide email signatures and legal disclaimers configured via Exchange Transport Rules.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the burden. 'Audits' implies a non-destructive read and 'tenant-wide' discloses scope, both useful. However, it never explicitly confirms read-only behavior, permissions needed, or how it differs operationally from a plain transport-rule listing.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single tight sentence with the scope ('tenant-wide') and source ('Exchange Transport Rules') front-loaded. Nothing is padding, though it is almost too terse to carry behavioral weight.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Output schema exists so return values need no explanation, and there are no parameters to cover. The description gives enough to understand what is audited and where the data comes from; only auth/permission context is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Zero parameters, so there is nothing to document and no semantics to misstate. Baseline of 4 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (audits) and resource (tenant-wide email signatures and legal disclaimers), plus the mechanism (Exchange Transport Rules). This meaningfully separates it from siblings like m365_list_transport_rules, though it doesn't explicitly declare that differentiation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance on when to use this versus m365_list_transport_rules or the other compliance_audit_* tools. The agent is left to infer usage from the one-line purpose statement alone.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
compliance_audit_litigation_holdC
Audits mailboxes with Litigation Hold or Retention Hold enabled for legal preservation.
| Name | Required | Description | Default |
|---|---|---|---|
| identity | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden. It implies a read-only audit but never states that it is non-destructive, what permissions or identity scope are required, or how results are limited. For a compliance audit tool with zero annotation coverage this is a meaningful gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One tight sentence with the scope constraints front-loaded and no filler. It is efficient, though its brevity is partly under-specification rather than pure conciseness.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be explained, and the purpose statement is clear. However, with no annotations, no usage guidance, and an undocumented parameter, the definition is only minimally sufficient for an agent to invoke it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The single 'identity' parameter has 0% schema description coverage and is never mentioned in the description, so an agent has no guidance on whose mailboxes are audited or what the null default means. Only the low parameter count keeps this from being a 1.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Audits') and a precise resource ('mailboxes with Litigation Hold or Retention Hold enabled for legal preservation'), which distinguishes it from compliance_audit_retention_policies and compliance_audit_recoverable_items by scope. It does not explicitly name or route away from any sibling, so it sits at 4 rather than 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives no when-to-use context, no prerequisites, and no alternatives among the many compliance_* and m365_* audit siblings. The agent must infer that this is the tool for litigation-hold investigations from the name alone.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
compliance_audit_recoverable_itemsC
Audits mailbox Recoverable Items and Deleted Items folder sizes and quotas for compliance and storage hygiene.
| Name | Required | Description | Default |
|---|---|---|---|
| identity | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden. It implies a read-only audit but never states that it is non-mutating, what permissions or role are required, or what identity formats are accepted. An output schema exists, so return values need not be described, but the operational profile is thin.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single efficient sentence with the resource and scope front-loaded and no filler. It is appropriately sized, though it could trade a few words for the missing parameter and usage detail.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a one-parameter read-only audit tool with an output schema, the description is close to sufficient, but it omits identity format and any permission/system requirements. Given its compliance context and many sibling audit tools, those gaps are meaningful.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% for the single required 'identity' parameter. The description mentions 'mailbox' implicitly tying identity to a mailbox, but does not clarify whether it expects a UPN, SMTP address, GUID, or alias, leaving the one parameter under-specified.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb (Audits) and specific resources (mailbox Recoverable Items and Deleted Items folder sizes and quotas), which clearly separates it from sibling audits like compliance_audit_retention_policies and compliance_audit_litigation_hold. It does not explicitly name or contrast with any sibling, keeping it just below a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The trailing phrase 'for compliance and storage hygiene' gestures at motivation but gives no when-to-use condition, prerequisites, or alternatives among the many sibling audit/compliance tools. Nothing tells an agent when this tool is the right pick versus retention or litigation-hold audits.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
compliance_audit_retention_policiesA
Audits retention policies and MRM tags configured in Exchange Online for compliance/GDPR/LGPD.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are supplied, so the description carries the full behavioral burden. The verb 'Audits' implies a non-destructive read, and an output schema exists to cover return values, but the description never explicitly confirms it is read-only, mentions permission/auth needs, or states its scope (tenant-wide vs per-mailbox).
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with the verb first and zero filler. It is appropriately sized, though one more clause on scope would not have hurt.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a no-parameter audit whose return shape is defined by an output schema, the description covers what is inspected and under which compliance regimes. The only real gaps are the absence of annotations and any statement of scope or safety, which slightly weakens completeness.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so the baseline of 4 applies. There is no parameter syntax the description could or should add.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Audits') plus a precise resource ('retention policies and MRM tags configured in Exchange Online') and a regulatory scope (compliance/GDPR/LGPD). This clearly distinguishes it from siblings like compliance_audit_litigation_hold or compliance_audit_recoverable_items, which audit different artifacts.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description names the compliance context but gives no when-to-use trigger, prerequisites, or comparison to the other compliance_audit_* tools. An agent must infer on its own when this audit is the right call versus litigation hold or recoverable items audits.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
gal_audit_external_mail_contactsB
Audits external MailContacts in Exchange Online directory, optionally filtered by target domain.
| Name | Required | Description | Default |
|---|---|---|---|
| domain_filter | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure but adds almost nothing beyond the tool name. It does not state whether the audit is read-only, what permissions are required, what the audit returns, or any rate limits, leaving significant behavioral gaps.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The definition is a single front-loaded sentence with no redundant or filler content. Every word contributes to the core purpose and filter behavior.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the low complexity (one optional parameter) and the presence of an output schema, the description is minimally adequate for an agent to understand the tool's basic function. However, it omits parameter format details and behavioral context such as read-only nature, which would be needed for confident invocation without annotations.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The description clarifies that the single parameter acts as an optional filter by target domain, which adds meaning beyond the bare schema name. However, with 0% schema description coverage, it does not specify the expected format (e.g., 'example.com' vs '@example.com'), so compensation is partial.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb (audits) and resource (external MailContacts in Exchange Online directory), which is more precise than most sibling tools. However, it does not explicitly distinguish itself from other audit tools such as compliance_audit_* or m365_audit_ex_employees, so it falls short of a clear 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description offers no guidance on when to use this tool versus alternatives, nor does it state any prerequisites or exclusions. It only mentions an optional filter, which is parameter-related, not usage context.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
gal_search_directory_usersC
Searches directory accounts in Microsoft Entra ID by displayName, UPN, or email address.
| Name | Required | Description | Default |
|---|---|---|---|
| top | No | ||
| search_term | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden, yet it discloses nothing about permission requirements, matching semantics (exact vs prefix vs substring), pagination, or result caps. It only implies read-only behavior via 'Searches'.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single tight sentence with the key searchable attributes front-loaded and no filler. It is efficient, though its brevity is partly what leaves the behavioral gaps.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool is simple (one required search term) and an output schema exists, so return values need no explanation. Still, with zero annotations and zero schema coverage, the definition omits matching rules and the unpaginated-ish 'top' behavior that an agent needs to call it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It usefully explains that search_term matches displayName, UPN, or email, but says nothing about the 'top' parameter or its default of 20, leaving half the parameters undocumented anywhere.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Searches) and resource (directory accounts in Microsoft Entra ID) and names the three attributes that can be matched. It is unambiguous, though it does not explicitly distinguish itself from the m365_* user-lookup siblings such as m365_get_user_profile.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance on when to reach for this tool versus the sibling lookup tools, and no stated prerequisites or exclusions. The agent must infer usage purely from the one-line purpose.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
graph_generate_daily_activity_summaryC
Analyzes mailbox messages and generates a categorized summary for daily operational reviews.
| Name | Required | Description | Default |
|---|---|---|---|
| user_upn | Yes | ||
| top_messages | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden. It implies a read/analysis operation but says nothing about which mailbox folders or time windows are scanned, what permissions are needed, whether it can be run repeatedly, or performance/rate implications for large mailboxes.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single tight sentence with the core action front-loaded and no filler. It is efficient, though brevity here edges toward under-specification rather than polish.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be explained, but with zero annotation coverage and zero schema description coverage, an agent lacks scope, permission, and parameter guidance for what is a multi-message analysis operation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, and the description never mentions either parameter. user_upn (required) and top_messages (default 25) are undocumented in both the schema and the description, leaving the agent to infer that user_upn is the target mailbox and that top_messages caps volume.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: "analyzes mailbox messages" and "generates a categorized summary." This clearly distinguishes it from literal-read siblings like graph_list_recent_messages, though it never names an alternative to sharpen the boundary.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase "for daily operational reviews" gives an implied use context, but there is no explicit when-to-use vs when-not, no prerequisites, and no pointer to sibling tools that fetch raw messages instead.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
graph_list_recent_messagesC
Fetches recent emails from a mailbox via Microsoft Graph API.
| Name | Required | Description | Default |
|---|---|---|---|
| top | No | ||
| user_upn | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden, yet it discloses nothing about required permissions, ordering, what 'recent' means quantitatively, or pagination limits. For a Graph API read tool this leaves key operational behavior undocumented.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single compact sentence with the resource and mechanism front-loaded and no wasted words. It is efficient, though brevity here comes at the cost of the missing detail noted elsewhere.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists so return values need not be explained, but with zero annotations, 0% parameter coverage, and no definition of 'recent' or result limits, the definition is too thin for an agent to call it correctly with confidence.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% for both parameters, so the description must compensate and does not. 'user_upn' is only obliquely implied by 'from a mailbox' and the expected UPN format is unstated, while 'top' (default 10) is never mentioned or explained.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (fetches) and resource (recent emails from a mailbox) plus the backing API, which clearly separates it from sibling writers like graph_send_email or config tools. It stops short of naming how it differs from adjacent readers such as m365_message_trace or m365_list_inbox_rules.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no statement of when to use this tool versus alternatives like m365_message_trace, m365_get_mailbox_info, or the other mailbox-related siblings. Usage must be inferred entirely from the name.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
graph_send_emailC
Sends corporate email with HTML formatting via Microsoft Graph API with application permissions.
| Name | Required | Description | Default |
|---|---|---|---|
| subject | Yes | ||
| from_user | Yes | ||
| html_body | Yes | ||
| cc_recipients | No | ||
| to_recipients | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden. It discloses the auth mode ('application permissions') and body format, but says nothing about send side effects (e.g., whether the message lands in Sent Items), Graph throttling limits, permission scope required, or whether from_user must be a real mailbox/UPN. For a send/mutation tool this is a substantial gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with no padding. It could be tighter ('corporate' and the Graph API branding add little), but it is efficiently structured.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be explained, but with 5 params at 0% schema coverage and no annotations, the description is too thin for a tool that performs an irreversible external send. It should at minimum cover recipient/from formats and permission prerequisites.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% across 5 parameters, and the description does not compensate. It hints that html_body expects HTML but gives no format or example for from_user (UPN vs address), to_recipients/cc_recipients list semantics, or subject constraints, leaving four required inputs undocumented.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Sends corporate email') plus the transport and auth model ('via Microsoft Graph API with application permissions'). This clearly distinguishes it from siblings like graph_send_teams_notification and graph_list_recent_messages.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no when-to-use guidance and no alternatives named. An agent cannot tell from the description whether to prefer this over graph_send_teams_notification for a given notification task, or what prerequisites (e.g., a provisioned service principal) are needed.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
graph_send_teams_notificationC
Sends rich formatted card notifications to Microsoft Teams channels via Incoming Webhooks.
| Name | Required | Description | Default |
|---|---|---|---|
| title | Yes | ||
| message | Yes | ||
| theme_color | No | 0078D7 | |
| webhook_url | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden but discloses little beyond the delivery mechanism (Incoming Webhooks). It does not state that this is a write/send action, permission or webhook-setup requirements, rate limits, or failure behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with no filler or repetition. It is efficient, though the brevity contributes to underspecification rather than being pure virtue.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a four-parameter write/send tool with no annotations and 0% parameter coverage, the description is too thin. The output schema covers return values, but the agent still lacks setup, webhook default, and configuration context needed to call it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate for all four parameters, yet it adds almost no parameter meaning. It hints that the card carries title/message/theme, but never explains the webhook_url default or the expected message/title format.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description gives a specific verb (Sends) plus resource (rich formatted card notifications) and target (Microsoft Teams channels via Incoming Webhooks), which separates it from the sibling graph_send_email. It is clear and concrete, though it never names an alternative tool explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no when-to-use, when-not-to-use, or alternative guidance. The agent learns what the tool does but nothing about choosing it over graph_send_email or any sibling.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
m365_add_sender_to_allow_listC
Adds a sender email or domain to the Tenant Allow List with an expiration period.
| Name | Required | Description | Default |
|---|---|---|---|
| note | No | Approved via M365 MCP | |
| sender_or_domain | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full behavioral burden and largely fails: it says nothing about required permissions, whether the entry overwrites an existing allow-list rule, how the expiration is set or renewed, or what happens if the sender is already listed. The only extra signal is the generic 'with an expiration period' phrase.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single, front-loaded sentence with no filler. It earns its brevity, though the brevity is partly the source of the missing behavioral detail rather than a virtue of precision.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists so return values need not be explained, but for a security-sensitive mutation with zero annotations and zero schema descriptions, the definition should cover permissions, idempotency, and the unexplained expiration behavior. The dangling 'expiration period' reference is left unresolved.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must compensate and does not. It loosely matches 'sender_or_domain' but gives no format guidance (bare address vs domain vs wildcard), and it never mentions the 'note' parameter. Worse, it references an 'expiration period' that does not exist as a parameter, creating ambiguity about how that value is supplied.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Adds a sender email or domain to the Tenant Allow List'), and the expiration qualifier narrows the behavior. It is clearly distinguishable from the sibling m365_list_tenant_allow_list, though it never names that sibling or other alternates.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no when-to-use or when-not-to-use guidance, no mention of prerequisites (e.g. mail-flow or spam-filtering context), and no pointer to the related list tool or removal flow. An agent must infer the usage context entirely from the name.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
m365_audit_ex_employeesA
Audits disabled or terminated accounts for NDR 550 5.1.10 risks and license optimization opportunities.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden. 'Audits' strongly implies a read-only operation, but that is never stated, nor are permission requirements, scope, or whether any remediation is performed. It does disclose the specific audit criteria (NDR 550 5.1.10, license optimization), which is useful context.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single, front-loaded sentence that names the action, target, and two audit dimensions with no filler. It is dense in domain jargon but every clause earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be described, and there are no parameters to document. However, for an audit tool with no annotations, the description leaves out whether it is read-only and what the agent should do next, leaving it only minimally complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so there is nothing for the description to clarify; the baseline of 4 applies. No param gaps exist to compensate for.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb (audits) and a specific resource (disabled or terminated accounts) plus the two concrete concerns it checks: NDR 550 5.1.10 risks and license optimization. This distinguishes it from auditing siblings like gal_audit_external_mail_contacts, though it does not name which sibling to pick instead.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Use is implied by the 'ex_employees' name and the audit framing, but there is no explicit statement of when to run this versus m365_convert_ex_employee_to_shared, m365_get_mailbox_info, or m365_list_license_skus. No prerequisites or triggers are given.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
m365_check_quarantineC
Checks quarantined messages in Microsoft Defender / EOP by recipient, sender, or hours back.
| Name | Required | Description | Default |
|---|---|---|---|
| page_size | No | ||
| hours_back | No | ||
| sender_address | No | ||
| recipient_address | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden. 'Checks' implies a read-only operation, but it omits any mention of permissions, throttling, whether the check is exhaustive or paginated, and what a null/default filter does. For a tool with zero annotation coverage this is a significant gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with no wasted words. It is appropriately terse, though the same brevity leaves required parameter context uncovered.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given 0% schema description coverage, no annotations, and four parameters, the definition is under-specified: page_size is unmentioned, filter formats and defaults are absent, and behavioral context (auth, read-only nature, result scope) is missing. An output schema exists, so return-value explanation is not required, but the input-side and usage gaps remain substantial.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate for all four parameters. It names recipient, sender, and hours back, but omits page_size entirely and adds no syntax, format, or default value details (e.g., hours_back defaults to 48). Partial compensation is not enough for a 3.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a clear verb ('Checks') and resource ('quarantined messages') scoped to Microsoft Defender / EOP, and names the filter dimensions (recipient, sender, hours back). It does not explicitly distinguish itself from the similarly themed sibling m365_message_trace, so it falls short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description hints at how it can be filtered but gives no guidance on when to use this tool versus alternatives like m365_message_trace. There are no exclusions, prerequisites, or context about the scenario that selects this tool.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
m365_configure_vacationC
Configures scheduled out-of-office auto-reply with HTML formatting and optional forwarding/redirection.
| Name | Required | Description | Default |
|---|---|---|---|
| end_date | Yes | ||
| user_upn | Yes | ||
| start_date | Yes | ||
| redirect_to | No | ||
| return_date | Yes | ||
| signature_name | Yes | ||
| inform_team_handling | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries full behavioral burden. It hints that HTML formatting and forwarding/redirection are supported, but says nothing about whether existing settings are overwritten, what permissions are required, or reversibility โ significant gaps for a mutation tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single efficient sentence with the core purpose front-loaded and no filler. Its brevity is appropriate structurally even though the content is thin elsewhere.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 7-parameter mutation tool with 5 required fields, no annotations, and 0% schema description coverage, the description is far too sparse. An output schema exists so return values need not be explained, but the input semantics and behavioral effects are left largely uncovered.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the schema explains none of the 7 parameters. The description alludes to forwarding/redirection (mapping loosely to redirect_to) but leaves user_upn, start_date, end_date, return_date, signature_name, and inform_team_handling entirely unexplained.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The verb 'Configures' plus resource 'out-of-office auto-reply' is specific and identifies the operation clearly. It does not explicitly distinguish itself from close siblings like m365_disable_vacation or m365_get_vacation_status, but the intent is unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no when-to-use guidance, no mention of prerequisites (permissions, mailbox targeting), and no reference to the sibling status/disable tools. The agent must infer usage from the name alone.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
m365_disable_vacationC
Immediately disables automatic out-of-office replies and clears temporary vacation redirect rules.
| Name | Required | Description | Default |
|---|---|---|---|
| user_upn | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden. It usefully discloses that redirect rules are cleared (a destructive side effect) and that the disable is immediate, but it omits permission requirements, reversibility, and whether existing configured out-of-office messages are affected.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence that states the primary effect and the cleanup action with no wasted words. It is appropriately sized for the tool's scope.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be explained, but for a mutation tool with no annotations and a 0%-documented parameter, the description should cover when to use it, permissions, and the user_upn parameter. These key context pieces are missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and the only parameter, user_upn, is not mentioned in the description. The description does not compensate by explaining the expected format, source, or role of user_upn.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('disables') and resource ('automatic out-of-office replies') plus a secondary effect ('clears temporary vacation redirect rules'), distinguishing it from read sibling m365_get_vacation_status. It does not explicitly name the alternative m365_configure_vacation, so sibling differentiation is implied rather than stated.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives no explicit when-to-use, when-not-to-use, or alternative routing. 'Immediately' hints at urgency but does not say which scenarios call for this tool over m365_configure_vacation or m365_get_vacation_status.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
m365_get_group_membersC
Retrieves the member list of a distribution group.
| Name | Required | Description | Default |
|---|---|---|---|
| group_identity | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden. Beyond 'retrieves', it says nothing about permissions required, behavior for large groups (pagination/truncation), transitive vs direct members, or error cases for an unknown group โ all relevant for a directory read on groups that can have thousands of members.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single short sentence with no filler and the object front-loaded. It is efficient, though the brevity comes at the cost of the missing detail noted elsewhere.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be described. For a one-parameter read tool this is close to adequate, but the identity format and group-size behavior remain undocumented, leaving a real gap an agent could stumble on.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, and the single parameter group_identity has no title-level explanation of accepted value forms (GUID, SMTP address, display name). The description narrows the target to a distribution group, which is marginal added meaning but does not compensate for the format ambiguity.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Retrieves) and resource (member list of a distribution group), so an agent can tell it apart from m365_list_distribution_groups and m365_manage_group_member by object and scope. It stops short of explicitly contrasting with those siblings, so it is clear but not fully differentiating.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No indication of when to use this tool versus m365_list_distribution_groups or m365_manage_group_member, and no prerequisites or context are given. The implied usage (read a group's membership) is inferable from the verb, but nothing is stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
m365_get_mailbox_infoC
Retrieves comprehensive mailbox properties (type, protocols, UPN, aliases, forwarding, quotas).
| Name | Required | Description | Default |
|---|---|---|---|
| identity | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are supplied, so the description carries the full behavioral burden. 'Retrieves' implies a read operation, but nothing states whether elevated permissions are needed, whether the call is rate-limited, or that 'comprehensive' means the full property set is always returned. With zero annotation coverage this is a significant gap for a mailbox-admin tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One front-loaded sentence with no filler; the parenthetical field list is compact and informative. It is appropriately sized, though slightly list-heavy rather than explanatory.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be described, and the tool is a simple one-parameter read. However, with no annotations and an undocumented identity parameter, the definition leaves the two things an agent most needs โ how to identify the mailbox and what authority the call requires โ unstated.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The single required parameter 'identity' has 0% schema description coverage and the description never explains it. An agent cannot tell from either source whether identity accepts a UPN, SMTP address, alias, display name, or GUID โ a real ambiguity for an Exchange-style identity parameter.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource (retrieves mailbox properties) and enumerates the property groups returned (type, protocols, UPN, aliases, forwarding, quotas), which cleanly separates it from m365_list_mailboxes (enumeration) and m365_get_user_profile (user, not mailbox). It stops short of naming those siblings explicitly, so it is clear but not fully differentiated.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no when-to-use statement, no prerequisites (e.g., required admin roles), and no routing to alternatives such as m365_list_mailboxes for discovery. Usage is only implied by the verb 'Retrieves'.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
m365_get_user_profileB
Retrieves a user's corporate profile (department, jobTitle, phone numbers, office, account status).
| Name | Required | Description | Default |
|---|---|---|---|
| user_upn | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden. 'Retrieves' implies a read, and the field list hints at the payload, but there is no disclosure of permission requirements, behavior for a missing/disabled user, or whether the lookup is scoped to the tenant directory.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single well-formed sentence with the purpose front-loaded and no filler. Every listed field earns its place by scoping the returned payload.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be re-explained, and the description appropriately names the key returned fields. For a one-parameter read tool the only real gap is the undocumented user_upn format and missing-user behavior.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% โ the single required parameter user_upn has no schema description at all. The description never mentions the identifier, its UPN format, or whether it accepts a GUID/email, so it does not compensate for the coverage gap; only the self-evident parameter name keeps this from scoring lower.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Retrieves) and resource (a user's corporate profile) and even enumerates the fields returned (department, jobTitle, phone numbers, office, account status). It is clearly distinguishable from the mutation sibling m365_update_user_profile, but does not name a sibling explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No when-to-use guidance, no prerequisites, and no mention of alternatives such as m365_get_mailbox_info or gal_search_directory_users. The description only says what the tool does, leaving the agent to infer when to prefer it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
m365_get_vacation_statusB
Checks the scheduled automatic reply (OOF / Out-of-Office) and vacation forwarding rules for a mailbox.
| Name | Required | Description | Default |
|---|---|---|---|
| user_upn | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden. It discloses what is checked (OOF and forwarding rules) and implies a read-only operation, but does not state permissions, safety, or that it has no side effects. The disclosure is minimally adequate for a read tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single, front-loaded sentence with no filler. It efficiently communicates the tool's scope and checked items.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The existence of an output schema means return values are covered, and the description adequately names the checked settings. However, it omits usage guidance and parameter context, leaving notable gaps for agent invocation despite the tool's simplicity.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate for the undocumented user_upn parameter. It says 'for a mailbox' but never names or explains the required UPN argument, adding no semantic detail beyond the schema's bare type.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Checks') and two distinct resources: OOF/automatic reply and vacation forwarding rules, scoped to a mailbox. The read-only nature implicitly distinguishes it from siblings like m365_configure_vacation and m365_disable_vacation, though it does not explicitly name them.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description offers no when-to-use guidance, no conditions, and no alternatives. It does not tell the agent to use this tool for inspecting vacation settings vs. m365_configure_vacation for changing them.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
m365_list_connectorsA
Lists inbound and outbound email connectors configured in Exchange Online.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description must carry the behavioral burden. 'Lists' implies a read-only operation and the description states the scope (inbound/outbound, Exchange Online), but it omits permissions, rate limits, and pagination behavior. For a simple read-only list tool this is adequate but not rich.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, front-loaded sentence with no wasted words. It states the action and scope efficiently.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's simplicity (zero parameters, output schema present), the description covers what is listed and where. It omits guidance on when to choose this tool over sibling listing tools, but is otherwise sufficient for correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters, so the baseline score is 4. The description does not need to add parameter semantics beyond the empty schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description gives a specific verb ('Lists') and resource ('inbound and outbound email connectors configured in Exchange Online'), making the tool's purpose clear. It does not explicitly distinguish itself from sibling tools such as m365_list_transport_rules or m365_list_inbox_rules, so it falls short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use this tool versus alternatives. It neither mentions prerequisites nor names related tools like transport rule or inbox rule listings, leaving selection entirely to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
m365_list_distribution_groupsC
Lists distribution groups in Exchange Online with authentication and external delivery settings.
| Name | Required | Description | Default |
|---|---|---|---|
| result_size | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden. It discloses that results include authentication and external delivery settings, but says nothing about required permissions, pagination, rate limits, or what result_size does. For a zero-annotation tool this is thin.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single efficient sentence with the key scoping information front-loaded. No wasted words, though it is arguably too terse given the gaps elsewhere.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need no explanation, and the tool is a simple read-only list. However, with no annotations and an undocumented parameter, the definition leaves the agent guessing about permissions, result_size semantics, and when this beats sibling listing tools.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and the sole parameter (result_size, default 50) is never mentioned in the description. The description does nothing to compensate for the undocumented parameter or clarify its default/paging behavior.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Lists), resource (distribution groups), scope (Exchange Online), and even the attributes surfaced (authentication and external delivery settings). It is distinguishable from generic mailbox/group siblings, though it never names an alternative.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives no when-to-use guidance and no exclusions or alternatives. An agent must infer that this is the directory-style listing tool rather than something like gal_search_directory_users or m365_get_group_members.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
m365_list_inbox_rulesB
Lists all client-side and server-side inbox rules configured on a specific mailbox.
| Name | Required | Description | Default |
|---|---|---|---|
| user_upn | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full disclosure burden. It does add useful scope context by stating that both client-side and server-side rules are returned, implying a read-only inspection, but it omits permission requirements, result limits, and pagination behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with no filler; the resource and scope are stated immediately.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists so return values need not be described, and the tool is a simple read. However, for a tool sharing a crowded rule/audit namespace, the description never routes the agent away from siblings like m365_list_transport_rules or explains prerequisites.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% for a single required parameter, and the description only implies the mailbox selector via 'a specific mailbox.' Since only one, self-evidently named parameter (user_upn) exists, the gap is minor but the UPN format is never clarified.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Lists) and resource (client-side and server-side inbox rules) scoped to a specific mailbox. It does not distinguish itself from the sibling m365_list_transport_rules, which the 'rules' wording could easily be confused with.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no when-to-use guidance, no mention of when a transport-rule listing is preferable, and no prerequisites or permissions noted. The agent must infer context entirely from the name.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
m365_list_license_skusA
Lists all M365 license subscription SKUs showing enabled, consumed, and available seats.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It usefully discloses the returned fields (enabled, consumed, available seats), which hints at a read-only inventory operation, but never states that it is read-only, what permissions are required, or whether results are paginated.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One sentence, front-loaded with the verb and resource, and the returned metrics follow immediately. No filler or redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return format need not be restated, and with zero parameters the schema is fully self-describing. The only small gap is the absence of any read-only/permission framing for a no-annotation tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so the baseline is 4. There is nothing for the description to explain beyond what the empty schema already conveys.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Lists), a precise resource (M365 license subscription SKUs), and the payload contents (enabled, consumed, available seats). No sibling in the toolset covers license inventory, so the agent can distinguish it immediately from mailbox, group, or compliance tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The word 'Lists all' implies this is the inventory/audit entry point for license data, but there is no explicit when-to-use statement, no prerequisites, and no named alternative. Usage is inferred rather than guided.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
m365_list_mailboxesC
Lists mailboxes filtered by type (SharedMailbox, UserMailbox, RoomMailbox).
| Name | Required | Description | Default |
|---|---|---|---|
| result_size | No | ||
| recipient_type | No | SharedMailbox |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden. It discloses the filter dimension but says nothing about required permissions, pagination/result_size behavior, default scope, or ordering, which an agent needs for a listing tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with no filler. It is efficient, though its brevity is partly under-specification rather than pure tightness.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be described. For a 2-parameter read tool this is nearly adequate, but the unexplained result_size default and absent pagination/permission context leave a real gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It usefully enumerates the accepted recipient_type values (SharedMailbox, UserMailbox, RoomMailbox), which the schema itself does not. However, result_size and its default of 50 are left completely unexplained, so the compensation is only partial.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a clear verb ("Lists") and resource ("mailboxes") plus the filter dimension (type). It is unambiguous against read-oriented siblings like m365_get_mailbox_info or mutation tools like m365_convert_ex_employee_to_shared, but it never explicitly names or contrasts those alternatives.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to use this tool versus m365_get_mailbox_info or m365_list_distribution_groups, no prerequisites, and no exclusions. Usage is only implied by the word "Lists."
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
m365_list_tenant_allow_listB
Lists entries in the Microsoft Defender Tenant Allow/Block List (senders, domains, URLs).
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden. 'Lists' implies read-only, but the description does not disclose auth requirements, rate limits, pagination, or whether both allow and block entries are returned; it only names the entry categories.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with no filler. It names the verb, resource, and entry types efficiently, with every phrase earning its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a no-parameter list tool with an output schema, the description identifies the resource and entry categories sufficiently. A brief read-only or auth note would have improved safety clarity in the absence of annotations, but it is otherwise adequate.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters, so the baseline is 4. The description appropriately adds no parameter semantics because none exist, and the empty schema leaves nothing further to explain.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Lists') and exact resource ('Microsoft Defender Tenant Allow/Block List'), with a parenthetical naming entry types. It implicitly distinguishes itself from m365_add_sender_to_allow_list by verb, but does not explicitly name alternatives or siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives no when-to-use, when-not-to-use, prerequisites, or alternative routing. It implies inspection of the allow/block list, but leaves all usage context to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
m365_list_transport_rulesC
Lists tenant-wide mail flow and transport rules in Exchange Online.
| Name | Required | Description | Default |
|---|---|---|---|
| filter_state | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden, and it only implies a read-only listing via the verb 'Lists'. It says nothing about permissions required, pagination or result volume, or whether the filter is applied server-side, all of which matter for a tenant-wide directory listing.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with no filler or repetition; the scope qualifier sits before the resource. It is efficient, though it is arguably too terse to carry any of the usage or parameter context the tool needs.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be explained, but the description leaves the one parameter undefined, offers no usage guidance, and provides no behavioral context in the absence of annotations. For a compliance-relevant tenant-wide listing tool, this is thin.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and the single parameter filter_state is undocumented in both schema and description. The description gives no hint about what states can be filtered on or the accepted value format, so it fails to compensate for the coverage gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific verb ('Lists'), a specific resource ('transport rules') and a scope ('tenant-wide mail flow ... in Exchange Online'), which is enough to separate it from the mailbox-scoped m365_list_inbox_rules sibling without opening either schema. It does not, however, explicitly name that sibling, so the differentiation is left implicit.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'tenant-wide' hints at the scope, but there is no explicit when-to-use statement, no mention of when this is preferable to m365_list_inbox_rules, and no prerequisites or auth context. An agent has to infer applicability purely from the resource name.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
m365_manage_group_memberC
Adds or removes a user from a distribution group (action: 'add' or 'remove').
| Name | Required | Description | Default |
|---|---|---|---|
| action | No | add | |
| member_email | Yes | ||
| group_identity | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It implies a mutation but says nothing about required permissions (owner rights), whether removal is reversible, whether external/non-member addresses are allowed, or what happens on failure. Only the add/remove duality is disclosed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with the operation and allowed action values, no filler. It is efficient, though it could carry one more clause of behavioral context at the same cost.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The presence of an output schema means return values need not be explained. For a simple two-mode mutation the core intent is covered, but the absence of any permission, error, or eligibility context leaves an agent under-informed about safe invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0% and the schema declares no enum, so the description's mention of action values 'add'/'remove' adds genuine meaning the schema lacks. However, group_identity (name vs. address vs. ID) and member_email formats are left entirely undefined, so the description only partially compensates for the coverage gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (adds/removes) and resource (user from a distribution group), which clearly separates it from read-oriented siblings like m365_get_group_members and m365_list_distribution_groups. It does not explicitly name those siblings, but the action framing is unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description lists the two action values but gives no when-to-use guidance, prerequisites, or alternatives. There is no indication of when a user should choose this over m365_set_group_external_delivery or other group-management siblings.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
m365_message_traceC
Traces message delivery events in Exchange Online by sender, recipient, status, and time window.
| Name | Required | Description | Default |
|---|---|---|---|
| hours_back | No | ||
| result_size | No | ||
| status_filter | No | ||
| sender_address | No | ||
| recipient_address | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full disclosure burden, yet it only states what is filtered. It never says the operation is read-only, whether it requires admin/exchange permissions, latency of trace results, or that all parameters are optional with defaults.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single tight sentence that front-loads the verb and resource with no filler. It is efficient, though it could have spent a few more words on parameter/value detail without becoming bloated.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With an output schema present, return values need not be described, and the filter dimensions are covered. However, for a five-parameter tool with no annotations and 0% schema coverage, the definition omits result_size semantics and the read-only/permission profile, leaving real gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate; it names four of the five inputs (sender, recipient, status, and the time window implied by hours_back), which is genuinely useful. The fifth parameter, result_size, is not mentioned, and no format hints (e.g. address syntax, allowed status values) are given.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description gives a specific verb ('traces') plus resource ('message delivery events in Exchange Online') and enumerates the filter dimensions. No sibling tool overlaps with message tracing, so the agent can place it easily, though it never names a sibling to distinguish against.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no statement of when to use this versus alternatives such as graph_list_recent_messages or m365_check_quarantine, nor any prerequisite or when-not-to-use guidance. Usage must be inferred entirely from the purpose sentence.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
m365_remove_inbox_ruleC
Removes a specific inbox rule from a mailbox by its RuleIdentity or Name.
| Name | Required | Description | Default |
|---|---|---|---|
| user_upn | Yes | ||
| rule_name | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description must carry the full behavioral burden. It says only that the rule is removed; it omits that this is an irreversible mutation, what permissions are required, and any error behavior. A destructive operation with zero annotation coverage needs more disclosure.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single tight sentence with the action and key front-loaded. No wasted words, though brevity here contributes to the transparency gaps.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists so return values need not be described, but for a destructive mailbox mutation with no annotations and completely undocumented parameters, the description leaves the agent short of what it needs to invoke this safely and correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must compensate. It mentions identifying the rule by 'RuleIdentity or Name' (which loosely maps to rule_name and hints at user_upn's mailbox scope) but does not explain the expected UPN format or how the rule name is obtained, and the 'RuleIdentity' wording does not match the schema's single rule_name parameter.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description gives a specific verb (removes) and resource (inbox rule) and states the identifying key. It implicitly differentiates from m365_list_inbox_rules and m365_list_transport_rules by naming the object, but never explicitly contrasts with them.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance on when to use this tool versus alternatives, no prerequisites (e.g., listing rules first to obtain the rule name), and no note about what happens when the rule does not exist.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
m365_set_group_external_deliveryB
Configures whether a distribution group accepts messages from unauthenticated external senders.
| Name | Required | Description | Default |
|---|---|---|---|
| allow_external | No | ||
| group_identity | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden. It does clarify the security-relevant meaning of the toggle (accepting messages from unauthenticated external senders), which is genuine context beyond the raw schema, but it omits permission requirements, reversibility, and the effect of applying the change.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with no filler; the setting and its scope appear immediately. Nothing is redundant.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be explained, and the tool is small (2 params). However, for a mutation of a security-sensitive setting with no annotations, the description should at least mention the default and identity format; as written it is minimally adequate.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It implicitly explains what allow_external controls, but never names either parameter, does not state the default (allow_external defaults to true), and gives no format guidance for group_identity (name vs SMTP vs GUID).
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific verb (Configures) and a precise resource/setting (whether a distribution group accepts messages from unauthenticated external senders), which clearly separates it from sibling group tools like m365_manage_group_member and m365_list_distribution_groups. It stops short of explicitly naming an alternative or contrasting tool, so it is clear but not fully differentiated.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no when-to-use guidance, no stated prerequisites (e.g., required admin role for external-delivery changes), and no reference to related tools such as m365_list_distribution_groups or transport-rule tools. The agent must infer the context entirely.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
m365_update_user_profileC
Updates user directory attributes (department, job title, mobile phone, business phone).
| Name | Required | Description | Default |
|---|---|---|---|
| user_upn | Yes | ||
| job_title | No | ||
| department | No | ||
| mobile_phone | No | ||
| business_phone | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden of behavioral disclosure. It implies a mutation but does not state permission requirements, whether changes are reversible, or how the tool treats omitted parameters (e.g., whether they are left unchanged or cleared). For a directory-mutating operation, this is a significant gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single efficient sentence that front-loads the core action and enumerates the affected fields without waste.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The presence of an output schema means return values need not be described, but the overall definition is incomplete for a 5-parameter mutation tool with no annotations. Missing details include the required user identifier, update semantics, and any behavioral context beyond the field list.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It lists four updateable attributes (department, job title, mobile phone, business phone) but completely omits the required user_upn parameter, leaving the agent without guidance on how to identify the target user. It also does not explain the meaning of null defaults or the partial-update behavior.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a clear verb ('Updates') and resource ('user directory attributes') and enumerates the specific fields affected. It does not explicitly differentiate from the read-only sibling m365_get_user_profile, but the verb makes the intent unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is given on when to use this tool versus alternatives, nor any prerequisites or exclusions. The description provides only what the tool does, not when to invoke it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
m365_verify_dkim_statusB
Verifies DKIM signing status and active key configuration for tenant accepted domains.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden of behavioral disclosure. It does not state that the operation is read-only, what permissions are required, or any other behavioral trait beyond the basic purpose. The verb 'verifies' implies a read, but this is not made explicit.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, front-loaded sentence that contains no filler. It is appropriately sized for a simple verification tool and communicates the essential purpose immediately.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool is simple, has no parameters, and an output schema exists, so the description need not explain return values. It adequately states what is verified, but the absence of any annotation or behavioral note leaves a minor gap around read-only safety and permissions.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, and the schema is empty with full description coverage. Per the scoring rule, a zero-parameter tool has a baseline of 4, and there is no additional parameter semantics for the description to add.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb (verifies) and resource (DKIM signing status and active key configuration for tenant accepted domains), so an agent can understand the core action. However, it does not explicitly differentiate this tool from the potentially overlapping sibling compliance_audit_domain_signatures, which limits it to a 4 rather than a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use this tool versus alternatives, nor does it state any prerequisites or conditions. It simply declares what the tool does, leaving the agent to infer usage context from the name alone.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
34 tool updates
v1.0.0- First observed
compliance_audit_domain_signatures - First observed
compliance_audit_litigation_hold - First observed
compliance_audit_recoverable_items - First observed
compliance_audit_retention_policies - First observed
gal_audit_external_mail_contacts - First observed
gal_search_directory_users - First observed
graph_generate_daily_activity_summary - First observed
graph_list_recent_messages - First observed
graph_send_email - First observed
graph_send_teams_notification - First observed
m365_add_sender_to_allow_list - First observed
m365_audit_ex_employees - First observed
m365_check_quarantine - First observed
m365_configure_shared_mailbox - First observed
m365_configure_vacation - First observed
m365_convert_ex_employee_to_shared - First observed
m365_disable_vacation - First observed
m365_get_group_members - First observed
m365_get_mailbox_info - First observed
m365_get_user_profile - First observed
m365_get_vacation_status - First observed
m365_list_connectors - First observed
m365_list_distribution_groups - First observed
m365_list_inbox_rules - First observed
m365_list_license_skus - First observed
m365_list_mailboxes - First observed
m365_list_tenant_allow_list - First observed
m365_list_transport_rules - First observed
m365_manage_group_member - First observed
m365_message_trace - First observed
m365_remove_inbox_rule - First observed
m365_set_group_external_delivery - First observed
m365_update_user_profile - First observed
m365_verify_dkim_status
TDQS
Scored across 34 tools
Most tools target distinct resources and actions, such as vacation status vs. configure vs. disable, or message trace vs. quarantine checks. However, several audit and directory tools have adjacent purposes (e.g., compliance audits, GAL search vs. user profile lookup, transport rules vs. domain signature audit), so occasional confusion is possible.
Tools generally use snake_case verb_noun phrasing, which is readable and mostly predictable. But the server mixes multiple prefix families (m365_, graph_, gal_, compliance_) and some names are long audit-style descriptions rather than consistent action-oriented patterns.
At 34 tools, the surface is heavy for an MCP server and risks overwhelming tool selection. While Microsoft 365 administration is broad, many tools are narrow audit or configuration variants that could be consolidated or grouped.
The set covers a wide range of Exchange Online, Entra ID, Defender, compliance, and Graph operations, including many read and audit capabilities. However, it lacks several core lifecycle operations such as mailbox/user/group creation and deletion, license assignment, and broader Teams or SharePoint management.
Maintenance
Related MCP Connectors
Manage Microsoft 365 email, calendar, contacts and inbox rules via the Graph API with OAuth 2.0.
Gmail, Outlook, Drive, OneDrive and calendars for AI agents. Many accounts, one endpoint, audit log.
Protocol-native energy infrastructure orchestration for AI data centers. Provides 46 MCP tools across 8 grid protocols (IEC-61850, DNP3, Modbus, OCPP, OpenADR, IEEE 2030.5, IEC 60870-5-104, ICCP) with 5 core API primitives: connect, dispatch, settle, comply, and intel. Enables AI agents to programmatically interact with substations, grid interfaces, and energy assets for real-time workload-grid coordination.
Permissioned access to Outlook, OneDrive and Teams via the user's own Microsoft account
Related MCP Servers
- FlicenseNot gradedqualityDmaintenanceConnects AI assistants to Microsoft 365 accounts to manage emails, calendars, files, and Teams messages. It offers 71 tools and supports multi-user environments through a secure, customizable server architecture.55-
- AlicenseAqualityDmaintenanceMCP server for any Microsoft Exchange / OWA deployment. Gives LLM agents access to email, calendar, directory search, folders, availability, and meeting analytics via 30 tools.3026 PyPI8MIT
- FlicenseBqualityDmaintenanceEnables LLMs to manage your Microsoft 365 calendar, tasks, and email via Microsoft Graph API, acting as a personal secretary.14-
- FlicenseNot gradedqualityDmaintenanceConnects AI assistants to Microsoft 365 via the Graph API, enabling email search, attachment extraction, and OneDrive file reading through natural conversation.-