PrivacyFence
OfficialSearch and read Confluence pages and attachments; create and update pages.
Search and read Gmail messages, threads, and attachments; create drafts and replies, manage labels and filters, and archive messages. No tool sends email.
Read and write Google Apps Script project source; read the result of a run you started, without running scripts itself.
Read Google Calendar events, free/busy status, and rooms; create, update, and delete events; manage out-of-office and working location.
Search, read, download, upload, move, and write Google Drive files; edit and format Google Docs; read, write, and format Google Sheets.
Read, create, update, comment on, and transition Jira issues.
Read Salesforce records, search, and run reports in a read-only manner.
List and read Slack channels, DMs, and threads; search; send messages; start group chats.
Read and search Telegram chats; send messages.
PrivacyFence
AI access without giving AI the keys. Approve the sensitive. Automate the routine.
PrivacyFence is an open-source privacy and approval gateway between AI assistants and your business systems. It connects MCP-compatible assistants such as Claude Desktop and Claude Code to Gmail, Google Drive, Calendar, Slack, Salesforce, Jira, Confluence, Telegram, and more.
PrivacyFence enforces, independently of the AI, what an assistant may see and do. Sensitive reads and consequential actions require human approval, while routine requests can be automated by policy. Optional PII detection runs locally before personal data reaches the AI, and every decision is audited.
PrivacyFence runs on an employee’s own computer (macOS, Windows, or Linux) or as a central deployment on infrastructure the organization controls, allowing web clients such as claude.ai to connect as well. Connector credentials stay with PrivacyFence, never with the AI client, and no data passes through PrivacyFence-operated servers — there are none.
Website: privacyfence.eu · Download: privacyfence.eu/download · Docs: privacyfence.eu/docs
Why
Giving an AI assistant access to Gmail, Drive, Slack or Salesforce usually means giving it a standing permission and trusting it to use that permission well. Being allowed to read a record is not the same as wanting it sent to an AI, and a tool name in a client's prompt says little about what is about to change. PrivacyFence puts an independent control point between the assistant and your systems: the AI asks, PrivacyFence decides, and you see what is at stake before it happens.
Related MCP server: mcpgate
What it does
Human approval for sensitive reads and for writes, on cards that show the actual content or change, who is asking and why — not a raw tool name or JSON payload.
Local PII detection before a read reaches the AI: likely personal data is highlighted, and it sends the request to a card even when a rule would have let it through.
Policy-based automation: narrow always-allow rules (a sender domain, a Drive folder, a Slack channel, a Jira project) let routine requests run without a card.
An audit log of every accepted, denied and automatically approved request, chained so that an edit made without its key is detected.
Credentials stay with PrivacyFence. On a packaged install it runs under its own service account, so the AI client, which runs as you, cannot read them or approve its own request.
A defined set of connectors, each tool with a gate fixed in code. PrivacyFence is not a generic proxy for arbitrary MCP tools.
How it works
Claude Desktop connects through the PrivacyFence extension (PrivacyFence.mcpb); Claude Code and
other clients that speak Streamable HTTP connect to the local /mcp endpoint directly; in an
organization deployment, clients such as claude.ai and ChatGPT (Developer Mode) sign in with OAuth
through the organization's identity provider. How it works walks one read and one
write through, card by card.
Platforms
Install | Runs on |
macOS ( | macOS 13 or newer, Apple silicon |
Windows ( | Windows 10 / Windows Server 2016 or newer, x64 |
Linux ( | Ubuntu 24.04, Debian 13 or newer, amd64 |
Organization deployment | A Linux server with Python 3.11 or newer and systemd ( |
Tested with Claude Desktop and Claude Code on every install, with ChatGPT desktop on macOS, and with claude.ai and ChatGPT (Developer Mode) through an organization deployment; any MCP-compatible client can connect. See Platform support.
Connectors
Connector | What an AI client can do through it |
Gmail | Search and read messages, threads and attachments; create drafts and replies, labels, filters; archive. No tool sends email. |
Google Drive, Docs & Sheets | Search, read, download, upload, move and write files; edit and format Docs; read, write and format Sheets |
Google Calendar | Read events, free/busy and rooms; create, update and delete events; out-of-office and working location |
Google Contacts, Tasks | Read, create and update contacts and tasks |
Google Apps Script | Read and write project source; read the result of a run you started (PrivacyFence never runs scripts) |
Slack | List and read channels, DMs and threads; search; send messages; start group chats |
Telegram | Read and search chats; send messages |
Salesforce | Read records, search, run reports (read-only) |
Jira | Read, create, update, comment on and transition issues |
Confluence | Search and read pages and attachments; create and update pages |
Connectors summarizes what each one reviews and which writes need approval; the Tools reference lists every tool and its gate.
Quick start
Download the installer for your platform from privacyfence.eu/download and run it.
Sign out and back in once, so your account's new group membership takes effect.
Add a passkey when the companion app (menu bar, tray, or applications menu) asks, and keep the recovery code.
Open Settings from the companion and connect your services.
Connect Claude Desktop with
PrivacyFence.mcpb, or Claude Code with the/mcpendpoint.Ask your assistant for something, and approve it on the card.
Step by step for each platform: Getting started, then macOS, Windows or Linux. For claude.ai, ChatGPT or a whole team, see Organization deployment.
Documentation
Approvals and policy: cards, the PII check, always-allow rules, the privacy filter
How it works: the daemon, the MCP endpoint, PrivacyFence's own tools, unattended sessions
Configuration reference: every
settings.yamland organization-bundle keyOrganization deployment: running PrivacyFence centrally
Security and compliance: trust boundary, privilege separation, audit log
All published docs: privacyfence.eu/docs. Contributing: CONTRIBUTING.md. Changes per release: CHANGELOG.md.
Limitations
PrivacyFence is independent open-source software, not a certified compliance product. It has no
certification, business-continuity plan or SLA, and does not by itself make a deployment compliant
with any regulation. It does not protect against root or a local Administrator, or against local
code on an install that is not packaged (a source checkout or a pip install). A process running
as you can read the local review screen, though not approve from it, and in local mode the name an
AI client gives is never verified. The full list is
What PrivacyFence does not claim.
To report a vulnerability, see SECURITY.md.
License
Available Tools
122 toolsapps_script_get_contentARead-onlyIdempotent
Fetch the full source of a Google Apps Script project -- every file (.gs/.html) plus the appsscript.json manifest. Requires user approval.
| Name | Required | Description | Default |
|---|---|---|---|
| reason | Yes | One sentence: why are you calling this tool right now? | |
| script_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint and destructiveHint=false, so the safety profile is covered. The description adds the non-obvious behavioral fact that a user approval gate applies before the fetch executes, which is real context beyond the structured fields. It omits return size or rate-limit behavior, keeping it at a 4 rather than 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, zero filler, with the scope of the fetched payload front-loaded and the approval constraint at the end. Nothing needs to be trimmed.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description usefully enumerates what comes back (.gs/.html files plus the manifest), which is exactly the return-value context an agent needs. The only gap is that 'script_id' provenance and any payload-size caveats remain unaddressed.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 50%: 'reason' is documented in the schema, but 'script_id' has no schema description and the tool description supplies no format, origin, or syntax for it. Since coverage is at best half, the description should have compensated and does not.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Names a specific verb (Fetch) and resource (Google Apps Script project source) and pins down the scope precisely: 'every file (.gs/.html) plus the appsscript.json manifest.' This clearly separates it from siblings like apps_script_list_projects and apps_script_write_content.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description states a prerequisite ('Requires user approval') but never says when to reach for this tool versus apps_script_list_projects or apps_script_get_execution_log. Usage is only implied by the name and scope, so it is minimum-viable guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
apps_script_get_execution_logARead-onlyIdempotent
Read the result of the most recent run(s) of a script that the user triggered themselves outside PrivacyFence (status, duration, which function ran) -- not a live console.log transcript. Requires user approval.
| Name | Required | Description | Default |
|---|---|---|---|
| reason | Yes | One sentence: why are you calling this tool right now? | |
| script_id | Yes | ||
| max_results | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly/idempotent/non-destructive, so safety is covered. The description adds value beyond them: it discloses the approval requirement ('Requires user approval') and clarifies the data is post-hoc run results rather than a streaming log.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single dense sentence that front-loads the core action and appends the exclusion and the approval prerequisite. No filler, though the em-dash clause could be read as cramming three ideas together.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description usefully enumerates return fields (status, duration, function). It also flags the approval workflow. But it leaves two of three parameters undocumented, which is a real gap for a 3-param tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is only 33%: script_id and max_results have no descriptions in either schema or prose. The description does not compensate for these gaps, leaving the meaning of max_results (default 10, presumably truncating runs) unstated.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource (read the result of recent script runs) with scope qualifiers (status, duration, which function ran). It distinguishes itself from a live console.log transcript and implicitly from apps_script_get_content, though it doesn't name siblings explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Constrains usage to scripts 'the user triggered themselves outside PrivacyFence' and clarifies it is 'not a live console.log transcript,' which steers the agent away from a wrong assumption. It stops short of naming an alternative tool for live transcripts or for source inspection.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
apps_script_list_projectsARead-onlyIdempotent
List standalone Google Apps Script projects visible to the user (id, name, last-modified time). Container-bound scripts attached to a Sheet/Doc/Form are not returned. Auto-approved.
| Name | Required | Description | Default |
|---|---|---|---|
| reason | Yes | One sentence: why are you calling this tool right now? | |
| max_results | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint and destructiveHint=false, so the safety profile is covered structurally. The description adds useful non-structured context ('Auto-approved', container-bound exclusion), but says nothing about pagination behavior, result limits, or ordering despite a max_results parameter.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences with the primary purpose front-loaded and no filler. Efficient, though the trailing 'Auto-approved' is a fragment that sits slightly apart from the scoping content.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description helpfully enumerates returned fields (id, name, last-modified time) and states the scope boundary. The remaining gap is pagination/limit semantics around max_results, which is minor for a simple list operation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is only 50%: 'reason' is documented in the schema, but 'max_results' has only a type and default with no description. The description supplies no parameter meaning at all, so it fails to compensate for the coverage gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Specific verb 'List' plus resource 'standalone Google Apps Script projects', and it explicitly delimits scope by excluding container-bound scripts. An agent can distinguish this from apps_script_get_content or apps_script_get_execution_log without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The exclusion of container-bound scripts clearly defines when this tool applies versus when it does not, and 'Auto-approved' signals it can be called freely. It stops short of naming an alternative tool for container-bound cases, so it is strong context but not full routing.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
apps_script_write_contentA
Write new source to a Google Apps Script project. Replaces the project's entire file set -- there is no single-file/partial update, so always pass every file the project should have afterward, not just the ones you changed (fetch the current set with apps_script_get_content first if you need to preserve files you aren't touching). This only writes source -- PrivacyFence never runs the script; the user runs it themselves in the Apps Script editor once this write is approved. Requires user approval.
| Name | Required | Description | Default |
|---|---|---|---|
| files | Yes | JSON array of {"name": str, "type": "SERVER_JS"|"HTML"|"JSON", "source": str}, one entry per file (a JSON manifest file is named "appsscript" with type "JSON"). | |
| reason | Yes | One sentence: why are you calling this tool right now? | |
| script_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With annotations present the bar is lower, but the description adds substantial context the annotations do not carry: full-file replacement semantics (no partial update), the approval requirement, and the fact that the script is not executed by PrivacyFence. The only tension is that destructiveHint=false sits awkwardly beside 'replaces the project's entire file set', though the description's explicit warning makes the overwrite behavior transparent rather than hidden.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four sentences, each load-bearing: action, replacement semantics, workflow alternative, then approval/execution caveats. The destructive-ish constraint is front-loaded before the recovery guidance, so an agent that stops reading early still sees the critical rule.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a no-output-schema mutation tool with 3 required params, the description covers what the write does, what it clobbers, and the approval gate. It does not say what a successful response looks like or how failures surface, which is a minor remaining gap rather than a blocking one.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 67% and the schema already documents files and reason, while script_id is bare. The description compensates for the most consequential parameter (files) by explaining that the passed set is the complete post-write state, not a delta, which is meaningfully more than the schema's per-entry format description.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Write new source to a Google Apps Script project') and immediately constrains the scope of that write ('Replaces the project's entire file set -- there is no single-file/partial update'). This clearly distinguishes it from the sibling apps_script_get_content, which it names as the read-side counterpart.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly tells the agent when and how to use it ('always pass every file the project should have afterward, not just the ones you changed') and names the alternative workflow ('fetch the current set with apps_script_get_content first if you need to preserve files'). It also states the operational precondition ('Requires user approval') and the downstream limitation ('PrivacyFence never runs the script; the user runs it themselves').
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
calendar_create_eventC
Create a new calendar event. Requires user approval.
| Name | Required | Description | Default |
|---|---|---|---|
| color | No | Event color id (1-11) or name, e.g. "Tomato" -- see calendar_list_colors | |
| rooms | No | Comma-separated room resource email addresses to book | |
| title | Yes | ||
| reason | Yes | One sentence: why are you calling this tool right now? | |
| end_time | Yes | ISO 8601 datetime | |
| location | No | ||
| attendees | No | Comma-separated email addresses | |
| recurrence | No | RRULE line(s) to make this a recurring event, e.g. "RRULE:FREQ=WEEKLY;COUNT=10". One rule per line for more than one (RRULE plus EXDATE/RDATE/EXRULE). Omit for a non-recurring event. | |
| start_time | Yes | ISO 8601 datetime | |
| calendar_id | Yes | ||
| description | No | ||
| add_google_meet | No | Set to true to add a Google Meet video conference link |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare this is a non-read-only, non-idempotent, non-destructive write, so the safety profile is covered. The description adds one genuinely useful behavioral fact beyond annotations: the call is gated on user approval. It omits other meaningful side effects such as attendee invitations and room bookings.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences with the operation front-loaded and zero filler. It is efficient, though arguably terse for a 12-parameter mutating tool.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a creation tool with 12 parameters, no output schema, and partial annotation coverage, the description leaves out critical context: attendee notification behavior, room booking effects, timezone/format expectations, and recurrence handling. It is not adequate for the tool's complexity.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The description adds no parameter meaning at all. With only 67% schema coverage, several parameters (title, location, calendar_id, description) are undocumented in both places, and the description does nothing to compensate for that gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Create a new calendar event'), so the agent knows exactly what operation this performs. It does not distinguish itself from near-siblings such as calendar_create_out_of_office or calendar_update_event, but the core purpose is unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
'Requires user approval' is a prerequisite, not guidance on when to choose this tool over alternatives. There is no mention of when to use create vs. update_event, or when create_out_of_office is the better choice for OOO blocks.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
calendar_create_out_of_officeA
Create an out-of-office event on the primary calendar. Always auto-declines new conflicting meeting invitations that arrive while it's in effect — existing invitations already on the calendar are left alone. Requires user approval.
| Name | Required | Description | Default |
|---|---|---|---|
| title | No | Out of Office | |
| reason | Yes | One sentence: why are you calling this tool right now? | |
| end_time | Yes | ISO 8601 datetime | |
| start_time | Yes | ISO 8601 datetime | |
| decline_message | No | Message sent to organizers of auto-declined invitations |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only say readOnlyHint=false, idempotentHint=false, destructiveHint=false. The description adds the real behavioral payload: new conflicting invitations are auto-declined while in effect, existing ones are untouched, and user approval is required. Non-idempotency and what happens on overlapping/repeated calls are not disclosed, keeping it short of 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight sentences with no filler: the action is front-loaded, then the side effect, then the precondition. Every clause carries information an agent needs.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a mutation tool with no output schema, the description covers the important side effect and approval requirement. It omits edge behavior such as duplicate OOO events or what the call returns, which would be needed for full completeness.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 80%, so the schema already documents start_time, end_time, reason, and decline_message. The description contributes no parameter guidance (e.g., default title, decline_message semantics beyond the schema), so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Create an out-of-office event on the primary calendar'), which is enough to separate it from the generic calendar_create_event sibling. It doesn't explicitly name or contrast with that sibling, so 4 rather than 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives an implicit frame (OOO blocks with auto-decline) plus one hard precondition, 'Requires user approval,' which is genuinely useful. However, it never says when to prefer this over calendar_create_event or what to do if an OOO event already exists, so guidance stays implied.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
calendar_delete_eventADestructive
Delete a calendar event. For a recurring event, 'scope' controls what's deleted: 'this' (default) deletes only the given event_id; 'following' ends the series just before this instance, deleting it and every later occurrence but keeping earlier ones; 'all' deletes the entire series. Requires user approval.
| Name | Required | Description | Default |
|---|---|---|---|
| scope | No | For a recurring event: 'this', 'following', or 'all'. Ignored (has no other meaning) for a non-recurring event. | this |
| reason | Yes | One sentence: why are you calling this tool right now? | |
| event_id | Yes | ||
| calendar_id | Yes | ||
| send_updates | No | Who gets a notification email about this cancellation: 'none', 'all', or 'externalOnly'. Omit to use Calendar's own default. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true, idempotentHint=false and readOnlyHint=false, so the safety profile is covered. The description adds meaningful behavior beyond that: what each scope value destroys ('following' keeps earlier occurrences, deletes this and every later one) and that user approval is required. It stops short of covering notification side effects tied to send_updates.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the action, then the scope semantics, then the approval requirement. Every sentence carries information, though the parenthetical scope enumeration is dense against a 60%-covered schema.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a destructive mutation with annotations covering safety and no output schema, the description supplies the two things an agent most needs: what scope deletes and that approval is required. The gap is the unmentioned send_updates notification behavior, which affects a real side effect for a cancellation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 60%. The description meaningfully extends the 'scope' enum by explaining the semantics of 'following' and 'all', which the schema only names. However, it never mentions send_updates or the required 'reason' parameter, so it does not compensate for the uncovered half.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Delete a calendar event') and immediately disambiguates the only nontrivial dimension, recurring-event scope. An agent can tell this apart from calendar_update_event or calendar_get_event_details without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description details how to use the 'scope' parameter for recurring events but gives no explicit when-to-use vs alternatives (e.g. update vs delete, or which sibling handles series cancellation) and no prerequisites beyond 'requires user approval'. Usage is implied by the semantics rather than stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
calendar_get_event_detailsARead-onlyIdempotent
Fetch full details of a calendar event including attendees, description, conferencing links, and file attachments (e.g. the "Notes by Gemini" and transcript docs Google Meet attaches after a meeting ends). Each attachment's file_id can be passed to drive_get_file_content to read its content. Requires user approval.
| Name | Required | Description | Default |
|---|---|---|---|
| reason | Yes | One sentence: why are you calling this tool right now? | |
| event_id | Yes | ||
| calendar_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, and destructiveHint=false, so the safety profile is covered. The description adds value beyond them by disclosing that user approval is required and by noting the downstream pattern of feeding each attachment's file_id into drive_get_file_content, which an agent would not otherwise know.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core action and result set, followed by attachment specifics, an integration hint, and the approval requirement. Efficient, though the parenthetical example is slightly verbose. No wasted sentences.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description usefully enumerates the returned fields (attendees, description, links, attachments) and even explains the downstream use of attachment file_ids, which compensates for the missing output schema. The remaining gap is the undocumented calendar_id/event_id parameters, which is captured under parameter semantics.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is only 33% (only 'reason' is documented), and the description supplies no meaning for calendar_id or event_id — no format, ID source, or how to obtain them. With low coverage the description should compensate, but it addresses none of the parameters, so it fails the bar for this dimension.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Fetch) and resource (full details of a calendar event) and enumerates exactly what is retrieved: attendees, description, conferencing links, and attachments. This clearly distinguishes it from the listing sibling calendar_list_events, which returns summaries rather than full detail.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage (fetches detail beyond a listing) but never states when to prefer this over calendar_list_events or calendar_get_event_visibility. 'Requires user approval' is useful context but is not a when-to-use rule. Usage is inferable, not explicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
calendar_get_event_visibilityARead-onlyIdempotent
Get a calendar event's visibility setting (default, public, private, or confidential) without fetching its full details (attendees, description, etc.) the way calendar_get_event_details does. Auto-approved.
| Name | Required | Description | Default |
|---|---|---|---|
| reason | Yes | One sentence: why are you calling this tool right now? | |
| event_id | Yes | ||
| calendar_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint and destructiveHint=false, so the safety profile is covered. The description adds the 'Auto-approved' trait and the explicit contrast with the detail-fetching sibling, both of which are context beyond the structured fields, though it says nothing about return format or latency.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, front-loaded with the action and value domain, followed by the differentiator. No filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema, but the description tells the agent exactly what comes back (the four visibility values), how it differs from calendar_get_event_details, and that it is auto-approved. The only real gap is the undocumented calendar_id/event_id parameters.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is only 33% – only 'reason' is documented in the schema, while calendar_id and event_id have no descriptions. The description adds no information about any of the three parameters, so it fails to compensate for the coverage gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Get a calendar event's visibility setting') and enumerates the exact value domain (default, public, private, confidential). It explicitly distinguishes itself from the sibling calendar_get_event_details, so an agent can route correctly without opening either schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It clearly signals the selecting condition: use this when you only need visibility rather than full details, and names calendar_get_event_details as the heavier alternative. It stops short of stating when-not to use it or any prerequisite beyond that contrast.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
calendar_get_free_busyARead-onlyIdempotent
Query colleagues' schedules for a time range. For each email, tries to fetch full event details (title, time, status) when the authenticated user has calendar access; falls back to free/busy slots only when access is unavailable. Use this for meeting scheduling. Auto-approved.
| Name | Required | Description | Default |
|---|---|---|---|
| emails | Yes | Comma-separated list of email addresses | |
| reason | Yes | One sentence: why are you calling this tool right now? | |
| time_max | Yes | ISO 8601 datetime | |
| time_min | Yes | ISO 8601 datetime |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover the safety profile (readOnly/idempotent/non-destructive), and the description adds genuine behavioral context beyond that: the access-dependent fallback between full event details and free/busy slots, plus an 'auto-approved' note relevant to the approval workflow. It does not describe result shape or rate limits, but the added fallback semantics are substantive.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tightly written sentences, front-loaded with the core purpose before the fallback nuance and usage note. Every sentence carries information, though the fallback clause is dense. No padding or repetition of the schema.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only query tool with full annotation coverage, full schema coverage and no output schema, the description supplies the access-dependent return behavior an agent would otherwise have to guess, which is the main gap it needed to close. Only minor omissions (result shape specifics, pagination) remain.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all four parameters (emails, time_min, time_max, reason) are already documented in the schema. The description references emails and a time range generically but adds no format or interpretation detail beyond the schema. Baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Query colleagues' schedules for a time range') and scopes it to colleagues rather than the user's own calendar, which is the key distinction from siblings like calendar_list_events and calendar_get_event_details. An agent can identify what it does without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides the intended context ('Use this for meeting scheduling') but gives no when-not guidance and never names an alternative tool for related lookups (e.g., calendar_get_event_details, calendar_get_event_visibility). Usage is implied rather than contrasted.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
calendar_list_calendarsBRead-onlyIdempotent
List all Google Calendars for the authenticated user. Auto-approved.
| Name | Required | Description | Default |
|---|---|---|---|
| reason | Yes | One sentence: why are you calling this tool right now? |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, covering the safety profile. The description's 'Auto-approved' adds a behavioral hint that the call skips human approval, which is context beyond the annotations, but it is unexplained and no return format or pagination behavior is described.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences, front-loaded with the actual operation. Nothing is padded, though 'Auto-approved' is vague enough that it slightly under-delivers for the space it takes.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple read-only list tool with full schema coverage and annotations covering safety, this is close to sufficient. No output schema exists, so return values needn't be explained, but the unexplained 'Auto-approved' is a small loose end.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
There is one parameter ('reason'), and schema coverage is 100%, so the schema fully documents its purpose. The description adds nothing about the parameter, which is acceptable given full schema coverage — baseline 3.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('List all Google Calendars') with a clear scope qualifier ('for the authenticated user'). It implicitly distinguishes itself from sibling list tools like calendar_list_events, calendar_list_rooms, and calendar_list_colors by naming the distinct resource, though it never names those siblings directly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance on when to use this vs. alternatives or what a caller should do with the result. The trailing 'Auto-approved' is a cryptic note rather than usage direction. With over a hundred sibling tools, some routing context would be valuable.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
calendar_list_colorsARead-onlyIdempotent
List Calendar's fixed event color palette: each color's id, name (e.g. "Tomato", "Sage"), and hex background/foreground. Use a color's id or name as the color argument to calendar_create_event, calendar_update_event, or calendar_set_event_color instead of guessing a numeric id. Auto-approved.
| Name | Required | Description | Default |
|---|---|---|---|
| reason | Yes | One sentence: why are you calling this tool right now? |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, and destructiveHint=false, so the safety profile is covered structurally. The description adds genuine context beyond that: the palette is 'fixed' (safe to treat as static/cacheable) and the call is 'Auto-approved', which matters in a toolset that otherwise contains privacyfence approval workflows.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences: what it returns, when to use it, and a note on approval. The most decision-relevant information (use this instead of guessing an id) is front-loaded and there is no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, but the description enumerates the returned fields (id, name, hex background/foreground), giving examples of names. For a simple read-only enumeration tool this covers everything an agent needs to call it and use the result.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With a single parameter and 100% schema description coverage, the schema already documents 'reason' fully, so the baseline for a one-param tool is high. The description adds no further detail about the argument, but none is needed.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('List Calendar's fixed event color palette') and spells out the returned fields (id, name, hex background/foreground). An agent can distinguish it from sibling list tools like calendar_list_rooms or calendar_list_calendars without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly tells the agent when to call it: before passing a color to calendar_create_event, calendar_update_event, or calendar_set_event_color, and why ('instead of guessing a numeric id'). This names the downstream tools and the failure mode it prevents, so no inference is needed.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
calendar_list_eventsBRead-onlyIdempotent
List events from a calendar (id, title, start_time, end_time, all_day, status). No attendees, description, or links returned. Auto-approved.
| Name | Required | Description | Default |
|---|---|---|---|
| query | No | ||
| reason | Yes | One sentence: why are you calling this tool right now? | |
| time_max | No | ||
| time_min | No | ||
| calendar_id | Yes | ||
| max_results | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already establish read-only, idempotent, non-destructive behavior, so the bar is lower. The description still adds context beyond them: it names the exact output field subset (no attendees, description, or links) and discloses that calls are auto-approved, which is policy information the annotations do not carry.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, front-loaded with the core action and no filler. It could arguably merge the output-scope note more tightly, but every sentence carries information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 6-parameter tool with low schema coverage and no output schema, the description covers the return-field shape but leaves parameter semantics (time range format, max_results/pagination, required calendar_id and reason) unaddressed. Adequate but with clear gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is only 17% (just the 'reason' param), leaving calendar_id, query, time_min, time_max, and max_results undocumented. The description lists returned event fields rather than explaining any parameter's format, accepted time syntax, or default behavior, so it fails to compensate for the coverage gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('List events from a calendar') and enumerates the returned fields (id, title, start_time, end_time, all_day, status), which precisely bounds the tool's scope. It does not explicitly contrast with the closest sibling calendar_get_event_details, so it stops short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no statement of when to use this versus calendar_get_event_details, calendar_get_free_busy, or calendar_list_calendars. The note about what is not returned is a scope hint, not usage guidance, and 'Auto-approved' says nothing about selection criteria.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
calendar_list_roomsARead-onlyIdempotent
List meeting rooms and resource calendars from the organization's room directory. This is a locally-cached list IT refreshes with scripts/sync_room_directory.py, not a live Workspace lookup — it may come back empty if IT hasn't synced one yet. Returns room name, email, building, floor, and capacity. To check whether a room is actually free before booking, call calendar_get_free_busy with its resource_email. Use the room email with calendar_create_event or calendar_update_event to book. Auto-approved.
| Name | Required | Description | Default |
|---|---|---|---|
| query | No | Optional substring filter matched against room name, building, floor, or description (case-insensitive) | |
| reason | Yes | One sentence: why are you calling this tool right now? |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare a safe read (readOnlyHint, idempotentHint, destructiveHint false), and the description adds meaningful context beyond them: the data is IT-synced cache rather than a live Workspace lookup, it can be empty, the refresh script is named, and it is auto-approved. This is exactly the extra behavioral context annotations cannot convey.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four tight sentences, front-loaded with what the tool returns, then the cache caveat, then the routing to sibling tools. No sentence is filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description compensates by enumerating returned fields, and it covers freshness/staleness, availability-check workflow, and booking workflow. An agent has everything needed to call and interpret this correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so both parameters including the query filter semantics are already documented in the schema. The description's list of returned fields (name, email, building, floor, capacity) describes output rather than input semantics, so it adds no parameter meaning. Baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('List meeting rooms and resource calendars') plus the scope ('organization's room directory'), and it is clearly distinguishable from calendar_list_calendars and calendar_list_events. An agent can identify the tool's output domain without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly routes to alternatives: it names calendar_get_free_busy for availability checks and calendar_create_event / calendar_update_event for booking, and it warns the list is a locally-cached snapshot that may be empty. When-to-use, when-not, and follow-up tools are all stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
calendar_set_event_colorA
Set a calendar event's color. Only the color changes — no other fields are affected. Accepts a color id (1-11) or name, e.g. "Tomato" -- see calendar_list_colors. Requires user approval.
| Name | Required | Description | Default |
|---|---|---|---|
| color | Yes | Event color id (1-11) or name, e.g. "Tomato" | |
| reason | Yes | One sentence: why are you calling this tool right now? | |
| event_id | Yes | ||
| calendar_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare non-read-only, non-idempotent, non-destructive behavior, so the safety profile is covered. The description adds genuinely useful context beyond that: the mutation is field-scoped (only color changes) and requires user approval, which is not expressed in annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three compact sentences, front-loaded with the action, followed by scope, accepted values, and the approval requirement. No filler and every sentence contributes.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a mutation tool with no output schema, the description covers the mutation scope, value domain reference, and approval gate. The main residual gap is the undocumented calendar_id/event_id semantics, which neither schema nor description clarify.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 50%: color and reason are documented in the schema, while calendar_id and event_id are bare. The description restates the color id range and name example already present in the schema, adding little, and leaves the two identifier parameters unexplained.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Set a calendar event's color') and immediately scopes it against the nearest sibling behavior by clarifying that no other fields change, which separates it from calendar_update_event and calendar_set_event_visibility.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Routes the agent to calendar_list_colors for valid color values and states the approval precondition, giving clear context for invocation. It does not, however, state when NOT to use it or name a competing alternative explicitly.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
calendar_set_event_visibilityA
Set a calendar event's visibility. 'default' follows the calendar's own sharing settings; 'public' makes it visible to anyone who can see the calendar; 'private' hides its details from viewers who aren't invited; 'confidential' is a legacy synonym the Calendar API still accepts for 'private'. Only visibility changes — no other fields are affected. Requires user approval.
| Name | Required | Description | Default |
|---|---|---|---|
| reason | Yes | One sentence: why are you calling this tool right now? | |
| event_id | Yes | ||
| visibility | Yes | 'default', 'public', 'private', or 'confidential' | |
| calendar_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare the write/safety profile (readOnlyHint=false, destructiveHint=false), and the description adds real value on top: it scopes the mutation ('Only visibility changes — no other fields are affected') and discloses a user-approval gate, which annotations do not convey. It stops short of describing failure or permission behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One tight paragraph, front-loaded with the action, followed by value semantics and scope limits. Every clause carries information an agent needs; nothing is padding.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 4-required-param mutation with no output schema, the description covers action, value semantics, scope containment, and the approval requirement. It omits what errors or confirmation look like, but the essentials for correct invocation are present.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 50% and the description compensates for the most important gap by spelling out the semantics of each 'visibility' value, including that 'confidential' is a legacy synonym for 'private'. calendar_id and event_id remain undocumented, but their meaning is self-evident.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Set a calendar event's visibility') that is unambiguously distinct from calendar_get_event_visibility and calendar_update_event. An agent can pick this tool without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives a clear context for invocation (changing visibility only, with a stated approval requirement) and explains what each mode means, but never routes the agent explicitly to or away from the sibling calendar_get_event_visibility or calendar_update_event.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
calendar_set_working_locationA
Set your working-location presence (office or home) for a single day on the primary calendar — the same picker Google Calendar's web UI exposes. Requires user approval.
| Name | Required | Description | Default |
|---|---|---|---|
| date | Yes | ISO 8601 date, e.g. 2026-07-10 | |
| label | No | Office label shown on Calendar (office only) | |
| reason | Yes | One sentence: why are you calling this tool right now? | |
| location | Yes | "office" or "home" | |
| building_id | No | Workspace building id (office only; see calendar_list_rooms) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=false, idempotentHint=false, destructiveHint=false, so the safety profile is known. The description adds genuinely new context: the operation is scoped to a single day on the primary calendar only, and it requires user approval. It still doesn't say what a re-call does or whether the location can be cleared.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, zero filler: the scope and behavior lead, and the approval requirement is a short trailing clause. Nothing could be removed without losing meaning.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a non-idempotent write with full annotation coverage and no output schema, the essentials (scope, target calendar, approval gate) are present. Minor gaps remain around ambiguity with the out-of-office sibling and the effect of repeating the call, but nothing blocking correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% with descriptions for every parameter including the office-only qualifiers for label and building_id. The description adds no parameter syntax or constraints beyond what the schema already provides, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Specific verb ('Set') plus resource ('working-location presence'), with explicit scope: office or home, single day, primary calendar. It also names the concrete analogue ('the same picker Google Calendar's web UI exposes'), which lets an agent distinguish it from the other calendar write tools like calendar_create_out_of_office or calendar_create_event.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage by tying it to the Google Calendar web-UI picker and notes 'Requires user approval,' which is a real precondition. However it never states when to prefer this over siblings (e.g. calendar_create_out_of_office) or when not to call it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
calendar_update_eventA
Update an existing calendar event. For a recurring event, 'scope' controls which occurrences this touches: 'this' (default) affects only the given event_id; 'following' splits the series so this instance and every later one get the changes, leaving earlier ones untouched; 'all' updates the entire series. Requires user approval.
| Name | Required | Description | Default |
|---|---|---|---|
| color | No | Event color id (1-11) or name, e.g. "Tomato" -- see calendar_list_colors | |
| rooms | No | Comma-separated room resource email addresses to book | |
| scope | No | For a recurring event: 'this', 'following', or 'all'. Ignored (has no other meaning) for a non-recurring event. | this |
| title | No | ||
| reason | Yes | One sentence: why are you calling this tool right now? | |
| end_time | No | ||
| event_id | Yes | ||
| location | No | ||
| start_time | No | ||
| calendar_id | Yes | ||
| description | No | ||
| send_updates | No | Who gets a notification email about this change: 'none', 'all', or 'externalOnly'. Omit to use Calendar's own default. | |
| add_google_meet | No | Set to true to add a Google Meet link (skipped if one already exists) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only say readOnly=false, idempotent=false, destructive=false; the description adds meaningful context beyond that by disclosing the approval gate and the series-splitting consequence of 'following' (earlier occurrences untouched). It still omits partial-update semantics (whether omitted fields are cleared or preserved), which matters for a 13-param mutation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Purpose is front-loaded in the first sentence, followed by the scope semantics and the approval requirement. Every sentence carries information the agent needs and there is no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 13-parameter mutation with no output schema, the description covers scope and approval but omits several decision-relevant behaviors: whether this is a partial or full replace of event fields, what send_updates defaults to, and what the result of a series split looks like. Adequate but with real gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 46%, below the threshold where the schema can carry the load. The description compensates well for the riskiest parameter by explaining the real consequence of 'following' (series split) that the schema entry does not, and clarifies the 'this' default; remaining undocumented params (start_time, title, event_id) are largely self-explanatory.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Update an existing calendar event'), which separates it from calendar_create_event and calendar_delete_event by implication. It does not explicitly distinguish itself from adjacent mutators like calendar_set_event_color or calendar_set_event_visibility, so the boundary for those overlapping edits is left to inference.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The scope enumeration effectively documents when each occurrence-mode applies, and 'Requires user approval' flags a prerequisite. However, there is no guidance on when to choose this general update over a narrower sibling such as calendar_set_event_color, nor any exclusion criteria.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
confluence_cql_searchBRead-onlyIdempotent
Search Confluence using CQL (Confluence Query Language). Auto-approved.
| Name | Required | Description | Default |
|---|---|---|---|
| cql | Yes | e.g. 'space = MYSPACE AND type = page' | |
| reason | Yes | One sentence: why are you calling this tool right now? | |
| max_results | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint and destructiveHint=false, so the safety profile is covered. The description adds a piece of genuinely new context — 'Auto-approved' signal that no human approval gate applies (relevant given the privacyfence approval siblings) — but says nothing about pagination, result limits, or query cost.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short, front-loaded sentences with no padding; the core action leads. The trailing 'Auto-approved' is curt and slightly cryptic, but nothing is wasted.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 3-parameter, no-output-schema search tool, the description omits what results look like and how max_results affects them. Annotations cover safety, but the return/limit behavior is left undocumented for an agent that must decide result sizing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 67%: cql has an example and reason has explanatory text, but max_results has only a default with no description. The description adds no syntax or format guidance for cql beyond naming the language, so the schema still does most of the work.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a clear verb+resource (search Confluence) and names the query language CQL, so the agent knows this is a CQL-driven search. It does not differentiate itself from the sibling confluence_search, which likely serves a different query mode, leaving the agent to infer the distinction from the tool name alone.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to use this over confluence_search, nor any stated prerequisites. 'Auto-approved' hints at the approval workflow but is not framed as a use condition.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
confluence_create_pageA
Create a new Confluence page in the given space. Body is HTML storage format. Requires user approval.
| Name | Required | Description | Default |
|---|---|---|---|
| body | Yes | HTML storage format body | |
| title | Yes | ||
| reason | Yes | One sentence: why are you calling this tool right now? | |
| parent_id | No | Optional parent page ID | |
| space_key | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare a non-readonly, non-idempotent, non-destructive mutation. The description adds meaningful context beyond them: the body must be in HTML storage format and the operation requires user approval, which is important behavioral information for an agent.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is three short, front-loaded sentences with no wasted words. It immediately states the action, then adds the body format and approval prerequisite in order of importance.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a create tool with annotations covering safety and no output schema, the description provides the key missing context: HTML storage format and the user approval requirement. It is largely complete, though it could clarify return behavior or approval failure conditions.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 60%, so the schema already documents several parameters. The description reinforces the body format and indicates that space_key refers to the target space, but it does not fully compensate for undocumented parameters like title or explain required parameters such as reason beyond what the schema provides.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: 'Create a new Confluence page in the given space.' The verb 'Create' clearly distinguishes it from sibling operations like confluence_update_page, confluence_get_page, and confluence_list_pages.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It mentions a prerequisite ('Requires user approval') but gives no guidance on when to select this tool versus alternatives such as confluence_update_page or other creation tools. Usage is only implied by the verb 'Create'.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
confluence_download_attachmentARead-onlyIdempotent
Download a Confluence page attachment's content. Identify the attachment by the name returned from confluence_list_attachments. On a local install: saved to destination_dir, and the saved file path is returned -- destination_dir is required, there is no default, so choose deliberately: pass ~/Downloads (or another path the user asked for) when this attachment is a deliverable the user should find afterward, or your own working/scratch directory when you're only downloading it to read or process it yourself. On an organization-managed install: destination_dir is ignored (there is no local filesystem you and the human share) -- a small attachment's bytes come back directly in this tool's result so you can read or hand it to the human yourself; a larger one comes back as a one-time link the human opens in their own signed-in browser tab instead. Requires user approval.
| Name | Required | Description | Default |
|---|---|---|---|
| reason | Yes | One sentence: why are you calling this tool right now? | |
| page_id | Yes | ||
| attachment_name | Yes | ||
| destination_dir | Yes | Local install only: where to save the attachment -- required, no default. On an organization-managed install it is ignored and nothing is saved to it; any value will do. Use ~/Downloads (or a path the user specified) if the user should find this file afterward; use your own working/scratch directory if it's only for you to read or process. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations cover the safety profile (readOnly, idempotent, non-destructive), and the description adds substantial context beyond them: local vs organization-managed install behavior, what is actually returned (file path, raw bytes, or a one-time signed-in link), and the user-approval requirement. This is exactly the behavioral depth annotations cannot convey.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core action, then the install-dependent branching, and every clause carries real information. It is on the verbose side, but the length is largely justified by the genuinely branching behavior.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, and the description explicitly describes the three possible return outcomes so an agent knows what to expect. Combined with the approval and destination guidance, nothing needed to invoke it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 50% schema coverage, the description compensates: it explains destination_dir's required-with-no-default semantics and the consequence of the choice, and clarifies that attachment_name comes from list_attachments. page_id and reason are undocumented here but self-evident or covered by the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Download) and resource (a Confluence page attachment's content), and it names the sibling confluence_list_attachments as the source of the attachment name. An agent can distinguish it from drive_download_file and gmail_download_attachment without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear context: obtain the attachment_name from confluence_list_attachments before calling, and choose destination_dir deliberately depending on whether the file is a user deliverable or scratch. It lacks an explicit when-not / alternative (e.g. when to prefer a page-read instead), so it falls just short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
confluence_get_pageARead-onlyIdempotent
Fetch the full content of a Confluence page by page ID. Returns the page body as HTML storage format. Requires user approval.
| Name | Required | Description | Default |
|---|---|---|---|
| reason | Yes | One sentence: why are you calling this tool right now? | |
| page_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, and destructiveHint=false. The description adds genuinely new context beyond them: the return is HTML storage format and the call 'Requires user approval', which is an important operational constraint not captured in annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight sentences with zero waste; the core action is front-loaded and the return format plus approval requirement follow logically. Every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, but the description covers the return value (HTML storage format) and the approval requirement. The main gap is the lack of any routing guidance against the numerous sibling retrieval tools.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 50%: the 'reason' parameter is documented in the schema, but 'page_id' has no schema description. The description only implies page_id's meaning ('by page ID') without adding format or source details, so coverage of the undocumented parameter remains thin. Baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Fetch') and resource ('full content of a Confluence page') scoped by page ID, and discloses the return format. It does not explicitly differentiate from the close sibling confluence_get_page_by_title, but the 'by page ID' qualifier implicitly distinguishes it.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No when-to-use or when-not-to-use guidance. It does not tell the agent when to prefer this over confluence_get_page_by_title, confluence_search, or confluence_cql_search, leaving routing to inference from the tool name alone.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
confluence_get_page_by_titleARead-onlyIdempotent
Fetch a Confluence page by space key and exact title. Requires user approval.
| Name | Required | Description | Default |
|---|---|---|---|
| title | Yes | ||
| reason | Yes | One sentence: why are you calling this tool right now? | |
| space_key | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover the safe-read profile (readOnly, idempotent, non-destructive), and the description adds a genuinely new behavioral trait: 'Requires user approval.' That approval requirement materially affects how an agent should invoke the tool and is not present in the structured data.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences, front-loaded with the lookup action and keys, followed immediately by the approval constraint. Zero wasted words.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple read tool with no output schema and annotations covering safety, the description is largely sufficient. However, it is silent on what happens when the title is not exact or multiple pages match, and parameter coverage is thin, leaving minor gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is only 33% (only 'reason' is documented), so the schema does not carry the burden. The description compensates partially by tying space_key and title to their meaning and stressing 'exact title,' but it gives no format guidance (space key syntax, partial-match behavior).
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Fetch), resource (Confluence page), and the lookup keys (space key and exact title), which distinguishes it from confluence_get_page (ID-based) and the search variants without opening a schema. It stops short of explicitly naming those siblings, so it is clear but not fully differentiated.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'by space key and exact title' implies when this lookup is appropriate, but there is no explicit when-to-use/when-not guidance and no mention of confluence_get_page or confluence_search as alternatives. Usage is inferable rather than stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
confluence_list_attachmentsARead-onlyIdempotent
List attachment names, media types, and sizes for a Confluence page. Auto-approved -- metadata only, no attachment content is returned. Use confluence_download_attachment to fetch the actual file.
| Name | Required | Description | Default |
|---|---|---|---|
| reason | Yes | One sentence: why are you calling this tool right now? | |
| page_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover readOnly/idempotent/non-destructive, and the description adds meaningful context beyond them: the call is auto-approved and returns metadata only, never attachment content. It doesn't discuss pagination or handling of pages with many attachments, so not a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, front-loaded with what is returned, then the safety/metadata note, then the alternative tool. No filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
A simple read-only list tool with annotations covering the safety profile and no output schema to explain; the description states what fields come back and which tool to use instead. Nothing needed to call it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is only 50% (page_id has no schema description) and the description only implies that page_id is a Confluence page rather than specifying its format (numeric ID vs. key). It adds marginal meaning over the schema, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Specific verb (list) + resource (attachments) + scope (for a Confluence page) + the exact fields returned (names, media types, sizes). It is immediately distinguishable from the sibling confluence_download_attachment.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly routes the agent to confluence_download_attachment when actual file content is needed, which is the key alternative in this cluster. No explicit exclusions or prerequisites are stated (e.g., permissions to view the page), so it falls just short of 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
confluence_list_pagesCRead-onlyIdempotent
List pages in a Confluence space (title, id, version). Auto-approved.
| Name | Required | Description | Default |
|---|---|---|---|
| reason | Yes | One sentence: why are you calling this tool right now? | |
| space_key | Yes | ||
| max_results | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, so the safety profile is covered structurally. The description adds "Auto-approved," which is mildly useful given the surrounding privacyfence/approval tooling, but says nothing about pagination, result caps, or ordering behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Very short and front-loaded: the core purpose leads, with the field list and approval note as trailing fragments. No padding, though the terse fragments do little work.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema and only 33% parameter coverage, the description should carry more of the burden. It omits anything about pagination, the max_results cap behavior, or how to obtain a valid space_key, leaving real gaps for a listing tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is only 33%: 'reason' is documented in the schema, but 'space_key' and 'max_results' carry no description anywhere. The description mentions returned fields rather than parameter meaning and does not explain the expected space_key format or how max_results interacts with result ordering/pagination.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Names a specific verb and resource ("List pages in a Confluence space") and even enumerates the returned fields (title, id, version). It is clear on its own, but it does not distinguish itself from siblings like confluence_search, confluence_cql_search, or confluence_get_page_by_title, which an agent must disambiguate.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no explicit when-to-use or when-not-to-use guidance and no named alternative. The only contextual hint is "Auto-approved," which does not tell the agent when this list tool beats the search or CQL siblings.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
confluence_list_spacesARead-onlyIdempotent
List Confluence spaces the user has access to (key, name, type, description). Auto-approved.
| Name | Required | Description | Default |
|---|---|---|---|
| reason | Yes | One sentence: why are you calling this tool right now? | |
| space_type | No | Filter to 'global' or 'personal'; omit/empty to return all types | |
| max_results | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly/idempotent/non-destructive, and the description adds two genuinely new traits: the access-scoping ('spaces the user has access to') and 'Auto-approved', which tells the agent the call will not block on human approval. It still omits pagination behavior tied to max_results.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences, zero waste, with the scoping constraint and returned fields front-loaded. Nothing extraneous.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only list tool with no output schema, explicitly enumerating the returned fields (key, name, type, description) is valuable. It is nearly complete, lacking only expected pagination/truncation behavior for max_results.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 67%; space_type and reason are documented in the schema, while max_results has no description. The description adds no parameter-level syntax or format detail beyond the schema, so it sits at the baseline for this coverage level.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
It names a specific verb and resource (list Confluence spaces) with a clear scope qualifier ('the user has access to') and enumerates the returned fields. It is distinguishable from confluence_search and confluence_list_pages by resource, though it never explicitly contrasts with them.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No when-to-use guidance or alternatives are given. An agent must infer that this is for browsing spaces rather than finding pages, and nothing tells it when to prefer confluence_search or confluence_cql_search instead.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
confluence_searchBRead-onlyIdempotent
Full-text search across Confluence content. Returns matching pages/blog posts with excerpts. Auto-approved.
| Name | Required | Description | Default |
|---|---|---|---|
| query | Yes | Plain-text search terms | |
| reason | Yes | One sentence: why are you calling this tool right now? | |
| max_results | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, and destructiveHint=false, so the safety profile is covered. 'Auto-approved' adds a governance fact beyond the annotations, which is genuine value, but the description says nothing about result limits, pagination, or scope restrictions.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight sentences with purpose and return value front-loaded and little waste. 'Auto-approved' is terse to the point of ambiguity about who or what approves, which keeps it from a 5.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple read-only search with no output schema, stating that excerpts are returned is adequate, but the tool omits result-count/pagination behavior tied to max_results and gives no distinction from the CQL search sibling, leaving gaps for correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 67%: query and reason are documented in the schema, while max_results (default 20) is undocumented in both schema and description. The description adds no parameter-level meaning beyond the schema, so the baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Full-text search across Confluence content') and names what it returns (pages/blog posts with excerpts). However, it does not distinguish itself from the sibling confluence_cql_search, which an agent needs to choose between, so it falls short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No when-to-use guidance, no conditions, and no mention of alternatives such as confluence_cql_search or confluence_list_pages. 'Auto-approved' describes governance behavior, not usage context.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
confluence_update_pageA
Update the title and/or body of an existing Confluence page. Body is HTML storage format. Requires user approval.
| Name | Required | Description | Default |
|---|---|---|---|
| body | Yes | New HTML storage format body | |
| title | Yes | ||
| reason | Yes | One sentence: why are you calling this tool right now? | |
| page_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare readOnlyHint=false, idempotentHint=false, and destructiveHint=false; the description adds two behavioral facts beyond those: the body format (HTML storage format) and the requirement for user approval. It does not explain what approval entails or whether page versions are preserved, but it adds meaningful context.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short, front-loaded sentences with no filler. The most important facts — what it does, the body format, and the approval requirement — are all stated directly.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a mutation tool with no output schema and annotations covering the safety profile, the description is minimally adequate. However, it leaves gaps about how page_id is obtained, whether updates overwrite content, and the mismatch between 'and/or' and the schema's required title/body.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is only 50% (page_id and title have no schema descriptions). The description only restates the body format (already in schema) and incorrectly says 'title and/or body' is updatable, while the schema marks both title and body as required. It provides no help for page_id or reason.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action (update) and resource (title and/or body of an existing Confluence page). The word 'existing' implicitly distinguishes it from confluence_create_page, but no sibling is explicitly named or contrasted.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied by 'update ... existing Confluence page', which suggests this tool is not for creating new pages. However, it gives no explicit when-to-use guidance, alternatives, or exclusions beyond the approval requirement.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
contacts_add_labelA
Add a label to a contact, creating the label if it doesn't already exist. Requires user approval.
| Name | Required | Description | Default |
|---|---|---|---|
| reason | Yes | One sentence: why are you calling this tool right now? | |
| label_name | Yes | ||
| resource_name | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations cover the safety profile (readOnlyHint=false, destructiveHint=false, idempotentHint=false), and the description adds genuinely new behavioral context: the label is created on demand as a side effect, and an approval step is required before the write. It stops short of describing failure modes or permission scope, but it clearly exceeds the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two dense sentences, front-loaded with the action and its upsert behavior, followed by the approval prerequisite. No filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a small mutation tool with annotations carrying the safety profile and no output schema, the description covers the key surprises (implicit label creation, approval gate). It could add parameter format detail, but nothing critical for correct invocation is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is only 33% – only 'reason' is documented in the schema. The description names the two concepts ('contact', 'label') but gives no format for resource_name (contact ID? email? resource path?) or label_name, so it does not compensate for the coverage gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Add a label to a contact') and adds the upsert nuance ('creating the label if it doesn't already exist'). An agent can immediately distinguish this from the inverse sibling contacts_remove_label.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The prerequisite 'Requires user approval' is stated, which is useful operational context. However, there is no explicit guidance on when to prefer this over contacts_update (which may also set labels) or on how the label should be identified, leaving usage to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
contacts_createA
Create a new contact in the user's Google address book. Requires user approval. emails and phones are JSON strings, e.g. '[{"value": "a@b.com", "type": "work"}]'. Contact deletion is not supported.
| Name | Required | Description | Default |
|---|---|---|---|
| notes | No | ||
| emails | No | JSON array of {value, type} dicts | |
| phones | No | JSON array of {value, type} dicts | |
| reason | Yes | One sentence: why are you calling this tool right now? | |
| job_title | No | ||
| display_name | Yes | ||
| organization | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already establish write semantics (readOnlyHint=false, destructiveHint=false, idempotentHint=false), and the description adds non-obvious behavior: an approval gate before execution and that deletion is unavailable. It omits what happens on duplicate display_name and whether the response carries the new contact's ID, which keeps it from a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, front-loaded with the core action, then precondition, then format hint. Nothing is padded, though the format example is embedded mid-sentence rather than separated, making it slightly harder to scan.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 7-parameter mutation tool with no output schema, the description covers the risky parts: approval requirement, non-destructive nature, and the JSON encoding of the two list-like fields. It lacks any note on duplicate handling or on the 'reason' argument, leaving small gaps but nothing that would cause a failed call.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With only 43% schema description coverage, the description compensates by clarifying the two trickiest fields, emails and phones, as strings and giving a literal example '[{"value": "a@b.com", "type": "work"}]'. The remaining fields (notes, job_title, organization, display_name) are self-evident, but the required 'reason' parameter gets no mention here.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Create a new contact in the user's Google address book'), which is unambiguous and clearly distinct from contacts_update, contacts_list, and contacts_search. It does not explicitly name a sibling alternative, but no other create-contact tool exists, so the risk of confusion is low.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description adds one real precondition ('Requires user approval') and one scope exclusion ('Contact deletion is not supported'), which help an agent decide whether the tool applies. However, it never says when to prefer this over contacts_update or how to handle an existing contact, so usage guidance is only implied.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
contacts_getARead-onlyIdempotent
Fetch a single contact by resource name (e.g. 'people/c12345'). 'source' asserts the expected kind of contact ('personal', 'directory', or 'both'/default); the call fails if the resource doesn't match. Auto-approved.
| Name | Required | Description | Default |
|---|---|---|---|
| reason | Yes | One sentence: why are you calling this tool right now? | |
| source | No | 'personal', 'directory', or 'both' (default). | both |
| resource_name | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly/idempotent/non-destructive, so the safety profile is covered. The description adds genuine behavioral detail beyond that: the call fails on a source/resource mismatch, and 'Auto-approved' discloses the approval flow. Return-shape details are absent, but no output schema exists.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight sentences, front-loaded with the core action, with no repetition of the schema's enum list or the tool name.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a single-resource read with two required params and no output schema, the description covers the identifier format, the source assertion behavior, and the approval mode. The only real omission is routing guidance against contacts_search.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 67% and the schema only lists the enum values for 'source'. The description goes further, explaining that 'source' asserts the expected contact kind and causes failure on mismatch, plus it documents the resource_name format by example. 'reason' is left to its self-explanatory schema text.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Fetch a single contact by resource name') and includes the canonical identifier format 'people/c12345'. This clearly distinguishes it from the sibling contacts_list and contacts_search tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage (fetch one contact when you already have its resource name) but never explicitly names contacts_search or contacts_list as the alternative for lookup-by-attribute. No when-not-to-use guidance is given beyond the source assertion semantics.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
contacts_listARead-onlyIdempotent
List contacts from the user's Google address book. Google blends personally-saved contacts together with Workspace directory profiles (colleagues) by default; use 'source' to split them apart. Returns display name, emails, phones, organization, job title, and a 'source' field ('personal', 'directory', or 'both' if the same person is both a saved contact and a colleague). Auto-approved.
| Name | Required | Description | Default |
|---|---|---|---|
| reason | Yes | One sentence: why are you calling this tool right now? | |
| source | No | 'personal' (saved contacts only), 'directory' (Workspace directory only), or 'both' (default). | both |
| max_results | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The annotations already declare readOnlyHint, idempotentHint, and destructiveHint=false, so the safety profile is covered. The description adds useful context beyond that: the default blending of personal and directory contacts, the meaning of the 'source' return field, and the 'Auto-approved' behavior. It stops short of describing pagination or rate limits, so a 4 is appropriate rather than a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is three tightly focused sentences: purpose first, then the source-blending behavior, then the return fields. Every sentence contributes useful information and nothing is redundant. It is concise but not maximally terse, so a 4 is warranted.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema, but the description lists the returned fields (display name, emails, phones, organization, job title, source), which compensates well for that absence. Annotations cover the read-only safety profile, and the source behavior is explained. The main omission is any detail about max_results or pagination, which keeps it from a 5.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 67%, with 'reason' and 'source' documented but 'max_results' lacking any description in either schema or tool text. The description does add meaningful rationale for the 'source' parameter by explaining the Google blending behavior, but it does not compensate for the missing 'max_results' semantics. This fits the baseline 3 for middling schema coverage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: 'List contacts from the user's Google address book.' It also clarifies the default blended scope and the source field, which makes the tool's purpose concrete. However, it does not explicitly differentiate itself from the sibling contacts_search tool, so it does not reach the sibling-differentiation standard of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives guidance for the 'source' parameter ('use source to split them apart'), which implies when to filter by personal vs directory. But it never states when to use this listing tool versus alternatives like contacts_search or contacts_get, nor does it name any alternative. Usage is therefore only implied, matching a 3.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
contacts_remove_labelB
Remove a label from a contact. Requires user approval.
| Name | Required | Description | Default |
|---|---|---|---|
| reason | Yes | One sentence: why are you calling this tool right now? | |
| label_name | Yes | ||
| resource_name | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations supply the safety profile (readOnlyHint=false, destructiveHint=false, idempotentHint=false), so the description need not restate them. It does add one genuinely useful behavioral fact beyond annotations: the call is gated on user approval. However, it says nothing about whether the removal is reversible, what happens if the label is absent, or what the response contains.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short, front-loaded sentences with zero filler. The action is stated first and the constraint second.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a mutation tool with no output schema and two undocumented required parameters, the description is thin: an agent knows what to do and that approval is needed, but not how to properly populate resource_name/label_name or what the outcome of the call is.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is only 33%: reason is described in the schema, but resource_name and label_name carry no description anywhere. The description provides no meaning for these two required parameters, so it fails to compensate for the coverage gap in a 3-parameter tool.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ("Remove a label from a contact"), which is immediately distinguishable from contacts_add_label and contacts_update among the siblings. It stops short of explicitly naming the inverse tool for differentiation, but the purpose is unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to use this tool versus contacts_update (which may also alter labels) or contacts_add_label. The only contextual cue is "Requires user approval," which is a precondition, not usage routing.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
contacts_searchARead-onlyIdempotent
Search contacts by name or email address. Use 'source' to search only personally-saved contacts, only Workspace directory contacts, or both (default). Note: 'directory' search only finds directory profiles you already have some contact history with; there is no full company-directory search under this app's permissions. Auto-approved.
| Name | Required | Description | Default |
|---|---|---|---|
| query | Yes | ||
| reason | Yes | One sentence: why are you calling this tool right now? | |
| source | No | 'personal', 'directory', or 'both' (default). | both |
| max_results | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare read-only, idempotent, non-destructive behavior. The description adds meaningful context beyond them: the permission constraint on directory search, the fact that no full-company search exists, and the 'Auto-approved' status. It omits pagination/return-format details, which is minor for a read tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, front-loaded with the core action before the source semantics and the caveat. Every sentence carries information; no filler or repetition.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only search with no output schema, the description supplies the essential scope rules and the directory-search limitation. The remaining gap (max_results behavior) is minor and not critical to correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 50%, so the description must carry weight. It clarifies what 'source' does and its default beyond the schema's bare enum list, but adds nothing for 'query' or 'max_results', leaving those undocumented in both places.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Search contacts by name or email address'), which clearly separates it from the sibling contacts_list (enumerate) and contacts_get (retrieve one). It does not explicitly name those siblings, so differentiation is implied rather than stated.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explains how to use 'source' to scope results and flags a key limitation ('directory' only finds profiles you already have contact history with). This tells the agent when directory search will and won't work, though it doesn't contrast against contacts_list.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
contacts_updateA
Update a contact's fields. Provide only the fields you want to change. Requires user approval. emails and phones are JSON strings, e.g. '[{"value": "a@b.com", "type": "work"}]'.
| Name | Required | Description | Default |
|---|---|---|---|
| notes | No | ||
| emails | No | JSON array of {value, type} dicts | |
| phones | No | JSON array of {value, type} dicts | |
| reason | Yes | One sentence: why are you calling this tool right now? | |
| job_title | No | ||
| display_name | No | ||
| organization | No | ||
| resource_name | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare non-read-only, non-idempotent, non-destructive, so the description adds value by disclosing the approval requirement and partial-update semantics. It does not detail auth scope, error behavior, or full side effects.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four short sentences, front-loaded with purpose, then usage, approval, and formatting. No filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Covers purpose, partial update, approval, and complex param format, but omits required resource_name/reason semantics and most mutable fields for an 8-param mutation with no output schema. Adequate but incomplete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is low at 38%, so description must compensate, but it only explains emails/phones JSON formatting (with example) and leaves resource_name, reason, notes, job_title, display_name, and organization undocumented. This is partial compensation, not enough for 8 params.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource: update a contact's fields. Clear against siblings like contacts_create/get/list, though it does not explicitly differentiate or name alternatives.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives partial-update instruction ('Provide only the fields you want to change') and a prerequisite ('Requires user approval'). No when-to-use vs alternatives or exclusions, so usage is implied rather than fully guided.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
drive_add_commentB
Add a comment to a Drive file. Requires user approval.
| Name | Required | Description | Default |
|---|---|---|---|
| reason | Yes | One sentence: why are you calling this tool right now? | |
| comment | Yes | ||
| file_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=false, idempotentHint=false, destructiveHint=false, so the agent knows this is a non-idempotent write. The description usefully adds that user approval is required, a behavioral constraint absent from the annotations, but says nothing about failure modes, duplicates, or post-write state.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences, action stated first and the precondition second. Zero filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 3-required-param write tool with no output schema, the definition covers the essential action and the approval gate, which is the key operational fact. It is thin on parameter detail and return behavior, but the gaps are modest given the self-explanatory parameter names.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is only 33%: only 'reason' is documented, while 'comment' and 'file_id' are bare. The description adds no parameter meaning at all, so it fails to compensate for the coverage gap, leaving format expectations unclear.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (add) and resource (comment on a Drive file), which is unambiguous. It does not differentiate itself from adjacent write tools like drive_write_file_content, but the action is distinct enough that no confusion arises.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
"Requires user approval" signals a precondition for invoking it, which is real usage guidance. However, there is no explicit when-to-use vs when-not, and no mention of alternatives for other comment-like targets (e.g. jira_add_comment).
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
drive_create_blank_fileB
Create a new blank Drive file. Auto-approved.
| Name | Required | Description | Default |
|---|---|---|---|
| name | Yes | ||
| reason | Yes | One sentence: why are you calling this tool right now? | |
| mime_type | Yes | ||
| parent_folder_id | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare it is a non-readOnly, non-idempotent, non-destructive mutation, so the safety profile is covered by structured data. 'Auto-approved' adds useful context about the privacy-fence approval flow, but the description never says whether repeated calls create duplicates or what the mutation returns.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences, front-loaded with the action, and no filler. Every word earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a non-idempotent creation tool with no output schema and 75% of parameters undocumented, the description is inadequate. An agent cannot tell what identifier the new file returns or how to place the file without guessing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is only 25% and three parameters (name, mime_type, parent_folder_id) are undocumented in both schema and description. The description adds nothing about expected formats or defaults, so it fails to compensate for the coverage gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Create a new blank Drive file'), and the word 'blank' usefully distinguishes it from drive_upload_file and drive_write_file_content in the sibling set. It stops short of explicitly naming which alternative to prefer, but the scope is unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to use this versus drive_upload_file, drive_write_file_content, or drive_sheets_create. 'Auto-approved' hints at the approval workflow but never states the conditions that make this the right tool.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
drive_docs_edit_contentA
Replace one occurrence of existing text in a Google Doc with new Markdown, without touching the rest of the document. find_text must match exactly one location in the document's plain, unformatted text — the words as typed, with no Markdown syntax in them at all (this is not the same as drive_get_file_content's output for a Doc, which now renders formatting as Markdown; strip any '**'/'#'/etc. markers back out of find_text first) — include enough surrounding context to make it unique, the same way a unique-match text editor requires; set replace_all=true to replace every occurrence instead. replace_markdown supports the same Markdown syntax as drive_write_doc_content, including GFM pipe tables. Requires user approval.
| Name | Required | Description | Default |
|---|---|---|---|
| reason | Yes | One sentence: why are you calling this tool right now? | |
| file_id | Yes | ||
| find_text | Yes | Exact plain-text substring to locate | |
| replace_all | No | ||
| replace_markdown | Yes | Markdown to insert in its place |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Adds context beyond the annotations: the change is scoped to a single match so the rest of the document is untouched, and it requires user approval. Annotations already declare non-readOnly/non-destructive/non-idempotent, which is consistent. It does not however state error behavior when find_text matches zero or multiple locations, which matters for a non-idempotent mutator.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Dense but nearly every clause earns its place, front-loading the core action and scope before the find_text caveats. The nested parenthetical about drive_get_file_content is somewhat run-on but conveys a real trap.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a non-idempotent mutation with no output schema and 5 params, it covers the critical ambiguity (find_text matching) and the approval requirement. Remaining gaps are error/return behavior on zero or ambiguous matches, which a complete definition would ideally mention.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 60% schema coverage, the description usefully compensates: it clarifies find_text semantics (exact plain-text, no Markdown markers, must be unique) and defines replace_all=true's effect that the schema leaves undocumented. It adds the syntax scope for replace_markdown. The 'reason' parameter is left to the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: replace one occurrence of existing text in a Google Doc with new Markdown. It explicitly delineates scope ('without touching the rest of the document') and names the sibling tools it differs from (drive_get_file_content, drive_write_doc_content), so an agent can distinguish it from the other drive_docs_* and drive_write_doc_content tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives concrete conditions: find_text must uniquely match, include enough surrounding context, and set replace_all=true to replace every occurrence. It also warns about a pitfall relative to drive_get_file_content's Markdown-rendered output. It stops short of stating when NOT to use this tool (e.g., for wholesale rewrites vs drive_write_doc_content).
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
drive_docs_format_contentA
Apply formatting (bold, italic, highlight, text color) to existing text in a Google Doc, located the same way as drive_docs_edit_content, without changing the text itself. Every formatting parameter is opt-in — its default means 'leave that aspect unchanged', so a call that only sets highlight_color never touches bold/italic already on the matched text. Requires user approval.
| Name | Required | Description | Default |
|---|---|---|---|
| bold | No | 'true' or 'false'; omit to leave unchanged | |
| italic | No | 'true' or 'false'; omit to leave unchanged | |
| reason | Yes | One sentence: why are you calling this tool right now? | |
| file_id | Yes | ||
| find_text | Yes | Exact plain-text substring to locate | |
| text_color | No | hex color e.g. '#000000'; omit to leave unchanged | |
| replace_all | No | ||
| highlight_color | No | hex color e.g. '#fff59d'; omit to leave unchanged |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare non-read-only, non-idempotent, non-destructive. The description adds real behavioral context beyond them: the opt-in semantics where each default means 'leave unchanged', and the fact that a call requires user approval. This is exactly the kind of mutation/approval detail annotations cannot express.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, front-loaded with the action, then the scoping/opt-in caveat, then the approval requirement. Efficient, though the opt-in explanation is slightly verbose by restating the same 'leave unchanged' idea twice.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For an 8-parameter mutation tool with no output schema, the description covers purpose, scoping, the critical opt-in default behavior, and the approval requirement. It leaves replace_all semantics and what the return/result looks like unstated, but overall the agent has enough to invoke correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 75% and the schema already labels each color/boolean param as 'omit to leave unchanged'. The description generalizes this into an interaction rule – setting only highlight_color never touches bold/italic on the matched text – which is semantic value not present in the per-parameter schema text.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Names a specific verb and resource ('Apply formatting ... to existing text in a Google Doc') and explicitly scopes it against the sibling drive_docs_edit_content by noting the text itself is not changed. An agent can distinguish it from drive_docs_edit_content and drive_write_doc_content without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
States the locating mechanism by reference ('located the same way as drive_docs_edit_content') and flags the approval prerequisite, giving clear context for use. It stops short of an explicit when-not boundary (e.g. use edit_content when you also need to replace text), but the contrast with 'without changing the text' effectively implies it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
drive_download_fileARead-onlyIdempotent
Download a Drive file. Google Workspace documents are exported as text/CSV. On a local install: saved to destination_dir, and the saved file path is returned -- destination_dir is required, there is no default, so choose deliberately: pass ~/Downloads (or another path the user asked for) when this file is a deliverable the user should find afterward, or your own working/scratch directory when you're only downloading it to read or process it yourself. On an organization-managed install: destination_dir is ignored (there is no local filesystem you and the human share) -- a small file's bytes come back directly in this tool's result so you can read or hand it to the human yourself; a larger file comes back as a one-time link the human opens in their own signed-in browser tab instead. Requires user approval.
| Name | Required | Description | Default |
|---|---|---|---|
| reason | Yes | One sentence: why are you calling this tool right now? | |
| file_id | Yes | ||
| destination_dir | Yes | Local install only: where to save the file -- required, no default. On an organization-managed install it is ignored and nothing is saved to it; any value will do. Use ~/Downloads (or a path the user specified) if the user should find this file afterward; use your own working/scratch directory if it's only for you to read or process. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations cover only the safety profile (readOnly/idempotent/non-destructive); the description adds substantial context they don't: user approval is required, organization-managed installs ignore destination_dir, small files return inline bytes while large files return a one-time link, and Workspace docs are converted. This is exactly the beyond-annotations value the dimension rewards.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core action, then the two install-mode branches laid out in sequence. It is long, and the destination_dir guidance is repeated almost verbatim in the schema, so some trimming is possible, but every sentence carries information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description shoulders the return-value burden and does so: saved path, inline bytes, or one-time link depending on install mode. Combined with approval and export disclosure, an agent has everything needed to call it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 67% and the two parameters that matter are described in the schema itself; the description's destination_dir guidance largely paraphrases the schema text. The required `reason` parameter is never mentioned in the prose. Baseline 3 is appropriate when the schema already carries the parameter detail.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource (download a Drive file) and immediately qualifies it with the export behavior for Workspace docs (text/CSV), which is behavior an agent would not guess. An agent can distinguish this from the read-oriented siblings without opening a schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit contextual guidance: pick ~/Downloads when the file is a user deliverable, a scratch directory when it's only for the agent to read. It does not, however, position this against the obvious alternative drive_get_file_content, leaving that selection to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
drive_get_file_contentARead-onlyIdempotent
Fetch the content of a Drive file by id. A Google Doc comes back as Markdown (headings, bold, italic, strikethrough, underline, code, link, ==highlight==, '---' dividers, nested bullet/numbered lists with a 2-space indent per level, GFM pipe tables with alignment), the syntax drive_write_doc_content accepts, so it round-trips into it or drive_docs_edit_content. An exact highlight or text color Markdown can't carry comes back in 'highlights'/'text_colors' ({text, hex} lists). A Sheet comes back as CSV, Slides as plain text. A PDF, .docx, .pptx or .xlsx comes back as its extracted text (a scanned PDF has none), with 'truncated': true when cut to fit; a .zip as its list of entries. Other files, such as images, only get a placeholder: use drive_download_file for those. Requires user approval.
| Name | Required | Description | Default |
|---|---|---|---|
| reason | Yes | One sentence: why are you calling this tool right now? | |
| file_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already establish readOnly/idempotent/non-destructive, but the description adds substantive behavior beyond them: it requires user approval, marks output with 'truncated': true when cut to fit, notes a scanned PDF yields no text, and flags that exact highlight/text color is returned separately in 'highlights'/'text_colors'. This is unusually rich disclosure for a read tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The Markdown syntax enumeration is long, but it is front-loaded after the core purpose and every clause carries format/round-trip information an agent needs. It is dense rather than padded, though the inline example markup could be trimmed.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a complex multi-format reader with no output schema and only 2 params, the description fully covers what is returned per file type, truncation, edge cases (scanned PDFs, images), and the approval requirement. Nothing essential to calling it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 50% schema coverage, the baseline is 3. The description clarifies that file_id identifies the Drive file and that 'reason' is the approval justification implicitly, but it adds no syntax for file_id beyond 'by id'; the schema already documents the 'reason' parameter's one-sentence format.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Fetch the content of a Drive file by id') and immediately differentiates itself by output format per file type (Doc→Markdown, Sheet→CSV, Slides→text, PDF/docx→extracted text, zip→entries). An agent can distinguish it from the many sibling drive_* tools without opening a schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly routes the agent elsewhere when appropriate: 'Other files, such as images, only get a placeholder: use drive_download_file for those.' It also names the round-trip targets (drive_write_doc_content, drive_docs_edit_content) so the agent understands this is the read counterpart to those write tools.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
drive_get_file_metadataARead-onlyIdempotent
Fetch metadata for a single Drive file by id (name, owners, times, sharing status). Auto-approved.
| Name | Required | Description | Default |
|---|---|---|---|
| reason | Yes | One sentence: why are you calling this tool right now? | |
| file_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, and destructiveHint=false, so safety is covered. The description adds a non-annotation behavioral fact — 'Auto-approved' — which matters in a catalog containing privacyfence_await_approval, telling the agent no gate will block this call. It does not describe error behavior for a bad id, keeping it short of a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One front-loaded sentence naming the operation, key, and returned fields, followed by a short 'Auto-approved.' fragment that carries real approval-flow information. Efficient overall, though the trailing fragment is slightly abrupt rather than fully integrated.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple read-only tool with rich annotations and no output schema, the description is nearly sufficient: it enumerates the main metadata fields so the agent knows what comes back, and flags auto-approval. Missing only edge-case behavior (invalid id, permissions on shared files).
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is only 50%: 'reason' is documented in the schema, 'file_id' is not. The description partially compensates by specifying 'by id', clarifying what file_id means and that it targets one file, but adds no format or example beyond that. Baseline 3 for a low-coverage two-param schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Fetch), resource (metadata for a single Drive file), lookup key (by id), and enumerates what the metadata includes (name, owners, times, sharing status). It is clearly distinguishable from drive_list_files (plural, filtered listing) and drive_get_file_content (content, not metadata).
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The 'single file by id' framing implies you must already have the file id, and 'Auto-approved' hints no approval round-trip is needed, but there is no explicit when-to-use, when-not-to-use, or named alternative (e.g. drive_list_files for discovery). Usage is inferable but not spelled out.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
drive_list_filesBRead-onlyIdempotent
Search Google Drive and return matching file metadata (id, name, mime_type, owners, sharing status). Auto-approved.
| Name | Required | Description | Default |
|---|---|---|---|
| query | Yes | Drive API 'q' search query syntax, not a plain-text search string, e.g. "name contains 'Foo'" or "fullText contains 'Foo'". See https://developers.google.com/drive/api/guides/search-files | |
| reason | Yes | One sentence: why are you calling this tool right now? | |
| max_results | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, destructiveHint=false, so the safety profile is covered. The description adds the returned field list and the 'Auto-approved' note, but that note is ambiguous (approval-free execution? pre-vetted call?) and no pagination or result-limit behavior is disclosed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences with the action and the return payload front-loaded; no filler. 'Auto-approved' is terse but its meaning is murky, slightly weakening an otherwise efficient structure.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, enumerating the returned fields is valuable and correctly done. But for a search tool it omits result-limit/pagination behavior, the meaning of the required 'reason' field, and any routing to sibling Drive tools, leaving clear gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 67%; the query parameter is well documented in the schema itself (Drive 'q' syntax with examples and a doc link), so the description need not repeat it. However max_results has no description anywhere, and the description adds nothing about parameter semantics, so it sits at the baseline.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Search') and resource ('Google Drive files') and even enumerates the returned metadata fields (id, name, mime_type, owners, sharing status). It is distinguishable from drive_get_file_metadata and drive_list_folder by the 'search' framing, though it never explicitly contrasts itself with those siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives no when-to-use guidance, no prerequisites, and no pointer to alternatives such as drive_list_folder for browsing a known folder or drive_get_file_metadata for a known id. The agent must infer the selection from the name alone.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
drive_list_folderBRead-onlyIdempotent
List the direct children of a Drive folder by id. Auto-approved.
| Name | Required | Description | Default |
|---|---|---|---|
| reason | Yes | One sentence: why are you calling this tool right now? | |
| folder_id | Yes | ||
| max_results | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint and destructiveHint=false, so safety is covered. The description adds two genuinely useful behavioral facts — that it returns only direct children (not recursive) and that it is auto-approved — but says nothing about pagination, result ordering, or how max_results interacts with the returned set.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight, front-loaded sentences with no filler; the core behavior leads. 'Auto-approved' is terse to the point of being cryptic, but nothing here is wasted.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 3-parameter, no-output-schema read tool, the description gives the essential scope but omits pagination behavior and result format, which an agent needs to call it correctly when a folder has many children. Adequate but with clear gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is only 33%: folder_id and max_results have no descriptions at all. The description says the folder is identified 'by id' and clarifies the output is direct children, but leaves max_results (default 50) and the expected id format entirely unexplained, so it does not compensate for the coverage gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (List), resource (Drive folder children), scope (direct children, by id). An agent can tell it apart from drive_list_shared_drives, but the distinction from the very similar drive_list_files is not addressed, which is the sibling it most overlaps with.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no when-to-use guidance and no comparison to drive_list_files or drive_get_file_metadata, which are the natural alternatives for navigating Drive. 'Auto-approved' gestures at approval flow but does not tell the agent when this tool is the right choice.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
drive_move_fileA
Move a Drive file to a different folder. Requires user approval.
| Name | Required | Description | Default |
|---|---|---|---|
| reason | Yes | One sentence: why are you calling this tool right now? | |
| file_id | Yes | ||
| destination_folder_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=false, idempotentHint=false, and destructiveHint=false, and the description adds a genuinely useful trait not in the structured data: the call requires user approval. It still omits what happens to sharing/permissions after the move and why the operation is not idempotent, which keeps it below 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short, front-loaded sentences with no waste; the core action leads and the approval constraint follows. Nothing is padded or redundant.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 3-parameter mutation with no output schema and only one documented parameter, the description covers the action and the approval gate but leaves the parameter formats and post-move side effects unspecified. Adequate but with visible gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is only 33% (only 'reason' is documented), so the description must compensate. 'Move a Drive file to a different folder' implicitly maps file_id to the source and destination_folder_id to the target, but adds no format detail (ID vs. name, shared-drive restrictions), leaving half the compensation undone.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Move a Drive file to a different folder'), so the agent knows exactly what the tool does. It does not, however, differentiate itself from the many other drive_* siblings or clarify that this is the only move/parent-change operation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no explicit when-to-use or when-not-to-use guidance and no alternatives named, even though the sibling list is packed with drive tools (drive_list_folder, drive_get_file_metadata). 'Requires user approval' is an operational precondition, not usage guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
drive_sheets_add_sheetC
Add a new tab to an existing spreadsheet. Requires user approval.
| Name | Required | Description | Default |
|---|---|---|---|
| cols | No | ||
| rows | No | ||
| title | Yes | ||
| reason | Yes | One sentence: why are you calling this tool right now? | |
| spreadsheet_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=false, destructiveHint=false and idempotentHint=false, so the write/non-destructive profile is covered. The description adds genuinely useful context by disclosing the user-approval requirement, but says nothing about duplicate-title behavior or mutation side effects beyond that.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short, front-loaded sentences with no filler; the core action leads and the approval constraint follows. Efficient, though the brevity contributes to the documentation gaps noted elsewhere.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a mutation tool with no output schema and only 20% parameter coverage, the description should explain the required title/spreadsheet_id semantics and any uniqueness constraints. Only the approval requirement is covered, leaving most agent-relevant detail missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is only 20% – just 'reason' is documented. spreadsheet_id, title, rows and cols carry no explanation in either schema or description, and the description adds nothing (e.g. whether title must be unique, what rows/cols defaults mean). The required-vs-optional distinction is left entirely to the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('add a new tab to an existing spreadsheet'), which clearly distinguishes it from drive_sheets_create (new spreadsheet) and drive_sheets_rename_sheet. It does not explicitly name those siblings, but the action is unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The only usage signal is 'Requires user approval', which is a prerequisite rather than guidance. It never says when to add a sheet versus renaming or creating one, nor what happens if the sheet already exists.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
drive_sheets_createB
Create a new Google Sheets spreadsheet, optionally with named tabs. Auto-approved.
| Name | Required | Description | Default |
|---|---|---|---|
| name | Yes | ||
| reason | Yes | One sentence: why are you calling this tool right now? | |
| sheet_titles | No | JSON array of tab names, e.g. ["Q1","Q2"]. Defaults to a single 'Sheet1' tab if omitted. | |
| parent_folder_id | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare this is a non-read-only, non-idempotent, non-destructive write, so the safety profile is covered. 'Auto-approved' adds genuinely useful context about the approval path, but the description omits return details, ownership/location defaults, and quota behavior for a creation tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two compact sentences, no filler, with the core action front-loaded. Every clause carries information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema, so the description needn't explain return values, but for a creation tool it should indicate where the new spreadsheet is created by default and how the required reason parameter fits in. The approval note helps but the picture is incomplete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 50%: sheet_titles and parent_folder_id are documented in the schema, while name and the required reason are not. The description adds the concept of 'named tabs' (mapping to sheet_titles) but says nothing about where the file lands or what 'reason' is for. Partial compensation only.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: create a Google Sheets spreadsheet, with the optional tab-titling capability named. It is clear enough to separate from sibling reads/writes like drive_sheets_write_range or drive_create_blank_file, though it doesn't explicitly contrast with the generic blank-file creator.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives no when-to-use guidance or alternatives, and nothing about prerequisites such as folder placement or required Drive scopes. 'Auto-approved' hints at an approval flow but does not tell the agent when this tool is the right choice versus drive_create_blank_file.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
drive_sheets_delete_dimensionsADestructive
Delete rows or columns from a sheet tab, including any values, formulas, and formatting they contain. This is destructive — deleted cell content is not recoverable through PrivacyFence. Remaining rows/columns shift to close the gap. Requires user approval.
| Name | Required | Description | Default |
|---|---|---|---|
| count | No | ||
| reason | Yes | One sentence: why are you calling this tool right now? | |
| sheet_id | Yes | Numeric tab id, from drive_sheets_get_metadata | |
| dimension | Yes | 'ROWS' or 'COLUMNS' | |
| start_index | Yes | 0-based, inclusive of the first row/column removed | |
| spreadsheet_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true and idempotentHint=false, but the description goes well beyond them: it states the deleted content is not recoverable through PrivacyFence, that remaining rows/columns shift to close the gap, and that user approval is required. These are exactly the behavioral facts an agent needs before invoking a destructive mutation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, all front-loaded: what is deleted, that it is irreversible, and the side effect on remaining rows. Zero filler and no repetition of structured fields.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a destructive mutation with no output schema, the description covers the critical unknowns (irreversibility, gap-closing shift, approval gate). It omits any note on what the call returns or what happens on out-of-range indices, which is a minor residual gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 67% and the description adds no parameter-level meaning — it never explains that start_index is 0-based inclusive, how count interacts with the range, or what dimension accepts. With most params documented in the schema, the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (delete) and resource (rows or columns from a sheet tab), with the scope of what is removed ('including any values, formulas, and formatting'). An agent can distinguish this from drive_sheets_insert_dimensions without opening either schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description conveys that this is a destructive operation requiring user approval, which implies caution before use, but it never names the alternative (insert_dimensions) or states the condition that should select delete over other sheet-editing tools. Usage is implied rather than guided.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
drive_sheets_format_rangeA
Apply formatting to a range in a spreadsheet: bold/italic, colors, number format, horizontal/vertical alignment, text wrap, column width, frozen rows/columns, and merged cells. Every parameter is opt-in — its default means 'leave that aspect unchanged', so a call that only sets a background color never touches unrelated formatting already on the range. Requires user approval.
| Name | Required | Description | Default |
|---|---|---|---|
| bold | No | 'true' or 'false'; omit to leave unchanged | |
| italic | No | 'true' or 'false'; omit to leave unchanged | |
| reason | Yes | One sentence: why are you calling this tool right now? | |
| range_a1 | Yes | Plain A1 range scoped to sheet_id, e.g. 'A1:C10' (no sheet-name prefix, must be fully bounded) | |
| sheet_id | Yes | Numeric tab id, from drive_sheets_get_metadata | |
| merge_type | No | KEEP (default) / NONE (unmerge) / MERGE_ALL / MERGE_COLUMNS / MERGE_ROWS | KEEP |
| text_color | No | hex color e.g. '#ffffff'; omit to leave unchanged | |
| freeze_cols | No | Number of columns to freeze at the left (0 unfreezes); omit (-1) to leave unchanged | |
| freeze_rows | No | Number of rows to freeze at the top (0 unfreezes); omit (-1) to leave unchanged | |
| column_width | No | Pixel width for the range's columns; omit (-1) to leave unchanged | |
| number_format | No | Sheets number-format pattern, e.g. '0.00%', '$#,##0.00', 'yyyy-mm-dd'; omit to leave unchanged | |
| wrap_strategy | No | 'OVERFLOW_CELL' (overflow into empty neighboring cells) / 'CLIP' (cut off at the cell boundary) / 'WRAP' (line break to fit the cell); omit to leave unchanged | |
| spreadsheet_id | Yes | ||
| background_color | No | hex color e.g. '#ffcc00'; omit to leave unchanged | |
| vertical_alignment | No | 'TOP' / 'MIDDLE' / 'BOTTOM'; omit to leave unchanged | |
| horizontal_alignment | No | 'LEFT' / 'CENTER' / 'RIGHT'; omit to leave unchanged |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=false, destructiveHint=false, idempotentHint=false, so the mutation/safety profile is known. The description adds genuine context beyond that: partial-update semantics (unset parameters never touch unrelated existing formatting) and an approval requirement. It stops short of describing merge/unmerge reversibility or how the required 'reason' approval flow works.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, front-loaded with the purpose, then the crucial opt-in semantics, then the approval constraint. No filler; every sentence carries information an agent needs for 16 parameters.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 16-parameter mutation tool with no output schema, the description covers purpose, the non-obvious opt-in default behavior, and the approval gate, which is close to sufficient. Minor gaps remain around merge/unmerge specifics and how freeze settings interact with the target range.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 94%, so almost every parameter is already documented in the schema (including the 'omit to leave unchanged' note per field). The description's list of formatting facets is a helpful overview but adds no syntax or value-range detail beyond what the schema already provides, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (apply formatting) and resource (range in a spreadsheet), then enumerates exactly which formatting aspects are covered — bold/italic, colors, number format, alignment, wrap, width, freeze, merge. This distinguishes it clearly from siblings like drive_sheets_write_range (values) and drive_docs_format_content (docs).
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives real guidance on how to call it (each parameter is opt-in; defaults leave the aspect unchanged), which is useful context for invocation. However, it never states when to choose this tool over alternatives such as drive_sheets_write_range or drive_sheets_get_metadata, so tool-selection guidance is only implied.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
drive_sheets_get_metadataBRead-onlyIdempotent
List the tabs in a spreadsheet (id, title, index, row/column count). Auto-approved.
| Name | Required | Description | Default |
|---|---|---|---|
| reason | Yes | One sentence: why are you calling this tool right now? | |
| spreadsheet_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, and destructiveHint=false. The description adds 'Auto-approved,' which is useful approval context beyond annotations, but it does not disclose access requirements, rate limits, or what happens if the spreadsheet is inaccessible.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences, front-loaded with the core action and return fields; 'Auto-approved' is brief and earns its place as approval context. It is slightly cryptic but has little waste.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple read-only metadata tool whose annotations cover safety, the description states the returned fields, but it omits usage guidance, the required reason parameter's role, and any access constraints. Adequate but with clear gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 50%: reason is documented, spreadsheet_id is not. The description adds no syntax, format, or meaning beyond 'in a spreadsheet,' so it does not compensate for the underexplained parameter.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Lists the tabs in a spreadsheet and names the returned fields (id, title, index, row/column count), which is a specific verb+resource. It does not explicitly contrast with sibling read tools like drive_sheets_get_values or drive_get_file_metadata, so it falls short of 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No when-to-use guidance or alternatives are given. 'Auto-approved' is an approval attribute, not usage direction, and the description never says when to choose this over drive_sheets_get_values or drive_get_file_metadata.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
drive_sheets_get_valuesARead-onlyIdempotent
Read a range of cells from a spreadsheet: display values by default, or the underlying values, or formulas instead of computed results, and optionally cell formatting alongside them. Requires user approval.
| Name | Required | Description | Default |
|---|---|---|---|
| reason | Yes | One sentence: why are you calling this tool right now? | |
| range_a1 | Yes | A1 notation range, e.g. 'Sheet1!A1:C10' | |
| spreadsheet_id | Yes | ||
| include_formatting | No | Also fetch per-cell formatting for the range (bold/italic/text color/background color/number format/horizontal/vertical alignment/text wrap -- the same aspects drive_sheets_format_range can set). | |
| value_render_option | No | FORMATTED_VALUE (default): displayed strings, e.g. '$1.00'. UNFORMATTED_VALUE: the underlying value with no formatting, e.g. 1. FORMULA: the formula text itself, e.g. '=A1+A2', instead of its computed result -- use this to read formulas. | FORMATTED_VALUE |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint and destructiveHint, lowering the bar. The description nonetheless adds a meaningful behavioral fact not in the annotations — 'Requires user approval' — plus the effect of include_formatting on the returned payload, which is real value.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One dense sentence that front-loads the core action and the render-mode variation; nothing is wasted, though the nested 'display values... or formulas... and optionally formatting' phrasing is slightly heavy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema, but the description communicates what comes back (display values, underlying values, formulas, optional formatting) and flags the approval requirement. Combined with annotations that carry the safety profile, an agent has enough to call it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 80%, so the schema already documents spreadsheet_id, range_a1, reason, include_formatting and value_render_option in detail. The description largely restates the render options without adding syntax or constraints beyond the schema, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('read a range of cells from a spreadsheet') and enumerates the three render modes plus optional formatting, so an agent knows exactly what the tool returns. It does not explicitly differentiate itself from siblings like drive_sheets_get_metadata or drive_sheets_write_range, which caps it below a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied through the value_render_option guidance ('use this to read formulas') and the note that approval is required, but there is no explicit when-to-use-this-vs-alternatives statement relative to drive_sheets_get_metadata or the write tools.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
drive_sheets_insert_dimensionsA
Insert blank rows or columns into a sheet tab, shifting existing content after the insertion point. Values/formulas are untouched, only their position shifts; formulas referencing shifted cells are adjusted automatically. Requires user approval.
| Name | Required | Description | Default |
|---|---|---|---|
| count | No | ||
| reason | Yes | One sentence: why are you calling this tool right now? | |
| sheet_id | Yes | Numeric tab id, from drive_sheets_get_metadata | |
| dimension | Yes | 'ROWS' or 'COLUMNS' | |
| start_index | Yes | 0-based index to insert before | |
| spreadsheet_id | Yes | ||
| inherit_from_before | No | Copy formatting from the row/column before the insertion point (Sheets UI default) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare the mutation/safety profile (readOnlyHint=false, destructiveHint=false), and the description usefully adds that values and formulas are untouched while formulas referencing shifted cells are auto-adjusted, plus a user-approval requirement. These are non-obvious behavioral facts beyond the annotations, though the adjustment/approval mechanics are not detailed further.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight sentences: core action first, then behavioral consequences, then the approval constraint. Each sentence carries information with little waste.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 7-parameter mutation tool with no output schema, the description covers the essential semantics (positional shift, formula adjustment, approval). Remaining parameter details live in the schema, so an agent has enough to call it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 71%, so most but not all parameters are self-documented. The description's 'shifting existing content after the insertion point' loosely conveys the start_index semantics, but it adds nothing about count, dimension values, or inherit_from_before beyond what the schema already says. Baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (insert) and resource (blank rows or columns into a sheet tab) with the exact effect on existing content. An agent can immediately distinguish this from the sibling drive_sheets_delete_dimensions by the insert vs. delete action.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The clear purpose implies the usage scenario (need blank rows/columns mid-sheet), but the description never explicitly states when to prefer this over alternatives, nor names delete_dimensions as the inverse operation or add_sheet as a different approach. Usage is inferred rather than guided.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
drive_sheets_rename_sheetA
Rename an existing tab in a spreadsheet. There is no delete-sheet tool — to mark a tab for removal, rename it (e.g. to 'TO BE DELETED - ') and the user can delete it by hand in the Sheets UI. Requires user approval.
| Name | Required | Description | Default |
|---|---|---|---|
| reason | Yes | One sentence: why are you calling this tool right now? | |
| sheet_id | Yes | Numeric tab id, from drive_sheets_get_metadata | |
| new_title | Yes | ||
| spreadsheet_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=false, destructiveHint=false, and idempotentHint=false, so the mutation profile is covered. The description adds genuinely new context: 'Requires user approval' and the no-delete-tool workaround, which an agent cannot derive from the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, front-loaded with the core action, then the workaround, then the approval requirement. No filler and each sentence carries distinct information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a four-parameter mutation with no output schema, the description supplies the approval requirement and the deletion workaround, which are the key operational facts. A note on what happens if the title already exists would complete it, but nothing essential is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 50%: 'reason' and 'sheet_id' are documented in the schema, while 'spreadsheet_id' and 'new_title' are not. The description implies the new_title semantics via 'rename' but adds no format or constraint detail, so it only marginally compensates for the gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Rename an existing tab in a spreadsheet'), which clearly separates it from drive_sheets_add_sheet and the other sheet tools in the sibling list. The scope (existing tab) is unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly covers the deletion workaround — 'There is no delete-sheet tool — to mark a tab for removal, rename it... and the user can delete it by hand' — which tells the agent when to reach for this tool beyond a simple rename. It doesn't enumerate exclusion conditions for ordinary renames, but the routing guidance is concrete.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
drive_sheets_write_rangeA
Write values and/or formulas into a range of an existing spreadsheet. A cell string starting with '=' is evaluated as a formula, exactly as if typed into the Sheets UI — there is no separate tool for formulas. Writing an empty row/column clears those cells. Requires user approval.
| Name | Required | Description | Default |
|---|---|---|---|
| reason | Yes | One sentence: why are you calling this tool right now? | |
| values | Yes | JSON 2D array of rows, e.g. [["Name","Total"],["Alice","=B2*2"]] | |
| range_a1 | Yes | A1 notation range, e.g. 'Sheet1!A1:C10' | |
| spreadsheet_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare the mutation profile (readOnly=false, idempotent=false, destructive=false), and the description adds real behavior beyond them: how '=' strings are evaluated as formulas and that writing empty rows/columns clears those cells. The approval requirement is also surfaced. It does not describe return values, but no output schema exists to replace that.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four tightly scoped sentences, front-loaded with the core action and followed by the two most surprising behaviors (formula evaluation, clearing on empty) and the approval requirement. No filler sentences.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a non-idempotent write tool with no output schema, the description covers the salient semantics an agent needs: formula handling, overwrite/clearing effects, and approval. Remaining gaps (rate limits, error behavior) are minor and atypical to document here.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 75%, so most params are documented structurally, but the description adds meaning the schema lacks: the formula-evaluation semantics of cell strings, which is not conveyed by the values example. It doesn't mention the 'reason' parameter, a minor gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource: writing values/formulas into a range of an existing spreadsheet. It clearly contrasts with reading (drive_sheets_get_values) by scoping to 'existing spreadsheet', but does not explicitly name siblings such as drive_sheets_format_range or drive_write_file_content.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The '=' formula note effectively routes agents away from hunting for a separate formula tool, and 'Requires user approval' hints at the invocation context. However, there is no explicit when-to-use guidance versus siblings like drive_sheets_format_range or drive_write_file_content.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
drive_upload_fileA
Upload any file (e.g. a PDF or image) to Drive as a new file — use this instead of drive_write_file_content for any binary file, since that tool only writes UTF-8 text. Provide exactly one of local_path (a path on the user's computer — where Claude Desktop runs: absolute, or starting with ~/. Claude's own working or outputs directory is fine), content_base64 (base64-encoded file bytes, decoded by PrivacyFence itself — use this when you only have the file's bytes and not a local path; 'name' is then required), or upload_id (the id privacyfence_create_upload_slot returned after you PUT the file's bytes to its upload_url — use this if local_path fails with an error about PrivacyFence being unable to read files in your home folder directly, e.g. no PrivacyFence extension is installed). On an organization-managed install, local_path is read from wherever PrivacyFence's own server runs, not the user's machine — prefer content_base64 or upload_id there. Requires user approval.
| Name | Required | Description | Default |
|---|---|---|---|
| name | No | ||
| reason | Yes | One sentence: why are you calling this tool right now? | |
| upload_id | No | ||
| local_path | No | ||
| content_base64 | No | ||
| parent_folder_id | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations cover the safety profile (readOnly=false, idempotent=false, destructive=false), and the description adds context they cannot: 'Requires user approval', the fact that content_base64 is decoded by PrivacyFence itself, and that on org-managed installs local_path resolves on the server not the user's machine.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The purpose and the sibling comparison are front-loaded, and nearly every clause carries operational information. It is nonetheless a single dense block with stacked parentheticals that could be broken into shorter statements for faster scanning.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 6-param mutation tool with no output schema and thin schema coverage, the description covers the central usage paths, approval requirement, and install-mode caveat well. It stops short of documenting parent_folder_id, so an agent has no guidance on target-location behavior.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is only 17% (only 'reason' is documented), so the description must carry the load — and it does for three of six params, explaining exactly-one-of semantics, that 'name' is required with content_base64, and what upload_id refers to. parent_folder_id and the full meaning of name are never mentioned, leaving a real gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Upload any file ... to Drive as a new file') and immediately distinguishes itself from the sibling drive_write_file_content by scoping to binary vs UTF-8 text. An agent can select this over its nearest alternative without opening either schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicit routing: 'use this instead of drive_write_file_content for any binary file', plus a decision procedure for the three input modes (local_path, content_base64, upload_id) with the conditions that favor each. It even gives a fallback path when local_path errors and an org-managed-install caveat.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
drive_write_doc_contentA
Write Markdown content to a Google Doc with rich formatting: headings (# through ######), bold, italic, bold-italic, strikethrough, underline, code, ==highlight== (these five nest freely with each other, e.g. ==bold and highlighted==), link (escape a literal '[' or ']' in the link text as '['/']'), bullet/numbered lists (indent a sub-list 2 spaces per nesting level), GFM pipe tables (a '| --- |' separator row under the header; ':---'/'---:'/':---:' for left/right/center column alignment), and '---'/'***'/'___' on their own line as a horizontal-rule divider. Clears the existing document content before writing — use drive_docs_edit_content or drive_docs_format_content instead for a change that shouldn't touch the rest of the document. Use this instead of drive_write_file_content when the target is a Google Doc and you want formatted output. Requires user approval.
| Name | Required | Description | Default |
|---|---|---|---|
| reason | Yes | One sentence: why are you calling this tool right now? | |
| file_id | Yes | ||
| markdown | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Discloses the critical destructive side effect ('Clears the existing document content before writing') and an operational gate ('Requires user approval') that the annotations do not convey. It also specifies the exact supported Markdown grammar, which is the behavior an agent must satisfy. The only tension is destructiveHint=false, but that hint concerns deleting the resource, and here the document survives with its contents replaced, which the description states plainly rather than concealing.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Long but dense: the formatting grammar is front-loaded and every clause is functional rather than filler. It is a single packed sentence followed by routing and approval notes, so it reads well even if it could be broken into shorter segments.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema and three required parameters, the description covers the destructive behavior, the approval gate, sibling alternatives, and the input format thoroughly. The only omission is what the call returns (e.g., revision/status), which is a minor gap for an agent that mainly needs to know its write will replace the document.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is only 33% (only 'reason' is documented) and the description never mentions 'file_id'. However, it extensively defines the content contract for the 'markdown' parameter — the single most error-prone input — including escaping rules, list indentation depth, and table separator syntax, which more than compensates for the missing file_id note (self-evident as the target doc id).
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a precise verb+resource combination ('Write Markdown content to a Google Doc') and immediately scopes the output format. It also distinguishes itself from three named siblings (drive_write_file_content, drive_docs_edit_content, drive_docs_format_content), so an agent can route correctly without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly gives the when-to-use case ('when the target is a Google Doc and you want formatted output') and the when-not case ('use drive_docs_edit_content or drive_docs_format_content instead for a change that shouldn't touch the rest of the document'). Alternatives are named with the condition that selects each, which is the strongest form of guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
drive_write_file_contentB
Write content to an existing Drive file. Requires user approval.
| Name | Required | Description | Default |
|---|---|---|---|
| reason | Yes | One sentence: why are you calling this tool right now? | |
| content | Yes | ||
| file_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare readOnlyHint=false, idempotentHint=false, destructiveHint=false, so safety profile is partly covered. The description adds a real behavioral fact not in annotations: user approval is required. But it omits what happens to existing content (overwrite vs append), which matters for a non-idempotent write.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences, front-loaded with the action and followed by the key precondition. No filler, though the second sentence could be integrated more informatively.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a non-idempotent mutation with no output schema, the agent lacks critical details: overwrite semantics for 'content', accepted file types, and how the approval flow is triggered. The approval note is valuable but the description is not complete for a writer tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 33%; only 'reason' is documented in-schema. The description adds nothing about 'content' (format? full replacement?) or 'file_id' (must already exist – inferred from 'existing'). Baseline 3 given the modest coverage that the description fails to supplement.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (write) and resource (content to an existing Drive file), which is clear. However, it does not differentiate from close siblings like drive_write_doc_content, drive_docs_edit_content, apps_script_write_content, or drive_upload_file, leaving the agent to infer which file types this applies to.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No explicit when-to-use guidance or alternatives named. The sibling set contains several overlapping write tools (drive_write_doc_content, drive_docs_edit_content, apps_script_write_content) and the description gives no routing cues. 'Requires user approval' hints at a precondition but is not framed as usage guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
gmail_add_labelB
Add a label to a Gmail message. Requires user approval.
| Name | Required | Description | Default |
|---|---|---|---|
| reason | Yes | One sentence: why are you calling this tool right now? | |
| label_name | Yes | ||
| message_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare this is a non-read-only, non-idempotent, non-destructive mutation. The description adds one genuinely new behavioral fact beyond the annotations: an approval gate is required before execution. It says nothing about authentication scope, what happens if the label does not exist, or whether the operation is reversible.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences, front-loaded with the action and followed by the critical constraint. No filler or redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema and a 33% schema coverage, the description is thin for a mutation tool. It conveys the core action and the approval gate, which is the minimum viable, but omits error behavior and label-existence semantics.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is only 33% (only 'reason' is documented), so the description must compensate. It nominally implies a label and a message as inputs, but adds no meaning about label_name format (must the label pre-exist?), message_id format, or how the three parameters interact.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb (add) and resource (a label to a Gmail message), which is unambiguous. It does not, however, differentiate itself from close siblings like gmail_remove_label or gmail_create_label, so the agent must infer the distinction from tool names alone.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
'Requires user approval' is a prerequisite note, not usage guidance. There is no statement of when to use this versus gmail_remove_label, gmail_create_label, or gmail_archive_message, and no when-not condition.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
gmail_archive_messageA
Archive a Gmail message by removing it from the Inbox. The message is not deleted and remains searchable. Requires user approval.
| Name | Required | Description | Default |
|---|---|---|---|
| reason | Yes | One sentence: why are you calling this tool right now? | |
| message_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=false and readOnlyHint=false, and the description reinforces this by stating the message is 'not deleted and remains searchable' — valuable reassurance beyond the structured hints. It also surfaces an approval requirement an agent must plan around, though it omits idempotency behavior (annotations mark it non-idempotent without explanation).
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, front-loaded with the action and scope, then the non-destructive clarification, then the approval constraint. Every sentence carries distinct information with no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a two-parameter mutation tool with no output schema, the description covers what happens to the message and the approval gate. Minor gaps remain around idempotency (re-archiving an already-archived message) and whether the action is reversible, but nothing essential for a correct call is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 50%: 'reason' is fully documented in the schema, while 'message_id' is undocumented but self-evident from its name. The description adds no parameter-level detail, so it neither compensates nor detracts; baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (archive) and resource (Gmail message) plus the mechanism (removing it from the Inbox). No sibling tool overlaps with archiving, so an agent can identify it immediately. The clarifying clause distinguishes it from deletion in the same breath.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description adds the constraint 'Requires user approval,' which is real usage context. However, it never says when to archive versus alternatives like gmail_add_label/gmail_remove_label or leaving the message in place, so routing guidance is only implied.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
gmail_create_draftC
Create a Gmail draft. Requires user approval.
| Name | Required | Description | Default |
|---|---|---|---|
| cc | No | ||
| to | Yes | ||
| bcc | No | ||
| body | No | Plain-text body. May be omitted if body_markdown is given -- the plain-text alternative is then auto-derived from it. At least one of body/body_markdown is required. | |
| reason | Yes | One sentence: why are you calling this tool right now? | |
| send_as | No | Send from this Gmail send-as address instead of the account's default (sets From:, and the signature used is this address's own). Must be one of the account's configured send-as addresses -- anything else is rejected. | |
| subject | Yes | ||
| body_markdown | No | Optional Markdown body for a rich-text draft. Supports **bold**, *italic*, ==highlight==, [links](url), bullet/numbered lists, and `# Heading 1`/`## Heading 2` (rendered as Gmail's Large/Huge font-size presets, not raw heading tags -- no tables). When given, the draft is sent as plain text + HTML together, so it renders formatted in HTML-capable clients and as readable plain text everywhere else. | |
| include_signature | No | Append the user's Gmail signature (the one Gmail stores for the sending address) to the end of the body. Omit to use the user's 'Append Gmail signature to drafts' setting. Don't also write a sign-off block of your own when this is on. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare the write profile (readOnlyHint=false, idempotentHint=false, destructiveHint=false), so the safety baseline is covered. The description adds one genuinely useful trait beyond them: the operation requires user approval. It stops there, saying nothing about what happens on rejection or whether anything is sent.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short, front-loaded sentences with no filler; the purpose leads and the constraint follows. It is under-specified rather than padded, so the low information density is a completeness problem, not a conciseness one.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 9-parameter mutation tool with no output schema and non-trivial semantics (send-as validation, markdown-to-HTML derivation, signature handling), two sentences are insufficient. Nothing tells the agent what the call returns or how the approval step is triggered, leaving the structured fields to do nearly all the work.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 9 parameters and only 56% schema description coverage, the description carries real burden and adds nothing: it mentions no parameter, including the required 'reason' field, whose connection to the approval workflow goes unexplained. The schema documents body/body_markdown/send_as/include_signature well, but to/cc/bcc/subject are bare.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Create a Gmail draft'), so the agent knows exactly what it does. However, it does not differentiate from close siblings like gmail_reply_draft, gmail_reply_all_draft, or gmail_create_draft_with_attachments, so selection among those is left to the name alone.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives no when-to-use guidance or alternatives. It never says to use this for a brand-new draft rather than gmail_reply_draft for a response, nor when to prefer the with_attachments variant. 'Requires user approval' is a behavioral note, not usage guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
gmail_create_draft_with_attachmentsA
Create a Gmail draft with one or more local-file attachments. Parallel to gmail_create_draft -- use this variant only when there is something to attach; use gmail_create_draft when there isn't, so a draft doesn't need this tool's extra attachments argument for nothing. Requires user approval.
| Name | Required | Description | Default |
|---|---|---|---|
| cc | No | ||
| to | Yes | ||
| bcc | No | ||
| body | No | Plain-text body. May be omitted if body_markdown is given -- the plain-text alternative is then auto-derived from it. At least one of body/body_markdown is required. | |
| reason | Yes | One sentence: why are you calling this tool right now? | |
| send_as | No | Send from this Gmail send-as address instead of the account's default (sets From:, and the signature used is this address's own). Must be one of the account's configured send-as addresses -- anything else is rejected. | |
| subject | Yes | ||
| attachments | Yes | JSON array of local file paths to attach, e.g. ["/path/to/report.pdf"] -- each either a path on the user's computer (where Claude Desktop runs: absolute, or starting with ~/. Claude's own working or outputs directory is fine) or 'upload:<upload_id>', the id privacyfence_create_upload_slot returned after you PUT the file's bytes to its upload_url -- use that form if a plain path fails with an error about PrivacyFence being unable to read files in your home folder directly (e.g. no PrivacyFence extension is installed, or this is an organization-managed install: a plain path there is read from wherever PrivacyFence's own server runs, not the user's machine -- prefer 'upload:<upload_id>'). At least one required. | |
| body_markdown | No | Optional Markdown body for a rich-text draft. Supports **bold**, *italic*, ==highlight==, [links](url), bullet/numbered lists, and `# Heading 1`/`## Heading 2` (rendered as Gmail's Large/Huge font-size presets, not raw heading tags -- no tables). When given, the draft is sent as plain text + HTML together, so it renders formatted in HTML-capable clients and as readable plain text everywhere else. | |
| include_signature | No | Append the user's Gmail signature (the one Gmail stores for the sending address) to the end of the body. Omit to use the user's 'Append Gmail signature to drafts' setting. Don't also write a sign-off block of your own when this is on. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only declare readOnlyHint=false, idempotentHint=false, destructiveHint=false, which leave the approval/auth story unclear. The description adds 'Requires user approval', a meaningful behavioral fact the agent needs before calling. It does not, however, describe failure modes for attachment reading or how the approval is triggered, so it falls short of a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, front-loaded with the core purpose, then routing, then the approval caveat. Minor redundancy in 'so a draft doesn't need this tool's extra attachments argument for nothing' slightly dulls the efficiency.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 10-parameter mutation tool with no output schema, the description covers purpose, routing, and the approval requirement, and the schema carries the per-parameter detail. It doesn't address what a successful call returns (draft id/link) or attachment-read failures, leaving a small gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 60%, with the heavy parameters (attachments, body_markdown, send_as, include_signature) documented thoroughly in the schema itself. The description only gestures at the 'extra attachments argument' without adding format or constraint detail beyond the schema, so baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource+differentiator: creating a Gmail draft with local-file attachments. It explicitly contrasts with gmail_create_draft, so the agent can distinguish this tool from its closest sibling without reading either schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit routing: use this variant only when there is something to attach, use gmail_create_draft when there isn't. This is a clean when/when-not rule naming the alternative.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
gmail_create_filterA
Create a Gmail filter. Provide at least one criteria field (from_address, to_address, subject, query, has_attachment) and at least one action (add_label_names, archive, mark_as_read, star, forward_to). Requires user approval.
| Name | Required | Description | Default |
|---|---|---|---|
| star | No | ||
| query | No | Gmail search syntax; matches the filter's 'Has the words' field | |
| reason | Yes | One sentence: why are you calling this tool right now? | |
| archive | No | Skip the Inbox | |
| subject | No | ||
| forward_to | No | ||
| to_address | No | ||
| from_address | No | ||
| mark_as_read | No | ||
| has_attachment | No | ||
| add_label_names | No | Comma-separated label names to apply; created if missing |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare the safety profile (readOnlyHint=false, idempotentHint=false, destructiveHint=false), and the description adds material context: 'Requires user approval,' which the annotations do not convey. It does not disclose duplicate-filter handling or what happens on re-invocation, so it is not fully transparent.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences, front-loaded with the verb and resource, followed by the constraint and the approval caveat. No filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For an 11-parameter mutation tool with no output schema, the description covers the essential construction rule, the field taxonomy, and the approval requirement. It could be more complete on the required reason parameter and on idempotency/duplicate behavior, but nothing critical is missing for a correct call.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is only 36%, so the description must compensate, and it does so by grouping the 11 parameters into criteria fields (from_address, to_address, subject, query, has_attachment) and actions (add_label_names, archive, mark_as_read, star, forward_to) — semantic meaning the schema does not provide. It omits mention of the required 'reason' parameter, leaving one gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Create a Gmail filter') that is unambiguous. It does not, however, distinguish itself from the sibling tools gmail_update_filter or gmail_list_filters, so the agent must infer which filter operation applies from context alone.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives a concrete construction rule (at least one criteria field and at least one action), which is genuinely useful guidance. It stops short of saying when to create a new filter versus updating an existing one via gmail_update_filter, and lists no exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
gmail_create_labelA
Create a Gmail label. Use '/' to create nested labels (e.g. 'Work/Projects' creates 'Projects' nested under 'Work', creating 'Work' first if it doesn't already exist). Fails if the exact label name already exists. Requires user approval.
| Name | Required | Description | Default |
|---|---|---|---|
| reason | Yes | One sentence: why are you calling this tool right now? | |
| label_name | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare non-readOnly and non-idempotent; the description corroborates with 'Fails if the exact label name already exists' and adds meaningful context beyond annotations: nested-label creation side effects ('creating Work first if it doesn't exist') and the required user approval gate.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tightly packed sentences, front-loaded with the core action and followed by the nesting rule and the failure/approval conditions. No filler or restatement of the tool name.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a mutation tool with no output schema and thin parameter documentation, the description covers the critical behaviors: nesting semantics, duplicate-name failure, and approval requirement. It omits any mention of the response shape or permission scope, a minor gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is only 50%, with label_name entirely undocumented in the schema. The description compensates by explaining the label_name syntax and the '/' nesting convention, which is the key semantic an agent needs; the reason parameter is already self-documented in the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Create a Gmail label'), which is clear and unambiguous. It does not, however, differentiate itself from siblings like gmail_add_label (applies a label to a message) or gmail_create_filter, so an agent must infer the distinction.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides concrete operational guidance ('/' for nesting, fails on duplicate name, requires approval) that shapes correct invocation. It never states when to prefer this over gmail_add_label or how it relates to filters, so usage context is only implied.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
gmail_download_attachmentARead-onlyIdempotent
Download a Gmail attachment's content. Identify the attachment by the name returned from gmail_list_message_attachments. On a local install: saved to destination_dir, and the saved file path is returned -- destination_dir is required, there is no default, so choose deliberately: pass ~/Downloads (or another path the user asked for) when this attachment is a deliverable the user should find afterward, or your own working/scratch directory when you're only downloading it to read or process it yourself. On an organization-managed install: destination_dir is ignored (there is no local filesystem you and the human share) -- a small attachment's bytes come back directly in this tool's result so you can read or hand it to the human yourself; a larger one comes back as a one-time link the human opens in their own signed-in browser tab instead. Requires user approval.
| Name | Required | Description | Default |
|---|---|---|---|
| reason | Yes | One sentence: why are you calling this tool right now? | |
| message_id | Yes | ||
| attachment_name | Yes | ||
| destination_dir | Yes | Local install only: where to save the attachment -- required, no default. On an organization-managed install it is ignored and nothing is saved to it; any value will do. Use ~/Downloads (or a path the user specified) if the user should find this file afterward; use your own working/scratch directory if it's only for you to read or process. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint/idempotentHint/non-destructive, so the safety bar is low, yet the description adds substantial behavior: local vs organization-managed install handling, that destination_dir is ignored in org mode, small-attachment bytes vs a one-time human-opened link, and that user approval is required. This is rich context beyond the structured fields.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core action and identification method, then the conditional destination_dir guidance. Every sentence carries information, though the local/org branching is dense and there is some overlap with the schema's destination_dir description.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema and four required params, the description fills the gap by explaining what is returned (saved file path on local installs, inline bytes or a one-time link on org installs). Combined with the approval requirement and install-mode branching, an agent has everything needed to call it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 50%. The description meaningfully expands destination_dir (required, no default, ignored on org installs), matching the schema's own description. However message_id, attachment_name, and reason get no explanation in the description beyond context clues, so it only partially compensates for the coverage gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Download a Gmail attachment's content') and immediately distinguishes itself from the sibling gmail_list_message_attachments, which only identifies attachments. An agent can tell exactly what this does without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear operational context: it names gmail_list_message_attachments as the source of attachment_name and tells the agent how to choose destination_dir deliberately (~/Downloads for a deliverable vs scratch dir for processing). It lacks an explicit 'do not use this for X' exclusion or a named alternative download path, so it falls short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
gmail_get_messageARead-onlyIdempotent
Fetch a single Gmail message by id, including body, metadata, and attachment list. Requires user approval.
| Name | Required | Description | Default |
|---|---|---|---|
| reason | Yes | One sentence: why are you calling this tool right now? | |
| message_id | Yes | ||
| include_html | No | Also return the message's HTML part as body_html (empty for a plain-text-only message). Only when you need the markup itself, e.g. to check how a draft's formatting or signature was saved -- body_text is enough to read a message. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already establish readOnly, idempotent, non-destructive semantics, so the description's addition of 'Requires user approval' is genuine extra behavioral context an agent needs before calling. It also discloses the return shape (body, metadata, attachment list). It stops short of covering pagination/size limits or what approval flow is triggered.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences, no filler: what it returns first, then the approval requirement. Every clause earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 3-parameter read tool with no output schema, the description covers return contents, keying, and the approval prerequisite. The only real gap is the absence of sibling routing guidance, which is a usage concern more than a completeness one.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 67%: include_html is documented in exhaustive detail (when to use HTML vs body_text) and reason carries a description, but message_id has none. The description's mention of body/metadata/attachment list loosely maps to include_html but adds no syntax or format detail beyond the schema, so baseline 3 is right.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (fetch), resource (a single Gmail message), and key (by id), plus exactly what the payload contains (body, metadata, attachment list). This cleanly separates it from gmail_list_messages, gmail_list_threads, and gmail_get_thread without needing to name them.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance about when to use this versus gmail_get_thread, gmail_list_messages, or gmail_download_attachment. 'Requires user approval' is a prerequisite/behavioral note, not routing guidance, so usage must be inferred entirely.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
gmail_get_threadBRead-onlyIdempotent
Fetch a full Gmail thread by id, including all messages. Requires user approval.
| Name | Required | Description | Default |
|---|---|---|---|
| reason | Yes | One sentence: why are you calling this tool right now? | |
| thread_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint and destructiveHint, so the safety profile is covered. The description adds one genuinely non-redundant trait, 'Requires user approval,' but says nothing about response size, message ordering, or pagination for large threads. With annotations carrying most of the burden, a 3 is appropriate.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences with no filler, and the core behavior plus the approval constraint are both front-loaded. Every clause earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only fetch with no output schema, the description discloses what is returned ('all messages') and the approval gate, which is enough for an agent to call it correctly. The undocumented thread_id format is the only real gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 50%; thread_id has no schema description and reason carries its own. The description's 'by id' loosely maps to thread_id but adds no format detail, and it never explains the required reason field beyond what the schema already says. Baseline 3 fits given the partial coverage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Fetch a full Gmail thread by id, including all messages') and the phrase 'full thread... including all messages' distinguishes it from gmail_get_message. It does not explicitly name the sibling it is not, so it falls just short of the 5 bar.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No when-to-use or when-not-to-use guidance and no alternatives named. An agent is not told how this differs from gmail_get_message or gmail_list_threads, nor when a thread fetch is preferable to a single-message fetch. Only the approval prerequisite hints at usage context.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
gmail_list_filtersARead-onlyIdempotent
List all Gmail filters with their criteria and actions. Auto-approved -- filter rules only, no message content is returned.
| Name | Required | Description | Default |
|---|---|---|---|
| reason | Yes | One sentence: why are you calling this tool right now? |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover the read-only/idempotent/non-destructive profile, so this adds value by disclosing two things they do not: the call is 'auto-approved' (no approval round-trip) and no message content is returned, bounding the data exposure. It still says nothing about volume, pagination, or ordering of results.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences with zero waste; the core purpose is front-loaded and the caveat about returned content follows as supporting detail.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description appropriately names what comes back ('criteria and actions'), and the absence of pagination or result-count guidance is a minor gap for a simple read-only listing tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% and the single 'reason' parameter is fully documented in the schema, so the baseline is 3. The description adds no syntax, format, or constraint detail beyond what the schema already provides.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('List all Gmail filters') and describes the returned shape ('criteria and actions'), which lets an agent distinguish it from the sibling mutators gmail_create_filter and gmail_update_filter without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is only implied — an agent can infer you call this to inspect existing filters, but the description never states when to reach for this versus gmail_list_labels or the create/update siblings, and offers no exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
gmail_list_labelsARead-onlyIdempotent
List all Gmail labels (system and user-created). Nested labels have a '/' in their name (e.g. 'Work/Projects'). Auto-approved -- label metadata only.
| Name | Required | Description | Default |
|---|---|---|---|
| reason | Yes | One sentence: why are you calling this tool right now? |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, and destructiveHint=false, so the safety profile is covered. The description adds genuinely new context: 'Auto-approved -- label metadata only,' which tells the agent no approval gate applies and that only metadata (not message content) is touched. It stops short of describing pagination or return ordering.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences with zero filler; the core purpose is front-loaded and the nesting/approval notes follow as useful qualifiers rather than padding.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple read-only list tool with no output schema and annotations covering safety, the description supplies enough: what is listed, how nesting is encoded, and that it is auto-approved. Only minor gaps remain (ordering, pagination, whether dead/system labels are included), none of which block correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Only one parameter (reason) exists and schema description coverage is 100%, so the schema fully documents it. The description adds nothing about the reason argument, so this sits at the baseline for a fully-covered schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (list) and resource (Gmail labels) with an explicit scope: system and user-created. The clarification about '/' nesting makes the output shape predictable and clearly separates this from siblings like gmail_list_messages, gmail_list_filters, and gmail_create_label.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is only implied by the read-only listing nature; there is no explicit when-to-use statement or pointer to alternatives such as gmail_add_label or gmail_create_label. An agent can infer it is the enumeration step before labeling, but nothing in the text says so.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
gmail_list_message_attachmentsARead-onlyIdempotent
List attachment names, MIME types, and sizes for a Gmail message. Auto-approved — metadata only, no attachment content is returned. Use gmail_download_attachment to fetch the actual file.
| Name | Required | Description | Default |
|---|---|---|---|
| reason | Yes | One sentence: why are you calling this tool right now? | |
| message_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly/idempotent/non-destructive, but the description adds real context beyond them: 'Auto-approved' discloses the approval policy and 'no attachment content is returned' sets a hard expectation about output scope. It does not mention pagination or size limits, but the core behavioral traits are well covered.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, each earning its place: capability, safety/scope constraint, and the alternative routing. The most decision-relevant information is front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description carries the return-value burden, and it does so by naming the three fields returned and explicitly excluding content. Only the undocumented message_id parameter keeps this from being fully complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 50%: the 'reason' parameter is documented in the schema, but message_id has no description in either place. The phrase 'for a Gmail message' hints at message_id's role but adds no format or sourcing guidance, so the description only marginally compensates for the coverage gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (List) and resource (attachment metadata) with the exact returned fields: names, MIME types, sizes. It also implicitly scopes to a single Gmail message, distinguishing it cleanly from gmail_list_messages and gmail_download_attachment.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly routes the agent to the alternative: 'Use gmail_download_attachment to fetch the actual file.' Combined with 'metadata only', it clearly defines when to use this tool versus when to use the download sibling.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
gmail_list_messagesARead-onlyIdempotent
Search Gmail and return matching message summaries (id, thread_id, subject, sender, date). Auto-approved — no body content is returned.
| Name | Required | Description | Default |
|---|---|---|---|
| query | Yes | ||
| reason | Yes | One sentence: why are you calling this tool right now? | |
| max_results | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint and destructiveHint, so the safety profile is covered. The description adds real value beyond that: 'Auto-approved' tells the agent no approval round-trip occurs, and 'no body content is returned' scopes the response payload. It stops short of pagination or result-ordering behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences, zero filler, with the search action and the return-payload constraint front-loaded. Every clause earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, and the description does the minimum by listing returned fields and ruling out bodies. But it omits query syntax, result limits/pagination, and when to prefer a sibling, which is material for a search tool with undocumented parameters.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With schema description coverage at 33%, the description must compensate, and it does not: 'query' (the key parameter) gets no syntax or Gmail-operator guidance, and max_results is never mentioned. Only 'reason' is documented, and that is in the schema, not the description.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Search Gmail') and enumerates exactly what comes back (id, thread_id, subject, sender, date). That message-level return shape plus the 'no body content' clause implicitly separates it from gmail_list_threads and gmail_get_message, so an agent can route without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied (call it when you need to find matching messages via a query) but no alternative is ever named and no when-not condition is given. The 'no body content' note hints that gmail_get_message is the follow-up for full content, but the description leaves the agent to infer that.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
gmail_list_threadsARead-onlyIdempotent
Search Gmail and return matching thread summaries (id, snippet). Auto-approved — snippet is a short excerpt of the last message's body, subject to the same 'body' privacy category as gmail_get_message.
| Name | Required | Description | Default |
|---|---|---|---|
| query | Yes | ||
| reason | Yes | One sentence: why are you calling this tool right now? | |
| max_results | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already establish read-only, idempotent, non-destructive semantics, yet the description adds value beyond them: the auto-approval status and the fact that the snippet inherits the 'body' privacy category from gmail_get_message. This is meaningful policy context an agent cannot derive from annotations alone.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences with the action and return shape front-loaded, followed by the privacy/approval caveat. Nothing is wasted or buried.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description usefully specifies the return shape (id, snippet) and the privacy classification of the snippet. It is incomplete for a search tool, though, because it gives no query-format guidance and never mentions the max_results control.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is only 33% (just 'reason'), and the description says nothing about the 'query' parameter's Gmail search syntax or the 'max_results' default of 10. For a search tool whose whole usefulness hinges on query syntax, this is a significant omission rather than adequate compensation.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Search Gmail and return matching thread summaries') and even names the return fields (id, snippet), so an agent knows it retrieves threads rather than messages. It never explicitly names the sibling gmail_list_messages, so the threads-vs-messages distinction must be inferred from the noun.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The 'Auto-approved' note tells the agent this call needs no approval step, which is real operational guidance. However, there is no statement of when to prefer this over gmail_list_messages or gmail_get_thread, and no prerequisites or exclusions, so usage is only implied.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
gmail_remove_labelB
Remove a label from a Gmail message. Requires user approval.
| Name | Required | Description | Default |
|---|---|---|---|
| reason | Yes | One sentence: why are you calling this tool right now? | |
| label_name | Yes | ||
| message_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare readOnlyHint=false, idempotentHint=false, destructiveHint=false, and the description correctly reinforces a mutation-plus-approval flow with 'Requires user approval'. It does not disclose reversibility, whether other labels are affected, or permission requirements, so the added value beyond annotations is modest.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two very short sentences with the action front-loaded and the approval requirement immediately after; nothing is wasted or buried.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a three-required-parameter mutation tool with no output schema, the definition is thin: two of the three parameters are undocumented and behavior on approval/rejection, and the effect on the message's label set, is left unspecified.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is only 33% — only 'reason' is documented, and even that is a generic approval prompt. label_name and message_id carry no explanation in either the schema or the description, so the description fails to compensate for the coverage gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource (remove a label from a Gmail message), and the naming distinguishes it from the sibling gmail_add_label. It is clear but adds no scope detail beyond the operation itself.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied by the name and by the paired sibling gmail_add_label, and 'Requires user approval' flags a prerequisite, but there is no explicit when-to-use vs. alternative guidance or exclusion criteria.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
gmail_reply_all_draftA
Create a Gmail draft replying to all participants of a message (original sender plus To/Cc recipients, excluding yourself), staying in the same thread. Requires user approval.
| Name | Required | Description | Default |
|---|---|---|---|
| cc | No | ||
| bcc | No | ||
| body | No | Plain-text body. May be omitted if body_markdown is given -- the plain-text alternative is then auto-derived from it. At least one of body/body_markdown is required. | |
| reason | Yes | One sentence: why are you calling this tool right now? | |
| send_as | No | Send from this Gmail send-as address instead of the account's default (sets From:, and the signature used is this address's own). Must be one of the account's configured send-as addresses -- anything else is rejected. | |
| message_id | Yes | ||
| body_markdown | No | Optional Markdown body for a rich-text draft. Supports **bold**, *italic*, ==highlight==, [links](url), bullet/numbered lists, and `# Heading 1`/`## Heading 2` (rendered as Gmail's Large/Huge font-size presets, not raw heading tags -- no tables). When given, the draft is sent as plain text + HTML together, so it renders formatted in HTML-capable clients and as readable plain text everywhere else. | |
| include_signature | No | Append the user's Gmail signature (the one Gmail stores for the sending address) to the end of the body. Omit to use the user's 'Append Gmail signature to drafts' setting. Don't also write a sign-off block of your own when this is on. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations cover the safety profile (readOnlyHint=false, destructiveHint=false, idempotentHint=false), so the bar is lower. The description still adds real value: it discloses that a draft (not a send) is produced, that the user must approve it, and that it stays in the same thread with self excluded from recipients.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences with zero waste; the recipient/thread scoping comes first and the approval constraint follows.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a draft-creating tool with no output schema, the description covers the key behavioral facts (draft not send, approval gate, thread continuation, recipient computation). Missing only edge-case behavior on invalid message_id or rate limits, which is minor.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 63% and the description adds no parameter-level detail beyond the recipient rule, which maps only loosely to cc/bcc. With the schema documenting send_as, include_signature, and body variants itself, baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (create a draft) and resource (Gmail reply-all draft) and precisely defines the recipient semantics ('original sender plus To/Cc recipients, excluding yourself') plus thread behavior. This cleanly separates it from gmail_reply_draft and gmail_create_draft.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Implies the reply-all use case through its recipient definition and states a prerequisite ('Requires user approval'), but never explicitly contrasts with gmail_reply_draft or says when reply-all is preferred over a single-recipient reply. Usage is inferable but not spelled out.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
gmail_reply_all_draft_with_attachmentsA
Create a Gmail draft replying to all participants of a message (original sender plus To/Cc recipients, excluding yourself), staying in the same thread, with one or more local-file attachments. Parallel to gmail_reply_all_draft -- use this variant only when there is something to attach. Requires user approval.
| Name | Required | Description | Default |
|---|---|---|---|
| cc | No | ||
| bcc | No | ||
| body | No | Plain-text body. May be omitted if body_markdown is given -- the plain-text alternative is then auto-derived from it. At least one of body/body_markdown is required. | |
| reason | Yes | One sentence: why are you calling this tool right now? | |
| send_as | No | Send from this Gmail send-as address instead of the account's default (sets From:, and the signature used is this address's own). Must be one of the account's configured send-as addresses -- anything else is rejected. | |
| message_id | Yes | ||
| attachments | Yes | JSON array of local file paths to attach, e.g. ["/path/to/report.pdf"] -- each either a path on the user's computer (where Claude Desktop runs: absolute, or starting with ~/. Claude's own working or outputs directory is fine) or 'upload:<upload_id>', the id privacyfence_create_upload_slot returned after you PUT the file's bytes to its upload_url -- use that form if a plain path fails with an error about PrivacyFence being unable to read files in your home folder directly (e.g. no PrivacyFence extension is installed, or this is an organization-managed install: a plain path there is read from wherever PrivacyFence's own server runs, not the user's machine -- prefer 'upload:<upload_id>'). At least one required. | |
| body_markdown | No | Optional Markdown body for a rich-text draft. Supports **bold**, *italic*, ==highlight==, [links](url), bullet/numbered lists, and `# Heading 1`/`## Heading 2` (rendered as Gmail's Large/Huge font-size presets, not raw heading tags -- no tables). When given, the draft is sent as plain text + HTML together, so it renders formatted in HTML-capable clients and as readable plain text everywhere else. | |
| include_signature | No | Append the user's Gmail signature (the one Gmail stores for the sending address) to the end of the body. Omit to use the user's 'Append Gmail signature to drafts' setting. Don't also write a sign-off block of your own when this is on. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare readOnly=false, idempotent=false, destructive=false, but the description adds real context: the recipient set is derived (sender + To/Cc, self excluded), the draft stays in the original thread, and user approval is required. It does not cover whether repeated calls duplicate drafts or any rate limit, so it stops short of a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences with zero filler; the core action and recipient scope come first, followed by the sibling routing and the approval requirement.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 9-parameter, 3-required mutation with no output schema and no annotations on safety specifics, the description supplies the essential selection and threading context and flags approval. Remaining gaps (duplicate-draft behavior, response shape) are minor since no output schema exists.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 67%, so most parameters are already documented in the schema. The description only adds 'one or more local-file attachments' to the attachments parameter, which the schema itself explains in far greater detail (path vs upload:<upload_id>). Baseline 3 is appropriate when the schema does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (create draft), resource (Gmail draft), the exact recipient computation (original sender plus To/Cc, excluding yourself), thread behavior, and attachment capability. It names the sibling gmail_reply_all_draft and how this differs, so an agent can select it without opening either schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly routes the agent: 'use this variant only when there is something to attach,' naming the alternative variant and the selecting condition. It also states the prerequisite 'Requires user approval,' covering when-not and gating in two sentences.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
gmail_reply_draftA
Create a Gmail draft replying to a single message, staying in the same thread (sets threadId plus In-Reply-To/References so it actually threads, unlike gmail_create_draft). Addressed only to the original sender. Requires user approval.
| Name | Required | Description | Default |
|---|---|---|---|
| cc | No | ||
| bcc | No | ||
| body | No | Plain-text body. May be omitted if body_markdown is given -- the plain-text alternative is then auto-derived from it. At least one of body/body_markdown is required. | |
| reason | Yes | One sentence: why are you calling this tool right now? | |
| send_as | No | Send from this Gmail send-as address instead of the account's default (sets From:, and the signature used is this address's own). Must be one of the account's configured send-as addresses -- anything else is rejected. | |
| message_id | Yes | ||
| body_markdown | No | Optional Markdown body for a rich-text draft. Supports **bold**, *italic*, ==highlight==, [links](url), bullet/numbered lists, and `# Heading 1`/`## Heading 2` (rendered as Gmail's Large/Huge font-size presets, not raw heading tags -- no tables). When given, the draft is sent as plain text + HTML together, so it renders formatted in HTML-capable clients and as readable plain text everywhere else. | |
| include_signature | No | Append the user's Gmail signature (the one Gmail stores for the sending address) to the end of the body. Omit to use the user's 'Append Gmail signature to drafts' setting. Don't also write a sign-off block of your own when this is on. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare the mutation profile (readOnlyHint=false, idempotentHint=false, destructiveHint=false), and the description adds real value beyond that: it explains the threading behavior (threadId plus In-Reply-To/References), the recipient restriction, and that user approval is required. It omits any duplicate-draft risk or whether drafts are re-created on repeated calls.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight sentences, front-loaded with the action and the thread behavior, and the sibling contrast is packed into a parenthetical with no filler. Only slight compression cost is that the approval note trails at the end.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a write tool with no output schema, the description covers the essential behavioral facts an agent needs (threading, single-sender addressing, approval gate). It leaves gaps around cc/bcc semantics and send-as/signature behavior, but the schema documents those, so the definition is largely complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 63%, so several parameters carry their own docs, but the description adds almost no parameter-level detail for the 8 params (send_as, include_signature, cc/bcc are untouched). Notably it says 'Addressed only to the original sender' while the schema exposes cc/bcc, which leaves the agent unsure how those fields interact with that claim.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a precise verb+resource ('Create a Gmail draft replying to a single message') and explicitly distinguishes itself from the sibling gmail_create_draft by naming the threading mechanism. An agent can tell it apart from gmail_reply_all_draft and gmail_create_draft without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear context for use ('replying to a single message', 'Addressed only to the original sender'), which implicitly routes multi-recipient replies to gmail_reply_all_draft. It does not name that sibling explicitly or state when-not-to-use, so it stops short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
gmail_reply_draft_with_attachmentsA
Create a Gmail draft replying to a single message, staying in the same thread, with one or more local-file attachments. Parallel to gmail_reply_draft -- use this variant only when there is something to attach. Addressed only to the original sender. Requires user approval.
| Name | Required | Description | Default |
|---|---|---|---|
| cc | No | ||
| bcc | No | ||
| body | No | Plain-text body. May be omitted if body_markdown is given -- the plain-text alternative is then auto-derived from it. At least one of body/body_markdown is required. | |
| reason | Yes | One sentence: why are you calling this tool right now? | |
| send_as | No | Send from this Gmail send-as address instead of the account's default (sets From:, and the signature used is this address's own). Must be one of the account's configured send-as addresses -- anything else is rejected. | |
| message_id | Yes | ||
| attachments | Yes | JSON array of local file paths to attach, e.g. ["/path/to/report.pdf"] -- each either a path on the user's computer (where Claude Desktop runs: absolute, or starting with ~/. Claude's own working or outputs directory is fine) or 'upload:<upload_id>', the id privacyfence_create_upload_slot returned after you PUT the file's bytes to its upload_url -- use that form if a plain path fails with an error about PrivacyFence being unable to read files in your home folder directly (e.g. no PrivacyFence extension is installed, or this is an organization-managed install: a plain path there is read from wherever PrivacyFence's own server runs, not the user's machine -- prefer 'upload:<upload_id>'). At least one required. | |
| body_markdown | No | Optional Markdown body for a rich-text draft. Supports **bold**, *italic*, ==highlight==, [links](url), bullet/numbered lists, and `# Heading 1`/`## Heading 2` (rendered as Gmail's Large/Huge font-size presets, not raw heading tags -- no tables). When given, the draft is sent as plain text + HTML together, so it renders formatted in HTML-capable clients and as readable plain text everywhere else. | |
| include_signature | No | Append the user's Gmail signature (the one Gmail stores for the sending address) to the end of the body. Omit to use the user's 'Append Gmail signature to drafts' setting. Don't also write a sign-off block of your own when this is on. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare readOnlyHint=false, idempotentHint=false, and destructiveHint=false, so the mutation/non-idempotent profile is already covered by structured data. The description adds one genuinely useful behavioral fact not in the annotations -- "Requires user approval" -- plus thread continuity, but it does not describe what the call returns or side effects on the account. Modest added value over annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four short sentences, each earning its place: purpose, sibling routing, recipient scope, approval gate. Core action is front-loaded and there is no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 9-param mutation tool with no output schema, the definition covers routing, thread behavior, recipient scope, and the approval requirement -- enough for correct invocation. It omits any hint of the draft result or how cc/bcc interact with the 'only to the original sender' framing, leaving a small gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 67%, with rich in-schema descriptions for attachments, body_markdown, send_as, and include_signature, so the schema carries most parameter meaning. The description only gestures at attachments ("one or more local-file attachments") without adding syntax or constraints, and says nothing about cc/bcc, which have no schema descriptions. Baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource (create a Gmail draft replying to a message) plus scope details (same thread, local-file attachments). It explicitly names the parallel sibling gmail_reply_draft and the condition that selects this variant, letting an agent distinguish it without opening either schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives an explicit when-to-use rule ("use this variant only when there is something to attach") and distinguishes the recipient axis from reply_all via "Addressed only to the original sender." Both routing decisions an agent must make are answered directly.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
gmail_update_filterA
Replace an existing Gmail filter's criteria and actions, identified by filter_id (from gmail_list_filters). Gmail's API has no native filter update, so this deletes the filter and creates a new one with the given fields, which gets a new id. Requires user approval.
| Name | Required | Description | Default |
|---|---|---|---|
| star | No | ||
| query | No | ||
| reason | Yes | One sentence: why are you calling this tool right now? | |
| archive | No | ||
| subject | No | ||
| filter_id | Yes | ||
| forward_to | No | ||
| to_address | No | ||
| from_address | No | ||
| mark_as_read | No | ||
| has_attachment | No | ||
| add_label_names | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Discloses a non-obvious implementation quirk that annotations cannot convey: Gmail has no native update, so the filter is deleted and recreated, yielding a new id. It also flags the approval requirement. However, the description leaves the annotation tension unaddressed - it says the filter is deleted while destructiveHint is false - and does not clarify whether existing criteria/actions are fully replaced or merged.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, front-loaded with the verb+resource, then the API quirk, then the approval gate. Zero filler; every sentence carries decision-relevant information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema, so the description must carry more weight, and it does cover the key operational fact (delete-and-recreate, new id) and the approval gate. For a 12-parameter mutating tool at 8% schema coverage, though, it omits replacement semantics and per-field meaning, leaving real gaps for correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is only 8% across 12 parameters; only filter_id (via the list tool) and reason get any treatment. The description does not resolve the most consequential ambiguity: with string defaults of "" and boolean defaults of false, does omitting a field clear the previous criterion/action, and are add_label_names comma-separated? Most field names are self-explanatory, which keeps this above 1, but the description fails to compensate for the near-total schema gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource (replace an existing Gmail filter's criteria and actions) and immediately disambiguates from the sibling set by naming gmail_list_filters as the source of filter_id and noting that this is distinct from gmail_create_filter. An agent can tell this apart from gmail_create_filter without opening either schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It says when to use it (modifying an existing filter, with filter_id obtained from gmail_list_filters) and adds a hard precondition ("Requires user approval"), which is genuinely actionable. It stops short of explicitly contrasting with gmail_create_filter ("use that instead when no filter exists yet"), so it is clear context without full when-not guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
jira_add_commentA
Add a comment to an existing Jira issue. Requires user approval.
| Name | Required | Description | Default |
|---|---|---|---|
| body | Yes | Comment text (plain text) | |
| reason | Yes | One sentence: why are you calling this tool right now? | |
| issue_key | Yes | e.g. PROJ-123 |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare the safety profile (readOnlyHint=false, idempotentHint=false, destructiveHint=false), so the description's job is to add context beyond that. 'Requires user approval' is a genuinely useful behavioral constraint the annotations do not convey, though return/confirmation behavior after posting is not described.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences with zero waste, and the core action is front-loaded ahead of the approval constraint. Every clause earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a single-mutation comment tool with full schema coverage and annotations covering the safety profile, the definition gives enough to call it correctly. No output schema means return values need not be explained; minor gaps remain around any length/rate constraints on comments.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all three parameters (issue_key, body, reason) are already documented in the schema, including format hints like 'PROJ-123'. The description adds no further parameter meaning, which is the expected baseline when the schema carries the full load.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Add') and resource ('a comment to an existing Jira issue'), and the target is clearly distinguishable from siblings like jira_create_issue and jira_update_issue. An agent can identify the operation without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
'Requires user approval' gives one useful precondition for invoking the tool, implying it should not be called silently or autonomously. However, it names no alternatives (e.g. when to comment vs. update an issue) and offers no explicit when-not-to-use guidance, leaving usage context largely implied.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
jira_create_issueC
Create a new Jira issue. Requires user approval.
| Name | Required | Description | Default |
|---|---|---|---|
| reason | Yes | One sentence: why are you calling this tool right now? | |
| summary | Yes | ||
| priority | No | e.g. High, Medium, Low | |
| issue_type | No | e.g. Task, Bug, Story | Task |
| description | No | ||
| project_key | Yes | e.g. MYPROJ |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=false, idempotentHint=false, and destructiveHint=false, so the mutation profile is known. The description adds the genuinely useful behavioral fact that user approval is required, but says nothing about what is created, whether it can be undone, or what permissions are needed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences with zero filler and the core action front-loaded. It is efficient, though the approval note could have been paired with a bit more context without losing tightness.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 6-parameter mutation tool with no output schema, the description is very thin: no post-creation behavior, no coverage of required fields, and no return expectations. Annotations cover safety only, leaving a real completeness gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is only 67%, below the threshold where the schema can carry the burden alone, and the description supplies no parameter meaning at all. The required fields (project_key, summary, reason) and their format expectations are left entirely to the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
'Create a new Jira issue' states a clear verb+resource, which is easy to distinguish from update/transition siblings. However, it does not name a sibling or scope what is being created (fields, defaults), so it stops short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The only guidance is 'Requires user approval,' which is a procedural constraint, not when-to-use guidance. There is no mention of alternatives (e.g. jira_update_issue for existing issues) or the conditions that select this tool.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
jira_get_issueARead-onlyIdempotent
Fetch full details of a Jira issue by key (e.g. PROJ-123), including description and comments. Requires user approval.
| Name | Required | Description | Default |
|---|---|---|---|
| reason | Yes | One sentence: why are you calling this tool right now? | |
| issue_key | Yes | e.g. PROJ-123 |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already establish this as a safe, idempotent read (readOnlyHint=true, destructiveHint=false). The description adds two pieces of context the annotations can't: the required approval step and the fact that comments are included in the result, which is useful for an agent deciding whether another call is needed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two compact sentences, front-loaded with the operation then the behavioral caveats (payload scope, approval). No filler or repetition.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple 2-param read tool with full schema coverage and no output schema, the description covers purpose, key format, returned fields, and the approval requirement. It could go further by mentioning error behavior for missing keys, but it is otherwise sufficient to call the tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% – both issue_key and reason are documented in the schema with examples. The description reinforces the key format ('PROJ-123') but adds no new semantic detail, so the baseline of 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb (Fetch), resource (Jira issue), and lookup key ('by key e.g. PROJ-123'), clearly distinguishing it from siblings like jira_search_issues and jira_list_projects. It also enumerates the payload (description and comments), so an agent knows exactly what it gets.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It implies the tool is for retrieving a single issue when the key is known, but it never says when to prefer this over jira_search_issues or how to obtain the key. No explicit exclusions or alternative routing are given, so usage is only inferable.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
jira_get_transitionsARead-onlyIdempotent
List the status transitions available for a Jira issue right now (name and target status), given its current workflow state. Use before jira_transition_issue to see what transition names are valid. Auto-approved.
| Name | Required | Description | Default |
|---|---|---|---|
| reason | Yes | One sentence: why are you calling this tool right now? | |
| issue_key | Yes | e.g. PROJ-123 |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint/idempotentHint/no destruction, so safety is covered. The description adds genuinely new behavioral context beyond that: the result is time-dependent and workflow-state-dependent ('right now', 'given its current workflow state'), it discloses the return payload shape (name and target status), and 'Auto-approved' signals no approval gate is required.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences; the core action is front-loaded and the routing guidance follows immediately. No filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a two-parameter read-only lookup with no output schema, the description tells the agent what it returns (transition names and target statuses), when to call it, and that no approval is needed. Nothing needed to invoke it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so both parameters (issue_key, reason) are already documented in the schema. The description adds no format or interpretation detail beyond what the schema provides; baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('List the status transitions available for a Jira issue right now') plus the scope qualifier 'given its current workflow state', which distinguishes it from the sibling jira_transition_issue that actually performs the change.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly routes the agent: 'Use before jira_transition_issue to see what transition names are valid.' It names the downstream sibling and the condition that makes this tool necessary, leaving nothing to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
jira_list_projectsBRead-onlyIdempotent
List Jira projects accessible to the user (key, name, type, lead). Auto-approved.
| Name | Required | Description | Default |
|---|---|---|---|
| reason | Yes | One sentence: why are you calling this tool right now? | |
| max_results | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint and destructiveHint, so safety is covered. The description adds 'Auto-approved', which is a genuine behavioral trait beyond the annotations — the agent learns no approval round-trip is needed. It still omits pagination/result-limit behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single compact sentence plus a two-word note; returned fields are front-loaded and there is zero filler. Slightly terse rather than bloated, so it earns a 4.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple two-parameter read tool with no output schema, listing the returned fields and noting auto-approval covers the essentials. However, the undocumented max_results and absent pagination semantics leave a visible gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is only 50%: the required 'reason' parameter is documented in the schema, but max_results has no description anywhere. The description mentions no parameters at all, so it does not compensate for the undocumented max_results or clarify the reason/approval coupling.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (List) and resource (Jira projects) plus the returned fields (key, name, type, lead), so the agent knows exactly what comes back. It is clear but does not differentiate itself from any sibling (e.g. apps_script_list_projects), which keeps it out of 5 territory.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no statement of when to use this tool, when not to, or which sibling to prefer for a project-like listing. The only context is 'accessible to the user', which is scope, not guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
jira_search_issuesBRead-onlyIdempotent
Search Jira issues using JQL. Returns summary info for matching issues. Auto-approved.
| Name | Required | Description | Default |
|---|---|---|---|
| jql | Yes | e.g. 'project = MYPROJ AND status = Open' | |
| reason | Yes | One sentence: why are you calling this tool right now? | |
| max_results | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, and destructiveHint=false, so the safety profile is covered. The description usefully adds the return scope ('summary info') and that calls are auto-approved, but says nothing about pagination, result caps, or JQL failure behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, front-loaded with the primary action and JQL requirement; no filler. Slightly clipped, but nothing wasted.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, and the description only gestures at returns ('summary info') without describing fields or limits. Combined with an unexplained required 'reason' parameter and a bare max_results, the definition is adequate but leaves real gaps for an agent to fill by trial.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 67%: jql and reason are described in the schema, while max_results carries only a default. The description adds no JQL syntax detail, tie to max_results, or explanation of the mandatory meta-parameter 'reason', so it does not compensate for the gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Search Jira issues using JQL') and adds the scope of the return ('summary info for matching issues'). It implicitly differs from jira_get_issue (single issue) and jira_list_projects, but never names a sibling to reinforce the distinction.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives no when-to-use guidance, no condition for choosing this over jira_get_issue or jira_list_projects, and no exclusions. The only usage-adjacent signal is the word 'Auto-approved', which speaks to approval flow rather than tool selection.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
jira_transition_issueA
Move a Jira issue to a new status by transition name (e.g. "Done", "In Progress") — call jira_get_transitions first to see what's valid from the issue's current status. Requires user approval.
| Name | Required | Description | Default |
|---|---|---|---|
| reason | Yes | One sentence: why are you calling this tool right now? | |
| issue_key | Yes | e.g. PROJ-123 | |
| transition_name | Yes | e.g. Done, In Progress |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=false, idempotentHint=false, and destructiveHint=false, so the safety profile is partly covered. The description adds genuinely useful context beyond that: the transition names are state-dependent (must be validated first) and the operation 'Requires user approval', which an agent needs to know before invoking.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight clauses with no filler: the action and mechanism come first, then the prerequisite and approval warning. Every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With annotations present and no output schema required, the description covers the main gaps an agent needs: dependency on jira_get_transitions and the approval requirement. It omits failure behavior for invalid transition names, a minor gap for a simple 3-parameter tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all three parameters (issue_key, transition_name, reason) are already documented in the schema, including the same 'Done'/'In Progress' examples. The description adds no format or constraint detail beyond the schema, so the baseline of 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Move a Jira issue to a new status') and the mechanism ('by transition name'), with concrete examples. It is clearly distinguishable from siblings like jira_update_issue and jira_get_transitions.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly tells the agent to call jira_get_transitions first to see which transitions are valid from the issue's current status, which is real operational guidance. It stops short of stating when this tool should be preferred over jira_update_issue, but the prerequisite call is the key routing information.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
jira_update_issueA
Update fields on an existing Jira issue (summary, description, priority, and/or custom fields). Requires user approval.
| Name | Required | Description | Default |
|---|---|---|---|
| reason | Yes | One sentence: why are you calling this tool right now? | |
| summary | No | ||
| priority | No | ||
| issue_key | Yes | ||
| description | No | ||
| custom_fields | No | JSON object mapping Jira Cloud custom field display names (as seen in the Jira UI, not their customfield_NNNNN id) to new values, e.g. {"Story Points": 5, "Sprint": "Sprint 12"} |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=false, idempotentHint=false, and destructiveHint=false. The description adds the approval requirement, which is useful behavioral context, but does not explain update semantics such as partial-field behavior or overwriting.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, front-loading the purpose and then the approval requirement. Every sentence is useful and no extraneous text is present.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 6-parameter mutation with no output schema and low schema coverage, the description covers purpose and approval but omits usage alternatives, required-parameter context, and update semantics. It is minimally adequate but leaves clear gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is only 33%, so the description should compensate. It lists summary, description, priority, and custom fields, but omits the required issue_key and reason parameters and adds no syntax or format detail beyond what the schema already provides for custom_fields.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Update fields on an existing Jira issue') and enumerates the updatable field types. This distinguishes it from sibling tools like jira_create_issue and jira_transition_issue without requiring the agent to inspect the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The only guidance is 'Requires user approval,' which is a prerequisite rather than a when-to-use statement. It does not say when to choose this over jira_transition_issue, jira_add_comment, or other Jira mutations.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
privacyfence_await_approvalARead-onlyIdempotent
Long-poll pending approvals from gated calls' {status: 'approval_pending', approval_id, ...} results; status only, never content. Before the first call for an approval, relay that result's message and url (binder_url if several) to the user -- never wait silently. If pending_count > 1, first issue your other ready gated calls, then pass all approval_ids in one call (one human pass). Keep timeout_seconds under your client's tool-call timeout. Returns {approval_id: status}: 'pending' (schedule a follow-up if you can, else call again), 'approved' (re-issue the ORIGINAL call with identical arguments -- the only way to get the data), 'denied' (a human said no -- re-issuing will not change that; don't retry, ask the user how to proceed unless denial_feedback says otherwise), 'expired' (re-issuing starts a fresh approval) or 'unknown' (no such id here). denial_feedback holds the user's instruction for a denial: follow it. Returns on any change or at the timeout. Prefer this over re-issuing the original call to poll.
| Name | Required | Description | Default |
|---|---|---|---|
| approval_ids | Yes | ||
| timeout_seconds | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the readOnly/idempotent/non-destructive annotations, it explains long-poll behavior, status-only return semantics, and the full status lifecycle: pending, approved, denied, expired, unknown. It also documents the critical follow-up behavior for each outcome, including re-issuing the original call only after approval and not retrying after denial unless denial_feedback says otherwise.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense but front-loaded: it opens with what the tool does and immediately covers user-relay requirements, batching, and timeout constraints. Every sentence contributes operational guidance, and the status-specific return handling is compactly enumerated.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a complex asynchronous approval workflow with no output schema, the description supplies the missing return semantics, follow-up rules, batching behavior, timeout caveat, and denial-feedback handling. An agent has enough information to invoke and act on the tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must carry parameter meaning. It explains that approval_ids come from gated-call results, should all be passed in one call when pending_count > 1, and that timeout_seconds should stay under the client's tool-call timeout. It does not document ID format or timeout units/default, so it is strong but not exhaustive.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a precise verb and resource: long-poll pending approvals from gated calls' results, and clarifies the payload is status only, never content. It distinguishes this tool from re-issuing the original call for polling and from broader status tools by specifying that it watches approval state.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives explicit sequencing: relay the message and URL before the first call, never wait silently, and if pending_count > 1, issue other ready gated calls first then pass all IDs in one call. It also names the alternative and when to prefer this tool over simply re-issuing the original call to poll.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
privacyfence_begin_unattended_sessionAIdempotent
Tell PrivacyFence this conversation is an unattended/scheduled Cowork run (e.g. a Routine firing on a schedule) with no human necessarily watching, for the rest of this connection. From then on, any gated tool call that isn't already covered by a configured auto-accept rule is denied immediately with a clear error, instead of PrivacyFence opening a native approval dialog that nobody will answer. Call this once at the start of a scheduled run, and pair it with privacyfence_check_policy to plan which steps are safe to attempt. Never changes what auto-accepts, only what happens when nothing does. Errors if an administrator hasn't enabled unattended sessions for this install. Do not call this during a normal interactive conversation -- it makes denials immediate instead of prompting. reason: one sentence on why this session is unattended (e.g. the Routine/schedule that triggered it) -- logged in the audit entry for this session change, since no popup is shown for it to appear in.
| Name | Required | Description | Default |
|---|---|---|---|
| reason | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Adds substantial behavior beyond the annotations: subsequent gated calls not covered by auto-accept rules are denied immediately with a clear error instead of opening a dialog, auto-accept configuration is never altered, the call errors if an admin hasn't enabled unattended sessions, and the change is audit-logged. These are exactly the side-effect and failure-mode details the annotations cannot convey.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The core instruction and the 'for the rest of this connection' scope are front-loaded, and each sentence carries new information (denial behavior, admin precondition, audit logging). It is somewhat dense and the trailing 'reason:' fragment is awkwardly concatenated, but there is little true filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With one parameter, no output schema, and annotations covering safety/idempotency, the description fills every remaining gap an agent needs: when to invoke, when not to, the resulting denial semantics, the admin-enabled precondition, the error case, and the audit behavior.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description carries the burden for the single 'reason' parameter. It explains what the value is for, the expected form ('one sentence on why this session is unattended'), an example trigger, and its consequence (logged in the audit entry since no popup is shown). That compensates well for the empty schema, though it doesn't state a length limit or validation rules.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific action on a specific resource: marking this conversation as an unattended/scheduled Cowork run for the remainder of the connection. It is immediately distinguishable from the sibling privacyfence_end_unattended_session, which does the inverse.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit when-to-use ('Call this once at the start of a scheduled run'), explicit when-not ('Do not call this during a normal interactive conversation'), and names a companion tool to pair it with (privacyfence_check_policy) for planning safe steps.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
privacyfence_check_policyARead-onlyIdempotent
Before calling a gated tool, ask whether that exact call would auto-accept or need a human. Pass the connector, tool and args you're about to call. Returns {gate, verdict, matched_rule, matched_rule_id, reason, pii_gate_may_apply}. verdict is 'auto_accept' (the real call will pass through identically), 'requires_review' (no configured rule can match these args), or 'unknown' (it depends on fetched content this can't see in advance). matched_rule_id is set only for 'auto_accept': the privacyfence_list_policy rule id that lets this call through, usable as privacyfence_propose_policy_change's rule_id. For 'review'-gated (read) tools pii_gate_may_apply is always true: the PII gate scans real content and can force a popup even when a rule matches, which can't be predicted. No external API call, no popup, no side effects -- call it freely while planning, especially before and during an unattended run. reason: one sentence on why you're checking now (logged, self-reported, unverified).
| Name | Required | Description | Default |
|---|---|---|---|
| args | No | ||
| tool | Yes | ||
| reason | Yes | ||
| connector | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover the safety profile (readOnly, idempotent, non-destructive), and the description goes well beyond them: it discloses no external API call, no popup, no side effects, and explains the non-obvious 'unknown' verdict (depends on unfetched content) and the PII-gate caveat that can force a popup even when a rule matches. This is exactly the extra behavioral context the annotations cannot carry.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
It is front-loaded with the primary instruction and every sentence carries content (when to use, return shape, caveats, side-effect profile). It is dense and run-on in places, but there is little fat to remove.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description fully documents the returned fields (gate, verdict, matched_rule_id, reason, pii_gate_may_apply) and their possible values, plus the PII-gate nuance. For a planning/dry-run tool of this complexity, an agent has everything needed to call and interpret it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must compensate, and it largely does: it explains connector/tool/args as the call about to be made and defines reason's purpose and semantics ('one sentence on why you're checking now, logged, self-reported, unverified'). The args object's internal shape is only implied rather than specified, keeping it from a 5.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific verb and resource ('check whether that exact call would auto-accept or need a human') and explicitly contrasts the tool with siblings privacyfence_list_policy and privacyfence_propose_policy_change. An agent can distinguish it from those relatives without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It states exactly when to call it ('Before calling a gated tool', 'while planning, especially before and during an unattended run') and what input to pass. The relationship to sibling tools is spelled out (matched_rule_id is usable as propose_policy_change's rule_id), leaving nothing to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
privacyfence_create_upload_slotA
Get a one-time URL to upload a local file to PrivacyFence, for a client without the PrivacyFence extension (e.g. Claude Code) -- use it when a tool's local_path/attachments parameter says PrivacyFence can't read your files directly. Returns {upload_id, upload_url, method: 'PUT', max_bytes, expires_at, example}: PUT the raw bytes to upload_url (e.g. the curl -T example); no Authorization header -- the URL is the credential. Then pass upload_id to the tool that needs the file (its upload_id parameter, e.g. drive_upload_file, or an 'upload:' attachments entry, e.g. gmail_*_with_attachments). Never fetch upload_url yourself or pass it as local_path. Single-use, expires in 10 minutes, claimable only by your next tool call in this conversation. This uploads, gates and approves nothing: that happens when the destination tool runs. reason: one sentence on why this file is needed now (logged, self-reported, unverified).
| Name | Required | Description | Default |
|---|---|---|---|
| reason | Yes | ||
| filename | Yes | ||
| size_bytes | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only assert non-readOnly, non-idempotent, non-destructive. The description adds life-cycle and security traits well beyond them: single-use, 10-minute expiry, claimable only by the next tool call, no Authorization header because the URL is the credential, and that it neither gates nor approves anything.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Dense but front-loaded: purpose, then trigger, then return shape and usage, then constraints. The single paragraph format for such a nuanced tool is acceptable, though the trailing 'reason:' clause and example could be trimmed for scannability.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, yet the description enumerates the return fields ({upload_id, upload_url, method, max_bytes, expires_at, example}), the required next step, and the security constraints. Nothing an agent needs to invoke and chain this tool correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must compensate. It thoroughly explains the meaning and intent of 'reason' (logged, self-reported, unverified) and implies 'filename', but leaves 'size_bytes' undocumented even though upload limits are relevant. Strong on the semantically important params, with one gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Get a one-time URL to upload a local file to PrivacyFence') and immediately scopes it to a client without the extension. An agent can distinguish this from siblings like drive_upload_file without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives an explicit trigger ('use it when a tool's local_path/attachments parameter says PrivacyFence can't read your files directly'), names the follow-up destinations (drive_upload_file, gmail_*_with_attachments), and states two negative rules ('Never fetch upload_url yourself or pass it as local_path'). When, how, and when-not are all covered.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
privacyfence_end_unattended_sessionAIdempotent
Clear the unattended-session flag set by privacyfence_begin_unattended_session for this connection, restoring normal interactive approval behavior. Call this when a scheduled run finishes. Not strictly required -- the flag also clears automatically when the connection closes -- but call it if this connection might be reused afterward for something interactive. reason: one sentence on why the unattended session is ending now -- logged the same way as privacyfence_begin_unattended_session's.
| Name | Required | Description | Default |
|---|---|---|---|
| reason | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only say readOnly=false, idempotent=true, destructive=false. The description adds substantial context beyond that: it names the flag being cleared, states the effect (restoring normal interactive approval behavior), discloses that the flag auto-clears on connection close, and marks the call as optional. The idempotent annotation is consistent with a flag-clear operation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The purpose is front-loaded in the first clause, and each subsequent sentence adds a distinct fact (timing, optionality, fallback, parameter meaning). Slightly dense — the parameter note is run-on — but nothing is redundant.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a single-param, no-output-schema flag-clearing tool, the description covers purpose, timing, optionality, the automatic fallback, and the lone parameter. An agent has everything needed to decide whether to call it and how to populate reason.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0% and the single parameter is required, so the description must carry the burden — and it does: reason is described as "one sentence on why the unattended session is ending now" and cross-referenced to how begin_unattended_session logs it. Format constraints (length caps, etc.) are left unstated, keeping it short of a 5.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb (Clear) and resource (the unattended-session flag), and explicitly ties it to its counterpart privacyfence_begin_unattended_session. An agent can distinguish it from all siblings, including the begin/await/status PrivacyFence tools, without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives explicit timing ("Call this when a scheduled run finishes") and an explicit conditional for the optional case ("call it if this connection might be reused afterward for something interactive"). It also names the alternative path — automatic clearing on connection close — so the agent knows when NOT to bother calling it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
privacyfence_list_policyARead-onlyIdempotent
List every configured auto-accept rule, plus the scope catalogue privacyfence_propose_policy_change accepts. Returns {rules, scope_groups}. Each rule has: id (pass as privacyfence_propose_policy_change's rule_id to update or remove it), sentence (human-readable), connector, scope_type, value, operations (internal keys, informational), verbs ({verb, family} pairs; family is 'read', 'write', 'send' or 'destructive'), conditions, and covered_tools (every tool name the rule can auto-accept). Each scope_groups entry has an id (pass as group), the verbs that scope can govern (pass a subset as verbs; any other verb is rejected before a popup), and whether it needs a value. Call this before proposing a change: an id or group only matches something real if you listed it rather than guessed. Read-only, no popup. reason: one sentence on why you're listing the policy now (logged, self-reported; this discloses the full rule set).
| Name | Required | Description | Default |
|---|---|---|---|
| reason | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, so the safety profile is covered. The description adds useful behavioral context beyond annotations: no popup, the reason is logged and self-reported, and the call discloses the full rule set. That is meaningful additional transparency, though not exhaustive about permissions or rate limits.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense and front-loads the core purpose and return shape before detailing fields. It is longer than strictly necessary, but the field-level detail is useful because there is no output schema. Most sentences earn their place by helping the agent use the returned data correctly.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema, so the description properly explains the return object {rules, scope_groups} and describes the important fields, including how ids and groups map to privacyfence_propose_policy_change. The single parameter is explained, and the read-only, no-popup behavior is covered with annotations. Nothing critical is missing for correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and there is one required parameter, so the description must carry the semantic burden. It explains that reason should be one sentence on why the policy is being listed now, that it is logged and self-reported, and that this call discloses the full rule set. This substantially compensates for the missing schema description.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: list every configured auto-accept rule plus the scope catalogue accepted by privacyfence_propose_policy_change. It distinguishes itself from the sibling mutation tool by naming that tool directly. An agent can tell what this tool does without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explicitly says to call this before proposing a change and explains why: an id or group only matches something real if it was listed rather than guessed. This gives a clear when-to-use condition tied to the sibling tool privacyfence_propose_policy_change. No alternative listing tool is relevant, so the guidance is complete.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
privacyfence_propose_policy_changeAIdempotent
Propose adding, updating or removing an auto-accept rule (a scope plus allowed verbs). ALWAYS blocks on a dialog a human must approve. If declined, or in an unattended session, it errors -- check the result, never assume success. Call privacyfence_list_policy first: group, verbs and rule_id must be listed, not guessed; a verb not in that group's scope_groups entry is rejected before any popup.
operation='add'/'update' need group (a scope_groups id, e.g. 'drive.folder'), value (resource ids/names the scope matches, e.g. a Drive folder id; omit when the group's needs_value is false) and verbs (a non-empty subset of the group's verbs, e.g. ['read', 'update']). 'update' also takes rule_id (that rule is removed, then re-added). 'remove' needs only rule_id.
Returns {confirmed, changed, description, rule_ids}: changed is false when confirmed but nothing differed; rule_ids names the rules affected.
reason: one sentence on why you're proposing this (logged, unverified).
| Name | Required | Description | Default |
|---|---|---|---|
| group | No | ||
| value | No | ||
| verbs | No | ||
| reason | Yes | ||
| rule_id | No | ||
| operation | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only provide readOnlyHint=false, idempotentHint=true, destructiveHint=false. The description adds substantial context beyond that: human-approval blocking, unattended-session errors, the need to check the result, verb validation happening before the popup, and 'update' being a remove-then-re-add. No contradiction with idempotentHint.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Dense but front-loaded with the core action and the critical 'always blocks / check the result' warning first, then parameter detail. Every sentence carries information, though the parameter paragraph is a tight wall of text that could be more scannable.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description supplies the return shape ({confirmed, changed, description, rule_ids}) and clarifies the changed=false case. For a blocking, approval-gated, 6-param tool this is complete enough to invoke correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% for 6 params, so the description must carry the load and does: operation semantics per value, group/value/verbs constraints with examples ('drive.folder', 'read','update'), rule_id usage per operation, and reason logging. Adds meaning well beyond the bare schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('propose adding, updating or removing an auto-accept rule') and defines what a rule is ('a scope plus allowed verbs'). It is clearly distinguishable from sibling read tools like privacyfence_list_policy and privacyfence_check_policy.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly instructs to call privacyfence_list_policy first and why (group, verbs, rule_id must be listed, not guessed), states it ALWAYS blocks on a human dialog, and explains failure conditions (declined or unattended session errors). Clear when-to-use and prerequisites.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
privacyfence_statusARead-onlyIdempotent
Check whether THIS PrivacyFence install is set up: call it before the first PrivacyFence-governed action in a conversation, or when asked why a connector (gmail_*, ...) is missing. An empty or partial tool list means connectors aren't authenticated yet, NOT that PrivacyFence is irrelevant -- this is the one tool guaranteed to exist even when every other tool is missing. Returns {mode ('local'/'org'), setup_complete (any connector authenticated), connectors [{name, enabled, authenticated, blocked_by: null, 'no_org_config', 'not_authenticated' or a reason}], next_step, message, sign_in_url (always null)}. If setup isn't complete, relay message to the human as-is. next_step 'open_privacyfence_companion' (local): the human opens PrivacyFence's companion app (menu-bar/tray icon; Linux: applications menu) and chooses Open Settings; you can't, and have no link. 'contact_your_administrator' (org): sign-in is via the org's IdP. Only side effect: an audit entry. reason: one sentence on why you're checking now (logged).
| Name | Required | Description | Default |
|---|---|---|---|
| reason | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Discloses behavior well beyond the annotations: the full return shape, how to relay message to the human, what next_step values mean (open_privacyfence_companion vs contact_your_administrator), the org-IdP sign-in path, and the audit-entry side effect. The audit entry is the one nuance worth noting against readOnlyHint=true, but it is honestly disclosed rather than hidden, so this reads as transparency rather than contradiction.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Long and dense, but front-loaded with the action triggers before the return-shape and next_step details. Nearly every sentence informs a decision the agent must make (when to call, how to relay, what next_step means), though the density could be tightened slightly.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema and no annotation coverage of returns, the description carries the full burden and delivers: mode, setup_complete, connectors with blocked_by reasons, next_step, message, and sign_in_url. Nothing an agent needs to interpret the result or act on it appears to be missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The single parameter has 0% schema description coverage, so the description must carry it, and it does: reason is defined as 'one sentence on why you're checking now (logged)', which conveys both format and that it is persisted. It adds real meaning beyond the bare {type: string} schema, though it stops short of any length/format constraint.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Opens with a specific verb+resource: verifying whether this PrivacyFence install is set up. It also carves out a scope no sibling covers ('this is the one tool guaranteed to exist even when every other tool is missing'), so it is unmistakable against privacyfence_check_policy and the other privacyfence_* tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit triggers: call before the first PrivacyFence-governed action in a conversation, or when a connector like gmail_* appears missing. It further explains what an empty/partial tool list does and does not imply, which is exactly the ambiguity an agent would face at call time.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
salesforce_get_recordARead-onlyIdempotent
Fetch a Salesforce record by object type and id. Requires user approval.
| Name | Required | Description | Default |
|---|---|---|---|
| reason | Yes | One sentence: why are you calling this tool right now? | |
| record_id | Yes | ||
| object_type | Yes | e.g. Account, Contact, Opportunity |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover the read-only, idempotent, non-destructive profile, so the added value here is 'Requires user approval' — a genuine behavioral constraint not derivable from the annotations and relevant given the privacyfence approval tools in the sibling set. It stops short of saying how approval is obtained or what happens if it is denied.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences with zero filler, and the core operation is front-loaded ahead of the approval caveat.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description does not indicate what a fetched record returns (field set, format, depth). The approval requirement is flagged but the mechanics are left implicit for a three-required-parameter tool with an unusual gating flow.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 67% and the description only restates the object_type and record_id parameters ('by object type and id') without adding format, casing, or id-syntax details. The 'reason' parameter is documented only in the schema, so the description adds little beyond the baseline.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Fetch) and resource (Salesforce record) with the lookup key (object type and id). It is distinguishable from salesforce_search because it fetches a single known record, though it does not name that sibling explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied by the 'by object type and id' framing, so an agent can infer this is for when the id is already known. However, there is no explicit when-to-use guidance and no comparison against salesforce_search or salesforce_run_report for the case where the id is unknown.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
salesforce_list_reportsBRead-onlyIdempotent
List Salesforce reports accessible to the user. Auto-approved.
| Name | Required | Description | Default |
|---|---|---|---|
| reason | Yes | One sentence: why are you calling this tool right now? |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint and destructiveHint=false, so the safety profile is covered. 'Auto-approved' adds a small piece of behavioral context (no approval gate), but pagination, result limits, and whether inaccessible reports are silently omitted are not disclosed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two very short sentences, front-loaded with the action and resource. Nothing is wasted, though 'Auto-approved' is a fragment whose meaning is not elaborated.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a list tool with no output schema, the description should at least hint at return shape or pagination behavior, which it does not. It is minimally adequate for a simple read operation but leaves return-size and format questions unanswered.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% and the single 'reason' parameter is fully documented in the schema, so the baseline of 3 applies. The description adds no additional parameter meaning beyond what the schema already provides.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('List Salesforce reports') with a scope qualifier ('accessible to the user'), which is enough to separate it from salesforce_run_report and salesforce_get_record. It stops short of explicitly naming those siblings, so it is clear but not fully differentiated.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to choose this over salesforce_run_report or salesforce_search, nor any prerequisite or context. The agent must infer usage entirely from the name.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
salesforce_run_reportBRead-onlyIdempotent
Run a Salesforce report by id and return the results. Requires user approval.
| Name | Required | Description | Default |
|---|---|---|---|
| reason | Yes | One sentence: why are you calling this tool right now? | |
| report_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint and destructiveHint=false, so the safety profile is covered. The description adds one useful behavioral fact — that user approval is required — but says nothing about result size, pagination, or execution time, which matter for a report run.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short, front-loaded sentences with zero filler; the core action comes first. It is efficient, though extremely terse given the gaps in other dimensions.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only tool with no output schema, the definition covers the basic action and the approval requirement, but leaves the agent without the id source, any expectation about return shape or volume, and no routing versus the list_reports sibling. Adequate but with clear gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 50%: 'reason' is documented in the schema, but 'report_id' has no description there. The brief phrase 'by id' does not tell the agent where to obtain a valid report id (e.g., from salesforce_list_reports) or its format, so the description fails to compensate for the coverage gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource+scope: run a Salesforce report by id and return results. However, it does not differentiate from the obvious sibling salesforce_list_reports, which an agent would need to distinguish this from (list vs run).
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No when-to-use guidance and no mention of alternatives such as salesforce_list_reports (to discover report ids) or salesforce_search. 'Requires user approval' hints at a workflow constraint but does not tell the agent when this tool is the right choice.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
salesforce_searchARead-onlyIdempotent
Search Salesforce by name or id across one or more object types — the same mechanism as the search bar at the top of the Salesforce UI. Returns lightweight Id/Name matches per object type; call salesforce_get_record for full field details on a match. Requires user approval.
| Name | Required | Description | Default |
|---|---|---|---|
| reason | Yes | One sentence: why are you calling this tool right now? | |
| account_id | No | Scope results to this Account's related records (e.g. its Opportunities). Requires object_types to be set, since not every object has an AccountId field. | |
| max_results | No | ||
| search_term | Yes | Name, partial name, or id to search for | |
| object_types | No | Comma-separated Salesforce object API names to restrict the search to, e.g. 'Opportunity,Contact'. Leave empty to search Salesforce's default globally-searchable objects. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover readOnly/idempotent/non-destructive, so the bar is lower. The description adds real context beyond them: it discloses the return shape ('lightweight Id/Name matches per object type') and the operational requirement ('Requires user approval'), which is not captured in the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, front-loaded with what the tool does, then result shape, then the handoff to salesforce_get_record and the approval requirement. No filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description helpfully summarizes the return ('lightweight Id/Name matches per object type') and the approval gate. It is nearly complete, only missing explicit guidance on the account_id scoping or result limits, which the schema partially covers.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 80%, so most parameters are self-documenting; the account_id and object_types scoping rules are well explained in the schema itself. The description only vaguely references 'one or more object types' and 'name or id' without adding syntax or format guidance, so baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Search Salesforce by name or id across one or more object types') and grounds it in a familiar mechanism (the UI search bar). It also implicitly distinguishes itself from salesforce_get_record by describing the lightweight result scope.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly routes the agent onward: 'call salesforce_get_record for full field details on a match.' That gives a clear next-step condition, though it does not state when to prefer this over other Salesforce siblings or when search is inappropriate.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
slack_create_group_chatA
Create (or reopen the existing) group-DM conversation with the given participants and return its channel id, ready for slack_send_message. Participants must already have a Slack user id (from slack_list_dms, slack_list_group_chats, or a message's user_id) -- this does not resolve email addresses or handles. Requires user approval.
| Name | Required | Description | Default |
|---|---|---|---|
| reason | Yes | One sentence: why are you calling this tool right now? | |
| participants | Yes | Comma-separated Slack user IDs to include (at least 2, e.g. 'U123,U456') |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only declare readOnlyHint=false, idempotentHint=false, destructiveHint=false. The description adds beyond that: it returns a channel id usable by slack_send_message, it reuses an existing conversation rather than always creating new, and it requires user approval. Minor tension: 'reopen the existing' reads as idempotent while idempotentHint=false, though this is a soft inconsistency rather than a clean contradiction.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, action front-loaded, then the input constraint, then the approval requirement. No filler; each sentence carries a distinct, necessary fact.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 2-parameter, no-output-schema tool, the description covers the mutation nature, the required input provenance, the return value, the reuse behavior, and the approval gate. Nothing an agent needs to call it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% and both params are documented, so baseline is 3. The description adds real meaning on top of the schema: it clarifies the participants value must be pre-resolved user IDs and explicitly states email addresses and handles are not accepted, which helps an agent avoid a bad call.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource (create/reopen the group-DM conversation) plus the outcome (returns channel id ready for slack_send_message). It positions itself against siblings by naming slack_list_dms, slack_list_group_chats, and slack_send_message as related tools. An agent can place it without opening any other schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear runtime context: you must already have user IDs (obtained from slack_list_dms, slack_list_group_chats, or a message's user_id), it does not resolve emails/handles, and it requires user approval. It also signals the follow-up tool (slack_send_message). It stops short of an explicit when-not-to-use (e.g., use slack_send_message for an existing channel, or that a 1-participant DM is not this tool).
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
slack_get_channel_historyARead-onlyIdempotent
Fetch recent messages in a Slack channel. Returns {messages: [...], has_more: bool}, plus a note when has_more is true -- more messages exist than were returned (a small/inactive channel, or a Slack-imposed cap; see docs/slack-setup.md) -- call again with a larger limit, or narrow the time range, to see the rest instead of assuming this is everything. Requires user approval.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | ||
| reason | Yes | One sentence: why are you calling this tool right now? | |
| channel_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Goes well beyond the readOnly/idempotent annotations by disclosing the return shape, the has_more flag semantics, why truncation happens (small channel or Slack-imposed cap), and that user approval is required. This is exactly the kind of pitfall disclosure annotations cannot carry.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with a short declarative sentence, then one dense sentence carrying the return contract and recovery guidance. The nested em-dash asides make it slightly hard to parse, but nothing is padding.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description usefully documents the return shape and the has_more contract, and notes the approval requirement. It leaves channel_id and the limit default/semantics to the schema, so it is complete enough without being exhaustive.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is only 33% (only 'reason' is documented). The description partially compensates by implying limit behavior ('call again with a larger limit') but says nothing about channel_id or the reason parameter, leaving the low-coverage gap only partly filled. Baseline for partial coverage plus marginal added meaning is a 3.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Fetch) and resource (recent messages in a Slack channel), with scope ('recent') that distinguishes it from slack_get_thread_replies and slack_search_messages. An agent can tell without opening the schema that this is a channel-scoped recent-message read.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives concrete next-step guidance for the truncated case (call again with a larger limit or narrow the time range) rather than assuming completeness. It does not explicitly exclude or route against siblings like slack_search_messages, so it is clear context but short of explicit when-not guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
slack_get_thread_repliesARead-onlyIdempotent
Fetch all replies in a Slack thread. Returns {messages: [...], has_more: bool}, plus a note when has_more is true -- more replies exist than were returned (see slack_get_channel_history's own note on why). Requires user approval.
| Name | Required | Description | Default |
|---|---|---|---|
| reason | Yes | One sentence: why are you calling this tool right now? | |
| thread_ts | Yes | ||
| channel_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, so safety is covered. The description adds useful behavioral context beyond that: the return shape, the has_more flag and its meaning, and the user approval requirement.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, front-loaded with purpose, then return shape, pagination note, and approval requirement. The parenthetical cross-reference is slightly indirect but there is no wasted text.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, so the description usefully supplies the return shape and approval requirement, while annotations cover safety. However, two of three parameters lack semantic detail and no usage alternatives are given, leaving gaps for correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is only 33%; only the reason parameter has a schema description. The description does not explain channel_id, thread_ts, or reason, so it fails to compensate for the low coverage even though the parameter names are fairly self-explanatory.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: fetch all replies in a Slack thread. It distinguishes the operation from channel-history retrieval by scoping to thread replies and references the sibling for pagination context.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description says what the tool fetches but gives no explicit when-to-use guidance, no when-not-to-use guidance, and no alternative tool routing. The only sibling mention is a cross-reference for a note, not an alternative for the agent to choose.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
slack_list_channelsARead-onlyIdempotent
List Slack channels visible to the user (id, name, privacy, topic, purpose, member count). Optionally filter to channels a specific participant belongs to. Auto-approved.
| Name | Required | Description | Default |
|---|---|---|---|
| reason | Yes | One sentence: why are you calling this tool right now? | |
| max_results | No | ||
| participant | No | Filter to channels this participant (user id, handle, or name) belongs to; comma-separated to require all of them as members of the same channel; empty returns all | |
| exclude_archived | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint and destructiveHint, so the safety profile is covered. The description adds real context beyond them: the result is limited to what the user can see, and the call is auto-approved, which matters in a privacyfence-governed tool set. It stops short of describing pagination or the max_results truncation behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, front-loaded with the resource and returned fields, then the optional filter. Nothing is padded, though the trailing 'Auto-approved.' reads as a fragment tacked onto the end rather than integrated into the flow.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With four parameters, 50% schema coverage, and no output schema, the description compensates partially by naming the returned fields. However, max_results and exclude_archived semantics remain undocumented, and archived-channel exclusion is a behavioral default an agent should know before calling.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 50%: reason and participant are documented in the schema, and the description merely restates the participant filter concept without adding syntax. max_results and exclude_archived (default true) carry no explanation in either place, so the description does not compensate for the gap. Baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('List Slack channels visible to the user') and enumerates the fields returned, which lets an agent tell it apart from slack_list_dms and slack_list_group_chats by the resource noun. It never explicitly names those siblings, so the differentiation is implied rather than stated.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description notes the optional participant filter and the 'visible to the user' scope, which implies usage context, but gives no explicit when-to-use guidance or exclusion conditions (e.g., when to prefer slack_list_dms or slack_list_group_chats). 'Auto-approved' hints the call needs no approval but does not explain the workflow.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
slack_list_dmsARead-onlyIdempotent
List 1:1 direct-message conversations visible to the user (id, other participant). Optionally filter to the DM with a specific participant (user id, handle, or display name). Auto-approved.
| Name | Required | Description | Default |
|---|---|---|---|
| reason | Yes | One sentence: why are you calling this tool right now? | |
| max_results | No | ||
| participant | No | Filter to the DM with this participant (user id, handle, or name); empty returns all |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, destructiveHint=false, so the safety profile is covered. The description adds useful context ('visible to the user', 'Auto-approved', which matters in a privacy-fence environment) but says nothing about pagination behavior or the max_results cap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two compact sentences, front-loaded with the resource and scope, then the optional filter. No filler or redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only list tool with annotations covering safety, the description is nearly complete: it names the returned fields even though no output schema exists. Only pagination and max_results behavior are left unaddressed.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 67%, so the participant parameter's accepted forms (user id, handle, display name) are already in the schema, and the description largely repeats them. The reason and max_results parameters are not addressed in the description, so no real value is added beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (List) and resource (1:1 direct-message conversations) with scope (visible to the user) and even names the returned fields. It clearly distinguishes itself from slack_list_channels and slack_list_group_chats by restricting to 1:1 DMs.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied by the resource scope rather than stated: an agent can infer this is for DMs but there is no explicit when-to-use vs slack_list_group_chats/telegram_list_chats, nor any exclusion guidance. The optional participant-filter hint provides some routing value.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
slack_list_group_chatsARead-onlyIdempotent
List group-DM conversations visible to the user (id, name, participants). Optionally filter to group chats containing a specific participant (user id, handle, or display name). Auto-approved.
| Name | Required | Description | Default |
|---|---|---|---|
| reason | Yes | One sentence: why are you calling this tool right now? | |
| max_results | No | ||
| participant | No | Filter to group chats containing this participant (user id, handle, or name); comma-separated to require all of them as members of the same group chat; empty returns all |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint and destructiveHint=false, so the safety profile is covered. The description still adds value beyond that: 'visible to the user' scopes the result set, it names the fields returned, and 'Auto-approved' discloses that no approval/policy gate is triggered. It does not discuss pagination or max_results behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, front-loaded with the operation and scope, then the optional filter, then the approval note. Every sentence carries distinct information with no padding.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema and a read-only, three-parameter tool, the description covers the resource, the returned fields, the visibility scope, and the optional filter. The only real gap is the meaning of max_results / result capping, which is undocumented in both schema and description.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 67%: participant and reason are documented in-schema, max_results is not. The description restates the participant filter semantics (user id, handle, or display name) but adds nothing beyond the schema, and nothing compensates for the undocumented max_results/pagination parameter. Baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description gives a specific verb and resource ('List group-DM conversations visible to the user') and even enumerates the returned fields (id, name, participants). The 'group-DM' qualifier inherently separates it from slack_list_dms and slack_list_channels, but no sibling is named explicitly, so it stops short of full differentiation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It states that the participant filter is optional and what it accepts, which implies the primary use case (browse group DMs, or narrow to one member). However it never says when to prefer this over slack_list_dms, slack_list_channels, or slack_search_messages, and gives no exclusions or prerequisites.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
slack_refresh_channel_cacheARead-onlyIdempotent
Force an immediate refresh of PrivacyFence's local cache of Slack channel/DM/group-DM names, used to resolve which conversation a message belongs to in search results and history/thread reads without a per-message conversations.info call. Refreshes automatically about once a week; call this after a new channel is created so it resolves by name right away. On a workspace with a lot of channels, one call may not finish the whole sync -- check the result's has_more flag and, if true, call this tool again (same args) to continue from where it left off. Auto-approved -- refreshes name lookups only, reads no message content.
| Name | Required | Description | Default |
|---|---|---|---|
| reason | Yes | One sentence: why are you calling this tool right now? |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover readOnly/idempotent/non-destructive, yet the description adds real context beyond them: 'Auto-approved', 'refreshes name lookups only', 'reads no message content', and the pagination behavior via has_more requiring a repeat call. This is exactly the value-add behavioral detail annotations cannot express.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the verb and purpose, and each sentence (usage timing, pagination, safety) earns its place. It runs slightly long for a single-parameter tool, keeping it just short of a 5.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema, but the description compensates by explaining the relevant return signal (has_more) and how to act on it. For a one-param, auto-approved cache refresh, nothing an agent needs to invoke it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Single required 'reason' parameter with 100% schema description coverage, so the schema already documents it fully. The description adds the pagination hint ('call this tool again (same args)') but nothing about the reason argument itself, so the baseline 3 for fully-covered params applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Force an immediate refresh of PrivacyFence's local cache of Slack channel/DM/group-DM names') and immediately explains the downstream purpose: resolving which conversation a message belongs to without a per-message conversations.info call. This distinguishes it cleanly from siblings like slack_list_channels and slack_get_channel_history.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly states the alternative baseline ('Refreshes automatically about once a week') and the condition that warrants a manual call ('call this after a new channel is created so it resolves by name right away'). It also gives the continuation condition using the has_more flag.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
slack_refresh_user_cacheARead-onlyIdempotent
Force an immediate refresh of PrivacyFence's local cache of Slack workspace member names/emails, used to resolve message authors in channel history, thread replies, and search results without a per-message users.info call. Refreshes automatically about once a week; call this when a teammate who joined recently isn't resolving correctly yet. Auto-approved -- refreshes name/email lookups only, reads no message content.
| Name | Required | Description | Default |
|---|---|---|---|
| reason | Yes | One sentence: why are you calling this tool right now? |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the annotations (readOnly/idempotent/non-destructive), the description discloses auto-refresh cadence, that the call is auto-approved, and precisely what is touched (name/email lookups only, no message content read). This directly reassures an agent weighing a cache-mutating operation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, each carrying real information (purpose, cadence/trigger, safety scope), with the core action front-loaded. Dense but no obvious filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, and none is needed for a refresh trigger. With a single fully-documented parameter and annotations covering the safety profile, the description supplies everything an agent needs to decide and invoke correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Only one parameter, and its schema description ("One sentence: why are you calling this tool right now?") has 100% coverage. The description adds no syntax or format guidance for the reason field, so the baseline 3 applies when the schema fully documents the parameter.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource: force refresh of the local Slack workspace member name/email cache. It explains the downstream purpose (resolving message authors in history/threads/search without a per-message users.info call), which distinguishes it from the sibling slack_refresh_channel_cache by scope.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives an explicit trigger condition ("call this when a teammate who joined recently isn't resolving correctly yet") and notes the automatic weekly cadence so the agent knows when manual invocation is warranted. It does not explicitly cross-reference slack_refresh_channel_cache as the alternative for channel data, so it stops just short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
slack_resolve_permalinkARead-onlyIdempotent
Parse a Slack message permalink (from a message's "Copy link") into the channel id, timestamp, and (if the link points at a threaded reply) thread root timestamp needed by slack_get_channel_history/slack_get_thread_replies. Reads no message content -- just decodes the link. Auto-approved.
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | A Slack message permalink, e.g. https://workspace.slack.com/archives/C0123/p1700000000123456 | |
| reason | Yes | One sentence: why are you calling this tool right now? |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover the safety profile (readOnly, idempotent, non-destructive). The description adds genuinely new behavioral context: it reads no message content, it merely decodes, and it is auto-approved (no approval friction). That exceeds what the annotations convey.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core action and its outputs, then two short clarifying sentences with no filler. Slightly dense in the first sentence but every clause earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description compensates by naming the exact return values (channel id, timestamp, thread root timestamp for replies). Combined with the schema-documented params and safety annotations, an agent has everything needed to call and interpret this tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents both url and reason. The description implies the permalink input and clarifies the decoded outputs, but adds little syntax or format detail beyond the schema. Baseline 3 is appropriate when structured fields carry the parameter load.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (parse/decode) plus resource (Slack message permalink) and enumerates the outputs (channel id, timestamp, optional thread root timestamp). It explicitly names the sibling tools that consume the result, so an agent can distinguish it from slack_get_channel_history and slack_get_thread_replies without opening schemas.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description makes the use case clear: convert a permalink into the IDs needed by slack_get_channel_history/slack_get_thread_replies. It also scopes what it does not do ('Reads no message content'). It stops short of explicit when-not-to-use or alternative-tool routing, but the context is strong.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
slack_search_messagesARead-onlyIdempotent
Search Slack messages matching a query, a participant, or both. Prefer participant (a user id, handle, or display name) over a text-only query when looking for messages from or with someone -- e.g. 'Bob wrote me' is participant='Bob'; 'Bob in a chat with Jane' is participant='Bob,Jane' -- it reads the matching DM/group-chat conversation(s) directly instead of relying on Slack's search index, which is more reliable for participant-based lookups. Combine with query to also filter those conversations' text. Defaults to the last 90 days (about 3 months) so results on a workspace with long history aren't dominated by old, no-longer-relevant matches; widen or disable via days. Requires user approval.
| Name | Required | Description | Default |
|---|---|---|---|
| days | No | Only include messages from the last N days (default 90, ~3 months); set to 0 to search all history with no cutoff | |
| count | No | ||
| query | No | ||
| reason | Yes | One sentence: why are you calling this tool right now? | |
| participant | No | Filter to the DM/group-chat conversation(s) with this participant (user id, handle, or name); comma-separated to match a group chat containing all of them |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, and non-destructive behavior. The description adds valuable context beyond that: it discloses that participant mode reads DM/group-chat conversations directly rather than using Slack's search index, that the default 90-day window exists to keep results relevant, and that the tool requires user approval.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with a clear purpose statement, then flows into usage guidance and defaults. It is somewhat long but every sentence adds information; the examples and rationale for participant mode earn their space.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 5-parameter search tool with no output schema and annotations covering safety, the description covers the main usage patterns, approval requirement, and default date window. It does not explain the count parameter or result format, but these are minor gaps against the annotations and schema that already exist.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 60%. The description duplicates the participant explanation already in the schema and adds only minor semantic value for query (how it filters conversations' text) and days (rationale for the default). The count parameter is not described in either place, leaving a gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb (search) and resource (Slack messages) plus matching criteria (query, participant, or both). It does not explicitly contrast with sibling tools like slack_get_channel_history or slack_list_dms, so it lacks full sibling differentiation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives clear when-to-use guidance for participant vs query with concrete examples ('Bob wrote me' → participant='Bob'), and explains how to combine them. It does not state when not to use this tool or name alternative tools for channel-scoped searches.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
slack_send_messageA
Send a message to a Slack channel or DM. Requires user approval. Set mark_unread=true to leave the message unread after sending (useful when sending a DM to yourself as a note; requires the im:write scope on the user token for DMs).
| Name | Required | Description | Default |
|---|---|---|---|
| text | Yes | ||
| reason | Yes | One sentence: why are you calling this tool right now? | |
| thread_ts | No | ||
| channel_id | Yes | ||
| mark_unread | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations give the safety profile (readOnlyHint=false, destructiveHint=false, idempotentHint=false), and the description adds genuinely new context: human approval is required before the message goes out, and DM self-notes require the im:write scope. Those are behavioral traits not derivable from the annotations. It stops short of saying whether the message is immediately visible to recipients or how failures surface.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight sentences with the core action front-loaded, followed by the approval prerequisite and the one non-obvious parameter. No filler or restated boilerplate.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema, so return behavior would need coverage, and the tool is a non-idempotent mutation where the approval flow matters. The description covers approval and mark_unread well, but omits thread_ts semantics and what a successful send returns, leaving gaps for a 5-param write tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is only 20% (only 'reason' is documented), so the description must compensate. It does a good job on mark_unread, explaining the effect, the use case, and the scope requirement, but says nothing about thread_ts (threaded reply behavior) or channel_id vs DM addressing, leaving two of five parameters unexplained anywhere.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: 'Send a message to a Slack channel or DM.' An agent can immediately tell this is the outbound-message tool, distinct from slack_get_channel_history or slack_search_messages. It does not, however, explicitly contrast itself with sibling write tools like telegram_send_message or slack_create_group_chat.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
'Requires user approval' and the mark_unread-as-self-note guidance give real context for calling it. But there is no when-not guidance, no indication of which channel types are unsupported, and no routing between this and other messaging tools. Usage is implied rather than prescribed.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
tasks_complete_taskA
Mark a task as completed. Requires user approval.
| Name | Required | Description | Default |
|---|---|---|---|
| reason | Yes | One sentence: why are you calling this tool right now? | |
| task_id | Yes | ||
| task_list_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The annotations already declare readOnlyHint=false, idempotentHint=false, and destructiveHint=false, so the safety profile is covered. The description adds a useful approval prerequisite, but it does not explain side effects, reversibility, error behavior, or what changes after completion beyond what the annotations imply.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two short sentences, front-loads the purpose, and contains no filler. The approval requirement follows immediately and is easy to retain.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a three-required-parameter mutation with no output schema, the description states the core action and the approval prerequisite, and the annotations cover the safety profile. However, it omits alternative-tool guidance and does not compensate for the undocumented task_id/task_list_id parameters, leaving it only partially complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is only 33%, with just the 'reason' parameter described in the schema. The description adds no information about task_list_id or task_id, so the two undocumented required parameters remain unclarified and the description does not compensate for the low schema coverage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: 'Mark a task as completed.' This clearly distinguishes the operation from siblings such as tasks_uncomplete_task, tasks_update_task, and tasks_move_task, even though it does not explicitly name them.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It provides one important usage condition, 'Requires user approval,' but does not explain when to use this tool versus alternatives like tasks_update_task or tasks_uncomplete_task. Usage is implied by the operation name, leaving only minimal explicit guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
tasks_create_taskC
Create a new task. Requires user approval.
| Name | Required | Description | Default |
|---|---|---|---|
| due | No | Due date in RFC 3339 format | |
| notes | No | ||
| title | Yes | ||
| reason | Yes | One sentence: why are you calling this tool right now? | |
| task_list_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=false, idempotentHint=false, and destructiveHint=false, so the mutation profile is covered. The description adds one genuinely useful trait beyond annotations: the user-approval requirement. It says nothing about duplicate handling or the effect of the required 'reason' field, so it only partially extends the structured data.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short, front-loaded sentences with no filler; the purpose comes first and the constraint second. It is efficient, though the second sentence could carry more useful detail at the same length.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a non-idempotent mutation with no output schema and 40% schema coverage, the description omits what an agent most needs: which fields are mandatory (task_list_id, title, reason) and what the approval flow implies. Annotations cover the safety profile but not the argument requirements.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is only 40% (five parameters, three required), so the description must compensate and does not: it never mentions task_list_id, title, reason, due, or notes. The 'reason' parameter's self-justification semantics are left entirely to the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Create) and resource (task), so the operation is unambiguous and distinguishable from siblings like tasks_update_task or tasks_get_task. It does not explicitly contrast itself with those siblings, keeping it short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The only guidance is 'Requires user approval', which is a prerequisite rather than a when-to-use statement. There is no mention of when to prefer tasks_create_task over siblings such as tasks_update_task or tasks_move_task, and no context for what a valid invocation looks like.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
tasks_get_taskBRead-onlyIdempotent
Fetch a single task by id. Auto-approved.
| Name | Required | Description | Default |
|---|---|---|---|
| reason | Yes | One sentence: why are you calling this tool right now? | |
| task_id | Yes | ||
| task_list_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint and destructiveHint=false, so safety is covered by structured data. The description adds one genuine behavioral fact beyond annotations — that the call is auto-approved and needs no approval flow — but says nothing about return shape or error behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two very short, front-loaded sentences with no filler; the operation comes first and the approval note second. Efficient, though 'Auto-approved' is so terse it borders on cryptic.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple read tool with annotations present and no output schema, the definition is minimally adequate but leaves two of three parameters unexplained and gives no hint about what the returned task contains.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is only 33% — task_id and task_list_id have no descriptions anywhere. The phrase 'by id' is ambiguous when the schema requires two distinct ids, so the description does not compensate for the gap; only the reason parameter is explained, and by the schema, not the description.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Fetch') and resource ('a single task') and the lookup key ('by id'), so it is clearly distinguishable from the sibling tasks_list_tasks. It falls short of 5 only because it does not name the sibling or clarify that the 'id' requires both a task id and a task list id.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
"Auto-approved" gestures at the approval context but gives no when-to-use guidance, no prerequisites, and no routing to alternatives such as tasks_list_tasks for browsing. The agent must infer usage entirely.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
tasks_list_task_listsBRead-onlyIdempotent
List all Google Task lists. Auto-approved.
| Name | Required | Description | Default |
|---|---|---|---|
| reason | Yes | One sentence: why are you calling this tool right now? |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint and destructiveHint, so the safety profile is covered. 'Auto-approved' adds a small piece of behavioral context (no approval gate), but the description says nothing about pagination, result size, or auth requirements beyond what annotations imply.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two very short sentences, front-loaded with the resource and action, with no filler. The 'Auto-approved' fragment is terse rather than wasteful, though it is somewhat cryptic.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple read-only list operation with a fully documented single parameter and annotations covering the safety profile, the description covers what is needed. No output schema exists, but the return is self-evident for a list tool, so little is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% with a single 'reason' parameter fully documented in the schema, so the description need not explain it. Baseline 3 applies; the description adds no parameter meaning beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (List) and resource (Google Task lists), which is readily distinguishable from the sibling tasks_list_tasks (which lists tasks, not lists). It does not explicitly name a sibling or contrast scope, so it falls short of the top band.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no when-to-use context, no prerequisites, and no exclusions. 'Auto-approved' is an operational note, not guidance about when this tool is the right choice over tasks_list_tasks or other task tools.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
tasks_list_tasksCRead-onlyIdempotent
List tasks in a task list. Auto-approved.
| Name | Required | Description | Default |
|---|---|---|---|
| reason | Yes | One sentence: why are you calling this tool right now? | |
| task_list_id | Yes | ||
| show_completed | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, so the safety profile is covered. "Auto-approved" adds a small piece of behavioral context about the approval flow, but nothing about pagination, return volume, or filtering behavior is disclosed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Only two short sentences with no filler, but the brevity comes at the cost of specification rather than being tight and complete. "Auto-approved" reads as a note rather than front-loaded useful information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 3-parameter tool with no output schema and 33% coverage, the description should clarify the `show_completed` filter and the required `reason` parameter. It leaves an agent unable to know how to configure results or why `reason` is mandatory.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is only 33% — only `reason` is documented. The description fails to compensate: it never explains `show_completed` (a non-obvious filter with a default) or clarify `task_list_id` beyond the obvious. Two of three parameters carry meaning that is undocumented in both places.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource ("List tasks in a task list"), but it largely restates the tool name tasks_list_tasks. It does not differentiate meaningfully from siblings like tasks_list_task_lists or tasks_get_task beyond the obvious resource noun.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance on when to use this versus tasks_list_task_lists, tasks_get_task, or other task tools. There is no mention of prerequisites or when-not-to-use, only a bare statement of function.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
tasks_move_taskC
Move a task from one list to another. Requires user approval.
| Name | Required | Description | Default |
|---|---|---|---|
| reason | Yes | One sentence: why are you calling this tool right now? | |
| task_id | Yes | ||
| source_list_id | Yes | ||
| destination_list_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already establish this is a mutating, non-idempotent, non-destructive operation, and the description usefully adds that user approval is required. However, it does not explain what approval entails (async vs. blocking), nor what happens to the task's other properties on move.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short, front-loaded sentences with no filler; the primary action comes first. Efficient, though terse to the point of omitting needed detail.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a mutating 4-parameter tool with 25% schema coverage and no output schema, the description is too thin – it omits the approval mechanics, the move's effect on task metadata, and any return/confirmation behavior an agent would need.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is only 25% – just 'reason' is documented. The three ID parameters (task_id, source_list_id, destination_list_id) have no descriptions, and the description does not compensate by clarifying their expected format or relationships.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Move a task from one list to another'), which is clearly distinguishable from siblings like tasks_update_task or tasks_complete_task. It doesn't explicitly name alternatives, but the action itself is unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Only a bare precondition ('Requires user approval') is given; there is no guidance on when to choose this over tasks_update_task (e.g., changing list membership vs. editing fields). No when-not guidance is present.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
tasks_uncomplete_taskB
Mark a task as not completed. Requires user approval.
| Name | Required | Description | Default |
|---|---|---|---|
| reason | Yes | One sentence: why are you calling this tool right now? | |
| task_id | Yes | ||
| task_list_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare readOnlyHint=false, destructiveHint=false, idempotentHint=false. The description adds the approval requirement, which is useful behavioral context beyond annotations, but it does not explain side effects, reversibility, or the non-idempotent nature.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences, front-loaded with the action and then the approval condition. Every word earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a mutation tool with three required parameters, low schema coverage, and no output schema, the description is incomplete: it omits the meanings of task_list_id and task_id and provides no output or error context. The approval note is helpful but insufficient.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is only 33% (only the reason parameter has a description). The description mentions no parameters, so it fails to compensate for the undocumented task_list_id and task_id parameters.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource ('Mark a task as not completed'), which is clear. It does not explicitly differentiate itself from sibling tools like tasks_complete_task or tasks_update_task, so it stops short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives one usage prerequisite ('Requires user approval') but no when-to-use vs when-not guidance or named alternatives. The agent must infer the appropriate context from the tool name and surrounding tools.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
tasks_update_taskB
Update a task's title, notes, or due date. Requires user approval.
| Name | Required | Description | Default |
|---|---|---|---|
| due | No | ||
| notes | No | ||
| title | No | ||
| reason | Yes | One sentence: why are you calling this tool right now? | |
| task_id | Yes | ||
| task_list_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=false, idempotentHint=false, destructiveHint=false, so mutation/non-idempotence is known. The description adds genuinely new context with 'Requires user approval', but leaves open how approval is obtained, whether unspecified fields are preserved, and whether the update can be applied to nonexistent tasks.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences, the mutation scope is front-loaded, and the approval caveat follows immediately. No filler words or repetition of the tool name.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 6-parameter mutation tool with 17% schema coverage and no output schema, the description covers the field set and the approval gating but omits identifier semantics and rate/pagination behavior. Adequate but with clear gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is only 17% (just the reason field), so the description has to compensate, and it partially does by naming the title/notes/due fields. However it says nothing about task_id, task_list_id, or the reason parameter's purpose, leaving half the parameter surface unexplained in both schema and description.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (update) and resource (task) plus the three editable fields, which distinguishes it from siblings like tasks_complete_task and tasks_move_task. It does not name any sibling explicitly, but the operation is unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No when-to-use or when-not-to-use guidance, and no comparison against alternatives such as tasks_move_task or tasks_complete_task. The only usage-adjacent statement is 'Requires user approval', which is a prerequisite, not a selection rule.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
telegram_get_messagesCRead-onlyIdempotent
Fetch recent messages from a Telegram chat by chat id. Requires user approval.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | ||
| reason | Yes | One sentence: why are you calling this tool right now? | |
| chat_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true and destructiveHint=false, so the safety profile is covered. The description adds a genuinely useful workflow constraint (human approval required) but says nothing about ordering, pagination, or how many messages come back by default.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences, front-loaded with the action and resource before the approval caveat. Efficient, with no filler text.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 3-parameter read tool with no output schema and 33% schema coverage, the description leaves too much open: no return shape, no pagination/limit behavior, and no ordering semantics. The approval note is helpful but the definition is thin overall.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is only 33% and only chat_id is reflected in the description. The `limit` parameter (default 50) is undocumented in both the description and the schema, and `reason` semantics live only in the schema. The description does not compensate for the coverage gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Fetch) and resource (recent messages from a Telegram chat) qualified by chat id. It does not, however, distinguish itself from the sibling telegram_search_messages, leaving the retrieve-recent vs. search distinction to inference.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The only guidance is 'Requires user approval,' which is a precondition, not a when-to-use rule. Nothing tells the agent when to pick this over telegram_search_messages or telegram_list_chats.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
telegram_list_chatsBRead-onlyIdempotent
List Telegram chats (id, name, type, unread count). Auto-approved.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | ||
| reason | Yes | One sentence: why are you calling this tool right now? |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, and destructiveHint=false, so the safety profile is covered structurally. The description adds the useful fact that the call is auto-approved (no approval gate), but says nothing about pagination, ordering, or how `limit` bounds the result.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two compact, front-loaded clauses with no filler. It is terse rather than padded, though the parenthetical could be spent on parameter meaning instead.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only list tool with annotations covering safety and no output schema, naming the returned fields is a reasonable substitute for a return description. However, the undocumented `limit` parameter leaves a real gap for an agent deciding how to scope the call.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is only 50% and the undocumented parameter is `limit` (default 50), which is exactly the one that needs meaning. The description lists output fields rather than explaining what `limit` caps or how `reason` is consumed, so it fails to compensate for the coverage gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('List Telegram chats') and even enumerates the fields returned. It is clear what the tool does, but it does not distinguish itself from siblings like telegram_search_messages or telegram_refresh_chat_cache.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to use this list versus telegram_get_messages or telegram_search_messages. 'Auto-approved' describes an approval property, not a usage condition, so the agent gets no routing help.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
telegram_refresh_chat_cacheARead-onlyIdempotent
Force an immediate refresh of PrivacyFence's local cache of Telegram chat/group/channel names, used to resolve which chat a message belongs to in search results and chat history without a per-message lookup. Refreshes automatically about once a week; call this after a new chat starts so it resolves by name right away. Auto-approved -- refreshes name lookups only, reads no message content.
| Name | Required | Description | Default |
|---|---|---|---|
| reason | Yes | One sentence: why are you calling this tool right now? |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Goes beyond the readOnly/idempotent/non-destructive annotations by disclosing that it is auto-approved and that it reads only name lookups with no message content. This approval and data-scope context is exactly what an agent needs and is not in the structured fields.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, front-loaded with the action and its purpose, followed by the timing condition and safety note. No filler or repetition.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a single-parameter, read-only refresh tool with no output schema, the description covers purpose, timing, approval behavior, and data scope. Nothing an agent needs to call it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
There is one parameter with 100% schema description coverage, so the schema already documents 'reason' fully. The description adds no further parameter-level meaning, so the baseline of 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (refresh) and resource (PrivacyFence's local cache of Telegram chat names), and explains the downstream purpose: resolving chat names in search/history without per-message lookups. This clearly separates it from telegram_list_chats and slack_refresh_channel_cache.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives an explicit trigger ('call this after a new chat starts so it resolves by name right away') and implies when it is unnecessary ('refreshes automatically about once a week'). No named alternative tool is given, but the triggering condition is clear.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
telegram_search_messagesBRead-onlyIdempotent
Search messages across Telegram chats by keyword. Requires user approval.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | ||
| query | Yes | ||
| reason | Yes | One sentence: why are you calling this tool right now? |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint and destructiveHint=false, so safety is covered. The description's value-add is the approval gate, which is genuine behavioral context, but it says nothing about what denial looks like or pagination/result limits.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences, zero filler, with the core action front-loaded and the approval constraint immediately after.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only search with no output schema, the description covers the action and the approval requirement but leaves the return shape, result count behavior, and two of three parameters unexplained. Minimum viable, not complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is only 33%: query and limit are undocumented while reason has a description. The phrase 'by keyword' hints at query semantics, but the description adds essentially nothing about the limit parameter or the required reason string, so it does not compensate for the coverage gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (search) plus resource (messages) and scope (across Telegram chats by keyword), which cleanly separates it from telegram_get_messages and telegram_list_chats. It lacks any explicit naming of those siblings, so it stops short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Only guidance is the prerequisite 'Requires user approval'; there is no statement of when to reach for this versus telegram_get_messages or slack_search_messages, and no mention of when keyword search is inappropriate.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
telegram_send_messageB
Send a message to a Telegram chat or user by chat id. Requires user approval.
| Name | Required | Description | Default |
|---|---|---|---|
| text | Yes | ||
| reason | Yes | One sentence: why are you calling this tool right now? | |
| chat_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=false, idempotentHint=false, and destructiveHint=false, so the agent knows this is a non-idempotent write that is not destructive. The description adds a genuinely useful behavioral fact, the user-approval gate, but omits return behavior, rate limits, and what errors look like.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences, front-loaded with the action and identifier, with the approval note appended. Efficient and free of filler, though both sentences are quite terse relative to the tool's mutation risk.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a write tool with no output schema and 33% schema coverage, the description covers the action and approval requirement but leaves parameter meanings and failure/return behavior under-specified. Adequate but with clear gaps given the low parameter documentation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is only 33% (just the 'reason' parameter is documented in the schema). The description partially compensates by clarifying that chat_id resolves to 'a Telegram chat or user', but gives no meaning for the 'text' body parameter and no format hints for chat_id.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (send) and resource (message to a Telegram chat/user) with the key identifier (chat id). It does not explicitly contrast with the sibling slack_send_message, but the 'Telegram' qualifier makes the target system unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No when-to-use vs alternatives guidance and no exclusions. The only encoded context is the approval prerequisite ('Requires user approval'), which is a precondition rather than routing guidance for choosing this tool over slack_send_message or telegram_get_messages.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
122 tool updates
v5.3.0- First observed
apps_script_get_content - First observed
apps_script_get_execution_log - First observed
apps_script_list_projects - First observed
apps_script_write_content - First observed
calendar_create_event - First observed
calendar_create_out_of_office - First observed
calendar_delete_event - First observed
calendar_get_event_details - First observed
calendar_get_event_visibility - First observed
calendar_get_free_busy - First observed
calendar_list_calendars - First observed
calendar_list_colors - First observed
calendar_list_events - First observed
calendar_list_rooms - First observed
calendar_set_event_color - First observed
calendar_set_event_visibility - First observed
calendar_set_working_location - First observed
calendar_update_event - First observed
confluence_cql_search - First observed
confluence_create_page - First observed
confluence_download_attachment - First observed
confluence_get_page - First observed
confluence_get_page_by_title - First observed
confluence_list_attachments - First observed
confluence_list_pages - First observed
confluence_list_spaces - First observed
confluence_search - First observed
confluence_update_page - First observed
contacts_add_label - First observed
contacts_create - First observed
contacts_get - First observed
contacts_list - First observed
contacts_remove_label - First observed
contacts_search - First observed
contacts_update - First observed
drive_add_comment - First observed
drive_create_blank_file - First observed
drive_docs_edit_content - First observed
drive_docs_format_content - First observed
drive_download_file - First observed
drive_get_file_content - First observed
drive_get_file_metadata - First observed
drive_list_files - First observed
drive_list_folder - First observed
drive_list_shared_drives - First observed
drive_move_file - First observed
drive_sheets_add_sheet - First observed
drive_sheets_create - First observed
drive_sheets_delete_dimensions - First observed
drive_sheets_format_range - First observed
drive_sheets_get_metadata - First observed
drive_sheets_get_values - First observed
drive_sheets_insert_dimensions - First observed
drive_sheets_rename_sheet - First observed
drive_sheets_write_range - First observed
drive_upload_file - First observed
drive_write_doc_content - First observed
drive_write_file_content - First observed
gmail_add_label - First observed
gmail_archive_message - First observed
gmail_create_draft - First observed
gmail_create_draft_with_attachments - First observed
gmail_create_filter - First observed
gmail_create_label - First observed
gmail_download_attachment - First observed
gmail_get_message - First observed
gmail_get_thread - First observed
gmail_list_filters - First observed
gmail_list_labels - First observed
gmail_list_message_attachments - First observed
gmail_list_messages - First observed
gmail_list_threads - First observed
gmail_remove_label - First observed
gmail_reply_all_draft - First observed
gmail_reply_all_draft_with_attachments - First observed
gmail_reply_draft - First observed
gmail_reply_draft_with_attachments - First observed
gmail_update_filter - First observed
jira_add_comment - First observed
jira_create_issue - First observed
jira_get_issue - First observed
jira_get_transitions - First observed
jira_list_projects - First observed
jira_search_issues - First observed
jira_transition_issue - First observed
jira_update_issue - First observed
privacyfence_await_approval - First observed
privacyfence_begin_unattended_session - First observed
privacyfence_check_policy - First observed
privacyfence_create_upload_slot - First observed
privacyfence_end_unattended_session - First observed
privacyfence_list_policy - First observed
privacyfence_propose_policy_change - First observed
privacyfence_status - First observed
salesforce_get_record - First observed
salesforce_list_reports - First observed
salesforce_run_report - First observed
salesforce_search - First observed
slack_create_group_chat - First observed
slack_get_channel_history - First observed
slack_get_thread_replies - First observed
slack_list_channels - First observed
slack_list_dms - First observed
slack_list_group_chats - First observed
slack_refresh_channel_cache - First observed
slack_refresh_user_cache - First observed
slack_resolve_permalink - First observed
slack_search_messages - First observed
slack_send_message - First observed
tasks_complete_task - First observed
tasks_create_task - First observed
tasks_get_task - First observed
tasks_list_task_lists - First observed
tasks_list_tasks - First observed
tasks_move_task - First observed
tasks_uncomplete_task - First observed
tasks_update_task - First observed
telegram_get_messages - First observed
telegram_list_chats - First observed
telegram_refresh_chat_cache - First observed
telegram_search_messages - First observed
telegram_send_message
TDQS
Scored across 122 tools
Tools are namespaced per service (gmail_, drive_, slack_, calendar_, etc.) and the descriptions go out of their way to distinguish near-neighbors (e.g. drive_get_file_content vs drive_download_file, calendar_get_event_details vs calendar_get_event_visibility, drive_write_doc_content vs drive_docs_edit_content vs drive_docs_format_content). The main confusion risk is the combinatorial Gmail draft family (create/reply/reply_all each duplicated with a _with_attachments twin), where selection hinges on a single parameter rather than a distinct purpose.
Nearly everything is snake_case verb_noun with a stable service prefix, which makes the set predictable to scan. Minor deviations exist in prefix depth (drive_ vs drive_sheets_ vs apps_script_) and a few noun-first names (calendar_get_free_busy, slack_resolve_permalink), but nothing approaches mixed conventions.
122 tools is far past any comfortable surface for one server and bloats agent context, with clear combinatorial duplication (six Gmail draft variants, three sheets dimension tools) rather than genuinely distinct operations. It is partly excused by spanning ~12 services, but the total is still excessive relative to the scope.
Coverage across the domains is broad and deliberately gated: read/write/create/update exist for Calendar, Drive, Docs, Sheets, Gmail, Slack, Jira, Tasks, Contacts, Confluence, and Salesforce, with intentional read-only or draft-only surfaces (Gmail drafts instead of send). Known holes are explicitly documented (no contact deletion, no Sheets tab deletion, Salesforce read-only), which agents can work around.
Maintenance
Related MCP Connectors
Zero-secret MCP gateway for AI agents: risk-scored, audited calls with human-in-the-loop approval.
Zero-setup MCP gateway securely connecting AI to your tools with authentication and workflows
Security gateway for AI agents: policy, approval, and audited execution, no secrets shared.
MCP server connecting AI agents to 100+ apps (Gmail, Slack, Notion, GitHub) via one-click OAuth.
Related MCP Servers
- AlicenseNot gradedqualityDmaintenanceHuman-in-the-Loop authorization gateway for AI Agents. Securely pause MCP workflows and route high-risk actions to human approvers via Slack or Email.52 npm1MIT

mcpgateofficial
FlicenseNot gradedqualityCmaintenanceSelf-hosted MCP gateway that connects Claude, ChatGPT, and other AI agents to 20+ enterprise tools (GitLab, Jira, Notion, Google Workspace, Slack, Grafana, …) with OAuth, audit logs, and zero data leaving your infrastructure-
AllMCPofficial
AlicenseNot gradedqualityBmaintenanceOpen-source MCP hub providing a single endpoint for AI agents to access dozens of business integrations (CRMs, spreadsheets, telephony, ads) with multi-tenancy, OAuth, and context-efficient tool discovery.Apache 2.0- FlicenseNot gradedqualityBmaintenanceSelf-hosted MCP server that filters sensitive data from email and calendar before it reaches AI assistants, masking or withholding confidential content based on user policy.-