Kaption WhatsApp
Server Details
Search, summarize and organize your WhatsApp chats from Claude, ChatGPT or Cursor: contacts, groups, labels, notes, reminders and scheduled messages. Tool calls run in the Kaption extension for WhatsApp Web (Chrome, Edge) or the Kaption desktop app, which must be open. Sign-in is OAuth 2.1 with a code sent on WhatsApp. No send-now tool. Setup: https://kaptionai.com/mcp/
- Status
- Healthy
- OAuth
- Works in Glama
- Last Tested
- Transport
- Streamable HTTP · MCP 2025-11-25
- URL
TDQS
Scored across 17 tools
The general-purpose 'query' tool overlaps heavily with specialized tools like list_contacts, get_contact, list_groups, get_group, and get_contact_groups; for example, finding a contact's groups can be done via get_contact_groups or query with include_participants. get_analytics also duplicates export_contacts and export_chat sections. This creates unclear boundaries for common tasks.
Most tools follow a consistent verb_noun pattern (list_contacts, get_group, manage_chat), but 'call_recordings' is a noun phrase and 'query' is a bare verb, which are minor deviations from the otherwise predictable naming.
17 tools cover a broad WhatsApp management and analytics domain; slightly above the ideal 3-15 range but each functional area (contacts, groups, calls, analytics, management, reminders) has a reasonable set of tools.
Core read and management operations are well-covered, but there is no tool to send an immediate message (only scheduled messages) and no message deletion, forwarding, or reaction tools. This is a notable gap for a WhatsApp server.
Available Tools
17 toolscall_recordingsRead call recordingsARead-onlyIdempotentInspect
Read WhatsApp call recordings made by the Kaption extension and their transcripts, with every word timed. Each recording names its conversation (the other person, or the group for a group call), so it links to query, get_contact and get_group. Recordings of locked chats are never returned.
Actions: list - Recordings, newest first (optional conversation_id, date_from, date_to, limit) get - One recording and its transcript: turns by speaker ("you" or "contact") with start/end seconds (requires id; include_words adds each word's timing) search - Recordings whose name or transcript matches the search text, with the matching lines (requires search; same filters as list)
Examples: Recent calls: { action: "list", limit: 10 } Calls with one person: { action: "list", conversation_id: "5491157390064@c.us" } Read a transcript: { action: "get", id: "1790000000000-a1b2c3d4" } What was said about the budget: { action: "search", search: "budget" }
| Name | Required | Description | Default |
|---|---|---|---|
| id | No | Recording ID (required for get) | |
| limit | No | Max recordings to return (default 20, max 100) | |
| action | Yes | Call recordings action to perform | |
| search | No | Text to find in names and transcripts (required for search) | |
| date_to | No | Only recordings started on or before this date (ISO 8601) | |
| date_from | No | Only recordings started on or after this date (ISO 8601) | |
| include_words | No | For get: include each word with its start and end in seconds (default false) | |
| target_session | No | Session ID for multi-account routing | |
| conversation_id | No | Only recordings of this conversation (a contact or group ID) |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | No | The JSON-compatible result returned by the Kaption extension |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint and destructiveHint=false, so the safety profile is covered. Beyond that the description adds real behavioral facts: locked chats are never returned, list ordering is newest-first, and get returns turns by speaker with start/end seconds plus optional per-word timing. It does not mention paging behavior or what an empty result means.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded one-line purpose, then a compact action table and four one-line examples. Every sentence carries either a capability, a scope limit, or a usage cue; nothing is padding.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 9-parameter, 3-action tool with an output schema and full annotations, the description supplies everything an agent needs: action semantics, required inputs per action, return shape, ordering, and the locked-chat exclusion. Return-value details are correctly delegated to the output schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the description adds cross-parameter conditional semantics the per-field schema text only hints at: which filters apply to which action, that search reuses list's filters, and that include_words only affects get. It does not explain target_session, the one parameter that is purely schema-documented.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource (read WhatsApp call recordings and their timed transcripts) and then decomposes the tool into three concretely named actions with distinct outputs. It also distinguishes itself from siblings by naming where recordings link (query, get_contact, get_group), so an agent can place the tool in the graph without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Each action carries its selection condition (list for recent calls, get to read one transcript, search to find matching lines) and four worked examples cover the common cases. It stops short of explicit when-not guidance, e.g. it never says whether to prefer this over query or summarize_conversation for transcript questions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
download_mediaDownload message mediaARead-onlyIdempotentInspect
Download media content (image, video, audio, document, sticker) from a WhatsApp message. Returns base64-encoded media data with metadata.
Get message_id from query results. The message must be a media message.
Examples: Download an image: { message_id: "true_123@c.us_3EB0...", conversation_id: "123@c.us" }
| Name | Required | Description | Default |
|---|---|---|---|
| message_id | Yes | The message ID (from query results) | |
| target_session | No | Session ID for multi-account routing | |
| conversation_id | Yes | The conversation ID containing the message |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | No | The JSON-compatible result returned by the Kaption extension |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly, idempotent, non-destructive, closed-world behavior, so safety is covered. The description adds genuinely new context beyond that: the return is base64-encoded media plus metadata, and the call fails unless the target message is a media message. It omits size/timeout limits for large media, which would matter for base64 payloads.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Purpose and return format are front-loaded, followed by the precondition and a compact example. There is minor redundancy between 'Get message_id from query results' and the schema's own '(from query results)' note, but no wasted sentences.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With an output schema present, the description need not explain return fields, and annotations plus a 100%-covered schema handle safety and parameters. The stated media precondition and ID provenance close the main gaps; only edge cases like oversized media handling are unaddressed.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the description goes further by explaining where message_id comes from and providing a concrete example with a realistic ID format ('true_123@c.us_3EB0...') that the schema's abstract string type does not convey. target_session is left to the schema only.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Download) and resource (media content from a WhatsApp message), and enumerates the media kinds covered (image, video, audio, document, sticker), which no sibling tool overlaps with. An agent can identify the right tool without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives a clear precondition ('The message must be a media message') and routes the agent to the source of the ID ('Get message_id from query results'), naming the `query` sibling as the prerequisite step. It does not state when NOT to use it (e.g., non-media messages, or when only metadata is needed), so it falls short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
export_contactsExport contactsARead-onlyIdempotentInspect
Export all WhatsApp contacts as CSV (RFC 4180) or JSON. Deduped, sorted alphabetically by display name. Default format is CSV. JSON projects the requested fields. Available fields: jid, phone, name, pushname, is_my_contact, is_business. WhatsApp "@lid" privacy identifiers (opaque, non-dialable) are always excluded — only real phone-backed contacts are returned.
Examples: CSV all fields: { format: "csv" } JSON name + phone: { format: "json", fields: ["name", "phone"] } Filtered CSV: { format: "csv", query: "Argentina" } Saved contacts only: { format: "csv", is_my_contact: true }
| Name | Required | Description | Default |
|---|---|---|---|
| query | No | Optional case-insensitive substring filter (matches name, pushname, phone, JID) | |
| fields | No | Whitelist of fields to include. Defaults to all six: jid, phone, name, pushname, is_my_contact, is_business | |
| format | No | Output format. Default "csv" | |
| is_my_contact | No | If true, only contacts saved in the user address book. If false, only un-saved. Omit to include both. |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | No | The JSON-compatible result returned by the Kaption extension |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare the safe-read profile (readOnly, idempotent, non-destructive), so the bar is lower. The description still adds real behavior beyond them: deduplication, alphabetical sorting, the CSV default, and the important rule that opaque non-dialable '@lid' identifiers are always excluded. Return structure is covered by the output schema, so no penalty there.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with purpose, format, and guarantees in three short sentences, then examples. The examples are repetitive but each demonstrates a distinct parameter combination, so they earn their space rather than padding.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be explained. For a zero-required-parameter read tool, the description covers format choice, field projection, filtering, contact-savedness, and the '@lid' exclusion rule — everything needed to call it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, and the description goes slightly beyond the schema by enumerating the available field keys and demonstrating how format/fields/is_my_contact/query combine in practice. It does not add syntax the schema lacks, so it stays below the top of the scale.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Export all WhatsApp contacts') plus concrete output formats (CSV/RFC 4180, JSON) and processing guarantees (deduped, alphabetically sorted). It never names a sibling such as list_contacts or get_contact, so the agent gets no explicit routing signal to distinguish a bulk export from the per-contact retrieval tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Four worked examples show concrete invocation patterns (all fields, projected fields, filtered, saved-only), which gives clear context for how the tool is used. However, there is no explicit when-to-use versus when-to-use-something-else guidance, e.g. no statement about preferring this over list_contacts for bulk retrieval.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_analyticsAnalyze WhatsApp activityARead-onlyIdempotentInspect
Get WhatsApp analytics data: KPIs, activity patterns, rankings, response times, call stats, labels, emojis, words, countries, and more. Supports section-based drill-down, date range filtering, chat/label/community filters, pagination, and chat/contact exports.
Start with section="overview" (default) for a compact summary, then drill into specific sections.
Sections: overview — High-level summary with top 5 chats (~5KB) kpis — Core + account KPIs + account overview activity — Daily/hourly/weekday/monthly + sent/received + message types rankings — Top chats/groups/DMs/senders (paginated) response_times — Avg/median/fastest/slowest + by-hour + by-chat calls — Call statistics (total/answered/missed/video/voice) labels — Labels (business) or Lists (personal) with chat counts emojis — Top emojis (paginated) words — Top words (paginated) countries — Contact country distribution silences — Longest-inactive chats channels — Newsletter/channel details + subscriber counts communities — Community details + sub-groups conversation_starters — Who starts conversations, night msgs, unanswered streaks — Current/longest streak + last active date gaps — Conversation gaps (>1 day silence periods) organization — Pinned/archived/muted/unread chat lists chat_detail — Full analytics for ONE specific chat (requires chat_id) export_chat — Export chat messages in format (requires chat_id + format) export_contacts — Export contacts in format (requires format) community_growth — Community member count history over time channel_growth — Channel subscriber count history over time
Examples: Overview: {} Rankings: { section: "rankings", chat_type: "group", limit: 5 } Filter by label: { section: "activity", label: "Family" } Chat detail: { section: "chat_detail", chat_id: "120363406792713578@g.us" } Export CSV: { section: "export_chat", chat_id: "...", format: "csv", limit: 100 } Export contacts: { section: "export_contacts", format: "vcf" }
| Name | Required | Description | Default |
|---|---|---|---|
| label | No | Filter by label/list name or ID | |
| limit | No | Max items for paginated sections (default 20, max 500) | |
| query | No | Search keyword — filter analytics to only messages containing this text. Shows activity patterns for conversations mentioning a topic. | |
| format | No | Export format (required for export_chat/export_contacts) | |
| offset | No | Skip N items for pagination | |
| chat_id | No | Filter to specific chat. Required for chat_detail/export_chat | |
| date_to | No | Custom end date (ISO 8601) — overrides date_range | |
| section | No | Analytics section to retrieve | overview |
| chat_type | No | Filter rankings by chat type | |
| community | No | Filter by community name or ID | |
| date_from | No | Custom start date (ISO 8601) — overrides date_range | |
| date_range | No | Preset date range (default: "30d") | |
| target_session | No | Session ID for multi-account routing | |
| include_transcriptions | No | Include audio transcriptions in exports (default true) |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | No | The JSON-compatible result returned by the Kaption extension |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly/idempotent/non-destructive, so safety is covered. The description adds real behavioral context beyond them: per-section payload size (~5KB for overview), which sections are paginated, and which parameters each mode requires. It does not discuss auth, rate limits, or truncation behavior for large exports, which keeps it off a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the summary and usage pattern before the long section catalog, and every sentence earns its place. The 22-entry section list is long but is the necessary payload of this tool; slight redundancy between the section list and the examples.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists so return values needn't be described, and the description compensates for the tool's complexity by covering section semantics, defaults, requirements per mode, pagination, and concrete invocation examples. Nothing an agent needs to select and call this correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% so the baseline is 3, but the description genuinely adds meaning: it explains what each section value returns, that paginated sections honor limit/offset, and shows parameters in context via examples (chat_type only affects rankings, format is required for exports). That is more than restating the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Get WhatsApp analytics data') and enumerates exactly what data domains are covered. The section list makes it trivially distinguishable from siblings like summarize_conversation or query, since it owns the analytics surface.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly prescribes the workflow: 'Start with section="overview" (default) for a compact summary, then drill into specific sections.' It also names the conditions for sub-modes (chat_detail requires chat_id, exports require format) and provides six worked examples covering filtering, drill-down, and export paths.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_contactGet contactARead-onlyIdempotentInspect
Look up a single WhatsApp contact by JID or phone number. Pass either parameter — both work. If multiple raw contacts share the same phone (label dupes), the saved variant wins.
Examples: By JID: { jid: "5491155550001@c.us" } By phone with +: { phone: "+5491155550001" } By phone bare: { phone: "5491155550001" }
| Name | Required | Description | Default |
|---|---|---|---|
| jid | No | Full WhatsApp JID (e.g. "5491155550001@c.us") | |
| phone | No | Phone number — leading "+" and "00" are stripped during matching |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | No | The JSON-compatible result returned by the Kaption extension |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare read-only, idempotent, non-destructive behavior, so the safety profile is covered. The description adds genuine behavioral context beyond that: it explains the dedup resolution rule ("the saved variant wins") when multiple raw contacts share a phone, which the agent could not infer from structured fields.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded purpose followed by the dedup caveat and compact examples; nothing is padded. The three examples are slightly redundant with the schema's own format descriptions, but they remain tight and scannable.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need no explanation. For a two-param, read-only lookup, the description covers purpose, both keys, and the dedup edge case, leaving nothing an agent needs missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so both parameters are already documented. The description still adds value by supplying concrete call examples for each form (JID, phone with +, bare phone), reinforcing the accepted input shapes beyond the schema's format notes.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ("Look up") and resource ("a single WhatsApp contact") plus the two lookup keys (JID or phone). The word "single" cleanly distinguishes it from the sibling list_contacts, so an agent can route without opening either schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
"Pass either parameter — both work" tells the agent the two keys are interchangeable, which is useful. However, it never states when to reach for get_contact versus list_contacts or export_contacts, so the usage context is only implied.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_contact_groupsList contact groupsARead-onlyIdempotentInspect
List the WhatsApp groups a specific contact participates in. Reads from cached chat metadata — no network.
Examples: Groups for contact: { jid: "5491155550001@c.us" }
| Name | Required | Description | Default |
|---|---|---|---|
| jid | Yes | Contact JID — must be the full @c.us form (phone alone not accepted here) |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | No | The JSON-compatible result returned by the Kaption extension |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly, idempotent, non-destructive, and closed-world, so the safety profile is covered. The description adds a genuine trait beyond them: it reads from cached metadata with no network call, which tells the agent results may be stale or empty if the cache is cold.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded single sentence stating purpose, followed by a one-line behavioral note and a compact example. No filler, though the example largely restates the schema's format constraint.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need no explanation, and annotations cover the safety profile. Cache-backed sourcing and the JID format are both addressed, leaving only the sibling boundary (list_groups) unaddressed.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%; the schema already specifies the required `jid` and its @c.us format. The description's example ({ jid: "5491155550001@c.us" }) illustrates the format but adds no semantics beyond what the schema states, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource with scope: listing WhatsApp groups a given contact belongs to. This is clearly separable from siblings like list_groups (all groups) and get_contact (contact details), though it never names those siblings to make the boundary explicit.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives useful context ('reads from cached chat metadata — no network') that implies when this is cheap/safe to call, but offers no explicit when-to-use, when-not-to-use, or named alternative (e.g. list_groups vs this tool). Usage is implied rather than stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_groupGet groupARead-onlyIdempotentInspect
Fetch a single group with a LIVE participant list. Forces Store.GroupMetadata.update() against the WA backend before reading, so the result reflects current membership including recent joins/leaves.
Use this when accuracy matters; use list_groups for browsing.
Examples: Live group fetch: { jid: "120363421729019499@g.us" }
| Name | Required | Description | Default |
|---|---|---|---|
| jid | Yes | Group JID — must end in "@g.us" |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | No | The JSON-compatible result returned by the Kaption extension |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover the safety profile (readOnly, idempotent, openWorld, non-destructive), so the bar is lower, yet the description adds genuinely non-obvious behavior: it forces Store.GroupMetadata.update() against the WA backend to guarantee fresh membership. It stops short of quantifying the implied cost/latency of that forced refresh, which would have made it fully transparent.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the core action, follows with the behavioral rationale and the routing rule, then a compact example block. Every sentence earns its place with no redundant restatement of the name or schema.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be explained. Purpose, when-to-use, alternative, and the live-refresh behavior are all present, leaving nothing an agent needs in order to call this correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% and the schema already documents jid including the required '@g.us' suffix, so the baseline is 3. The description reinforces this with a concrete example value ('120363421729019499@g.us'), adding a little practical clarity over the schema alone.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Fetch) and resource (a single group) and adds meaningful scope detail: a LIVE participant list reflecting recent joins/leaves. The scope contrasts cleanly with the sibling list_groups, so an agent can distinguish the two without opening a schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly states when to use it ('when accuracy matters') and names the alternative ('use list_groups for browsing'), giving the condition that selects between them. Nothing is left to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_contactsList contactsARead-onlyIdempotentInspect
List WhatsApp contacts from the encrypted DBR3 cache (no network). Results are deduplicated by phone number — the same person across multiple labels collapses to one row. Saved contacts sort before unsaved, then alphabetically by display name. WhatsApp "@lid" privacy identifiers (opaque, non-dialable) are always excluded — only real phone-backed contacts are returned.
Examples: All contacts: {} Saved contacts only: { is_my_contact: true } Search by name: { query: "Maria" } Page 2 of 50: { limit: 50, offset: 50 }
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | Max contacts to return (default 50, max 500) | |
| query | No | Case-insensitive substring matched against name, pushname, phone, and JID | |
| offset | No | Skip N contacts for pagination (default 0) | |
| is_my_contact | No | If true, only contacts saved in the user address book. If false, only un-saved contacts. Omit to include both. |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | No | The JSON-compatible result returned by the Kaption extension |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Adds substantial behavior beyond annotations: it reads a local encrypted cache with no network, deduplicates by phone number across labels, defines sort order (saved before unsaved, then alphabetical), and excludes non-dialable @lid identifiers. These are concrete traits an agent could not infer from the readOnly/idempotent hints.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the key facts (source, dedup, sort, exclusion) before the examples, with no filler sentences. The examples section is slightly verbose but each maps to a distinct invocation pattern, so it largely earns its space.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be explained. The description covers source, deduplication, ordering, exclusions, and pagination semantics, leaving no ambiguity for a correctly-formed call.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3; however the examples show how parameters combine (is_my_contact for filtering, query for search, limit/offset for pagination), adding practical meaning beyond the individual field descriptions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ("List WhatsApp contacts") plus scope (from the encrypted DBR3 cache, no network). Clearly distinguishes it from get_contact (single) and export_contacts (bulk export) among siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The examples demonstrate invocation modes (all, saved-only, search, pagination) but never state when to choose this tool over alternatives like get_contact or export_contacts. Usage is implied through examples rather than explicit routing guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_groupsList groupsARead-onlyIdempotentInspect
List all WhatsApp groups the user belongs to. Reads from the cache — no network.
For a live snapshot of a single group, use get_group (it forces a fresh GroupMetadata.update).
Examples: All groups: {} Search by group name: { query: "family" } Top 10: { limit: 10 }
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | Max groups to return (default 50, max 500) | |
| query | No | Case-insensitive substring match against group name |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | No | The JSON-compatible result returned by the Kaption extension |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare the full safety profile (readOnly, idempotent, non-destructive, closed-world). The description adds valuable context beyond that — it reads from cache with no network, and notes get_group forces a fresh GroupMetadata.update, which explains staleness behavior the annotations don't.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Purpose is front-loaded, immediately followed by the cache/no-network caveat and the alternative routing. The examples are compact and each demonstrates a distinct scenario with no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema is present so return values needn't be explained, and annotations cover safety. Combined with cache-vs-live disclosure and param examples, an agent has everything needed to call it correctly; only pagination/ordering of results is left unspecified.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3, but the description goes further by showing each parameter in context (query substring match, limit as top-N). This directly illustrates intended usage of both optional params.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('List') and resource ('WhatsApp groups the user belongs to') with clear scope. It distinguishes itself from the sibling get_group by contrasting list-all behavior against a single-group live lookup.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly names the alternative (get_group) and the condition that selects it ('For a live snapshot of a single group'). The three examples map concrete intents to inputs, removing ambiguity about when to use which parameter.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
manage_chatManage chatADestructiveInspect
Manage chat state: archive, unarchive, mark as read/unread, pin, unpin, mute, unmute, set/clear draft.
Actions: archive - Archive a conversation unarchive - Unarchive a conversation mark_read - Mark a conversation as read mark_unread - Mark a conversation as unread pin - Pin a conversation (max 3 pinned) unpin - Unpin a conversation mute - Mute notifications (use mute_duration for duration) unmute - Unmute notifications set_draft - Set a draft message in the compose box (requires text) clear_draft - Clear the draft message
| Name | Required | Description | Default |
|---|---|---|---|
| text | No | Draft text (required for "set_draft" action) | |
| action | Yes | Chat action to perform | |
| mute_duration | No | Mute duration (only for "mute" action). Default: "forever" | |
| target_session | No | Session ID for multi-account routing | |
| conversation_id | Yes | The conversation ID |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | No | The JSON-compatible result returned by the Kaption extension |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly=false, destructive=true, idempotent=false, and openWorld=true, covering the safety profile. The description adds the 'max 3 pinned' constraint and the mute_duration linkage, but does not explain why actions like archive/pin carry destructiveHint, whether they are reversible, or what permissions are needed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A front-loaded one-sentence summary followed by a scannable action table; every line maps to a real enum value. The action list partly duplicates the schema enum, but the human-readable glosses justify the space.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With an output schema present, return values need not be described, and the action list plus annotations cover the operational essentials for a ten-action mutation tool. Remaining gaps are sibling routing and error/permission behavior, which are minor given the structured data available.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and the enum for action already lists every value, so the schema does the heavy lifting. The description restates the same action set and repeats constraints already in the schema ('requires text', mute_duration usage, default 'forever') without adding format or validation detail.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Manage') and resource ('chat state') and enumerates all ten sub-actions with one-line semantics, so an agent knows exactly what the tool covers. It separates itself from the manage_labels/manage_notes/manage_reminders siblings via the 'chat' resource, though the distinction is implicit rather than called out.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The action list tells the agent which action to pick for a desired state change, which is implied usage guidance, and it notes the mute_duration dependency and that set_draft requires text. It never states when this tool is preferable to a sibling or when an action should be avoided, so routing remains inferred.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
manage_labelsManage business labelsADestructiveInspect
Manage WhatsApp Business labels. Requires a WhatsApp Business account.
Actions: add - Add a label to a conversation (requires label_name/label_id + conversation_id) remove - Remove a label from a conversation (requires label_name/label_id + conversation_id) create - Create a new label (requires label_name) delete - Delete a label (requires label_name or label_id)
| Name | Required | Description | Default |
|---|---|---|---|
| action | Yes | Label action to perform | |
| label_id | No | Label ID (alternative to label_name for add/remove/delete) | |
| label_name | No | Label name (for add/remove/create/delete) | |
| conversation_id | No | Conversation ID or array of IDs (required for add/remove) |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | No | The JSON-compatible result returned by the Kaption extension |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true, idempotentHint=false and openWorldHint=true, so the safety profile is covered. The description adds the genuine prerequisite that a WhatsApp Business account is required, but does not explain what delete/remove actually destroy or whether they are reversible, leaving the mutation semantics largely to the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the purpose and prerequisite, then presents a scannable action list where every line carries actionable information. No redundant or filler sentences.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be described, and the description covers the account prerequisite and per-action requirements. It is nearly complete, though it omits any note on destructive-action confirmation or error behavior that would round it out.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3, but the description adds value by mapping required parameters to specific actions (e.g. create needs only label_name, add/remove need conversation_id), a cross-parameter constraint the flat schema does not express. It does not clarify label_id vs label_name precedence beyond the schema's own note.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ("Manage WhatsApp Business labels") and enumerates the four concrete actions with their effects, so an agent immediately understands scope. It does not name or contrast with siblings like manage_notes or manage_lists, so it falls short of the 5 bar for explicit sibling differentiation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Each action is annotated with the parameters it requires (add/remove need label_name/label_id + conversation_id; create needs label_name), which effectively tells the agent when to pick each action. There is no guidance on when not to use the tool or what alternative to prefer, so it stops short of 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
manage_listsManage chat listsADestructiveInspect
Manage personal chat lists (custom lists). These are the personal account equivalent of Business labels. Lists allow organizing chats into custom categories like "Family", "Work", etc. Not available on all accounts — check with action "list" first to see if lists are enabled.
Actions: list - List all custom lists (also shows if feature is enabled) get - Get a list and its associated chats (requires id or name) create - Create a new list (requires name, optional conversation_id for initial chats) edit - Edit a list name or replace its chats (requires id or name) delete - Delete a list (requires id or name) add_chat - Add conversation(s) to a list (requires id/name + conversation_id) remove_chat - Remove conversation(s) from a list (requires id/name + conversation_id)
Examples: List all: { action: "list" } Create: { action: "create", name: "Family", conversation_id: ["number@c.us"] } Add chat: { action: "add_chat", name: "Family", conversation_id: "number@c.us" }
| Name | Required | Description | Default |
|---|---|---|---|
| id | No | List ID (for get/edit/delete/add_chat/remove_chat) | |
| name | No | List name (for create/edit, or to resolve by name) | |
| action | Yes | List action to perform | |
| target_session | No | Session ID for multi-account routing | |
| conversation_id | No | Chat ID or array of IDs to add/remove |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | No | The JSON-compatible result returned by the Kaption extension |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true, openWorldHint=true, and non-idempotent behavior; the description adds genuinely new context by disclosing that the feature is account-gated and must be verified via 'list' before use. It stops short of stating that delete is permanent or what auth is required, which keeps it from a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded purpose followed by a scannable action table and concrete examples. Dense but each block earns its place; the examples are slightly redundant with the action list but aid invocation.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 7-action multi-mode tool, it covers feature availability, per-action parameter requirements, and worked examples, and an output schema exists so return shapes need not be described. Nothing an agent needs to invoke it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the description goes beyond the schema by mapping which parameters each action requires (id/name resolution, conversation_id for add/remove). That action-to-parameter mapping is real added meaning.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Manage personal chat lists'), and explicitly distinguishes them from Business labels, which routes the agent away from the sibling manage_labels tool. An agent knows exactly what domain this covers.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The action list gives per-action requirements (which params each action needs), and it explicitly warns to call action 'list' first to verify the feature is enabled. It does not, however, route to siblings like manage_labels for business labels, so context is clear but not fully exclusionary.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
manage_notesManage contact notesADestructiveIdempotentInspect
Manage contact notes. Requires a WhatsApp Business account with notes enabled.
Actions: get - Read the note for a contact set - Write/update the note for a contact
| Name | Required | Description | Default |
|---|---|---|---|
| note | No | Note text (required for "set" action) | |
| action | Yes | Note action to perform | |
| contact_id | Yes | The contact ID |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | No | The JSON-compatible result returned by the Kaption extension |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=false, destructiveHint=true and idempotentHint=true, and the description is consistent with them. The description adds the account/prerequisite requirement, but for a destructive write it never says that 'set' overwrites/replaces any existing note text, which is the most important behavioral fact an agent needs before calling it.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Extremely tight: one scope line, one prerequisite line, two action bullets. Front-loaded and every sentence earns its place with no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 3-parameter tool with an output schema, full schema coverage, and annotations, the description covers scope, prerequisite, and per-action intent. The only real omission is the overwrite/replacement behavior of 'set' on an existing note, which is a meaningful gap for a destructive operation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and the enum for 'action' is declared in the schema, so the schema carries parameter meaning. The description's only parameter-adjacent statement ('set - Write/update the note') restates the action rather than adding syntax, format, or length constraints for note text; baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
It states a concrete verb+resource (manage contact notes) and then enumerates the two discrete operations, 'get' and 'set', so an agent knows exactly what surface it covers. The note-oriented resource does not overlap with any sibling tool (get_contact, list_contacts, manage_labels, etc.), so there is no ambiguity about which tool to pick.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives a real prerequisite ('Requires a WhatsApp Business account with notes enabled') which is genuine usage context. However it never states when to use this versus a sibling such as get_contact, nor any exclusion conditions; the action list explains what each action does, not when to choose it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
manage_remindersManage remindersADestructiveInspect
Manage personal reminders. Reminders are stored in the cloud and trigger notifications via the Kaption extension.
Actions: list - List all reminders get - Get a specific reminder by ID create - Create a new reminder (requires title + datetime) update - Update a reminder (requires id, optional title/datetime) delete - Delete a reminder (requires id) complete - Mark a reminder as completed (requires id) uncomplete - Mark a reminder as not completed (requires id)
Examples: List all: { action: "list" } Create: { action: "create", title: "Follow up with client", datetime: "2026-03-07T14:00:00Z" } Complete: { action: "complete", id: "rem_abc123" }
| Name | Required | Description | Default |
|---|---|---|---|
| id | No | Reminder ID (required for get/update/delete/complete/uncomplete) | |
| title | No | Reminder text (required for create, optional for update) | |
| action | Yes | Reminder action to perform | |
| filter | No | Filter for list action. Default: "active" (non-completed only) | |
| datetime | No | ISO 8601 datetime for the reminder (required for create, optional for update) | |
| target_session | No | Session ID for multi-account routing | |
| notification_type | No | How to notify. Default: "automatic" |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | No | The JSON-compatible result returned by the Kaption extension |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true and readOnlyHint=false, so the safety profile is covered; the description usefully adds that reminders live in the cloud and fire notifications through the Kaption extension. It does not, however, flag that delete is irreversible while list/get are safe reads, nor mention auth or rate limits.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Purpose, action list, and examples are front-loaded and scannable, with no wasted prose. There is mild duplication in restating per-action required parameters that the schema already documents, which keeps it just short of a 5.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists so return values need not be explained, and 100% schema coverage documents filter, target_session, and notification_type. All actions and their required inputs are described, leaving only minor gaps such as the destructive nature of delete and multi-account routing behavior for target_session.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3, but the description goes further by stating which parameters each action requires (create needs title+datetime; update needs id plus optional title/datetime) and by giving concrete JSON examples. That combination of action-param mapping and examples adds practical meaning beyond the field-level schema text.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific resource ('personal reminders') and enumerates seven concrete operations (list/get/create/update/delete/complete/uncomplete), so an agent can tell exactly what the tool does. The reminder domain is distinct from every sibling (notes, labels, lists, scheduled messages), though the description never explicitly contrasts itself with them.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The action list with per-action parameter requirements and worked examples implies when each sub-action applies, and create/complete examples show the calling shape. However, there is no explicit guidance on when to prefer this tool over siblings like manage_notes or manage_scheduled_messages, and no when-not conditions or prerequisites (e.g., auth, extension availability) are stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
manage_scheduled_messagesManage scheduled messagesADestructiveInspect
Schedule WhatsApp messages to be sent automatically at a specific time, in one of two modes: bot (default) - sent from Kaption's WhatsApp number, not the user's own. Stored in Kaption's cloud and sent even when this computer is off. One-to-one chats only (no groups); one line, max 800 characters. local ("From this computer") - sent from the user's own WhatsApp number, as them, while this computer and WhatsApp are open. Stays on this device. Kaption keeps it within safe limits automatically (a few messages an hour and a day, minutes apart, only to chats where the other person has written, at most 3 a day to groups) and refuses what does not fit. The only mode that can send to groups (ones the user can post in and posted in within 30 days). Text only. A message whose time passes while this computer is off is missed, never sent late on its own. The local mode must first be turned on by the user in Kaption (its "From this computer" option, after reading the risks); an assistant cannot turn it on. A refused local request says why and ends with a reason code, e.g. "(reason: per-day)".
Actions: list - List scheduled messages of both modes (each has "mode"); pass mode to list only one get - Get a specific scheduled message by ID create - Schedule a new message (requires message + datetime + conversation_id) update - Change the message and/or datetime (requires id) delete - Cancel/delete a scheduled message (requires id); in local mode it cancels a waiting message and removes a finished one cancel - Local mode only: cancel a waiting, missed or failed message remove - Local mode only: remove a finished message from the list send_now - Local mode only: send a missed or failed message now (still within the limits) Pass mode "local" for every action on a local message.
Examples: List all: { action: "list" } Schedule (Kaption bot): { action: "create", message: "Hey, just following up!", datetime: "2026-03-07T09:00:00Z", conversation_id: "5491157390064@c.us" } Schedule from the user's own number: { action: "create", mode: "local", message: "Running 10 min late", datetime: "2026-03-07T09:00:00-03:00", conversation_id: "5491157390064@c.us" } Schedule to a group: { action: "create", mode: "local", message: "Standup moved to 10", datetime: "2026-03-07T09:00:00Z", conversation_id: "120363000000000000@g.us" } Cancel: { action: "delete", id: "msg_abc123" }
| Name | Required | Description | Default |
|---|---|---|---|
| id | No | Scheduled message ID (required for get/update/delete/cancel/remove/send_now) | |
| mode | No | "bot" (default): sent from Kaption's WhatsApp number. "local": sent from the user's own WhatsApp number while this computer and WhatsApp are open, within safe limits kept automatically; needed for groups. For list, omit to get both modes | |
| action | Yes | Scheduled message action to perform (cancel, remove and send_now need mode "local") | |
| filter | No | Filter for list action. Default: "pending" (not sent yet; in local mode also missed or failed ones waiting for the user) | |
| message | No | Message text to send (required for create, optional for update). Bot: one line, max 800 characters. Local: text only, up to 2000 characters | |
| datetime | No | ISO 8601 datetime when the message should be sent (required for create, optional for update) | |
| target_session | No | Session ID for multi-account routing | |
| conversation_id | No | Chat to send the message to (required for create): a person, or a group with mode "local" | |
| notification_type | No | Bot mode only. Notification type. Default: "automatic" |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | No | The JSON-compatible result returned by the Kaption extension |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the annotations (destructive=true, openWorld=true, non-idempotent), it discloses rate limits ('a few messages an hour and a day, minutes apart'), group eligibility (posted within 30 days), 800/2000 char caps, offline-miss behavior, retention ('stays on this device'), and refusal reason codes like '(reason: per-day)'. This is unusually rich operational context.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
It is long but front-loads the two modes before the action list and examples, and every section is functional. There is some redundancy between the prose mode description and the repeated per-action notes, but given the 8-action/9-param complexity it stays mostly disciplined.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need no prose. With modes, per-action requirements, limits, failure semantics, and worked examples all covered, an agent has everything needed to select an action and construct a valid call.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the description adds value beyond the schema by showing concrete example payloads and mapping mode requirements to specific actions (local needed for groups, delete vs cancel vs remove vs send_now distinctions). It largely restates schema-documented field semantics, hence not a 5.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource (schedule WhatsApp messages to be sent automatically) and enumerates the eight actions the tool dispatches. No sibling among the listed tools covers scheduled messaging, so it is cleanly distinguishable.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explicitly contrasts bot vs local mode, states the prerequisites for local mode ('must first be turned on by the user... an assistant cannot turn it on'), and ties each action to a mode constraint via per-action notes. When-to-use and when-not-to-use are both spelled out.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
querySearch WhatsAppARead-onlyIdempotentInspect
Query WhatsApp data: conversations, contacts, messages, transcriptions, labels, and communities. Supports listing, searching, filtering, and looking up by ID.
IMPORTANT: Multiple WhatsApp accounts may be connected (e.g. personal + business). Always query entity="session" FIRST to see all connected accounts and their session IDs. Then use target_session to route queries to the correct account. Each account has different conversations, contacts, and messages.
HOW TO READ MESSAGES: To get messages from a specific conversation, pass its id (e.g. "5491157390064@c.us"). This returns the conversation info WITH its messages. Use limit to control how many. Do NOT use entity="messages" for this — that is for global text search only.
AUDIO TRANSCRIPTIONS: To get audio transcriptions, use entity="transcriptions" with an optional query. Or pass a conversation id to see messages (audio messages include transcription text).
Examples: List sessions: { entity: "session" } List conversations: {} Target specific account: { entity: "conversations", target_session: "sess_abc123" } Read messages: { id: "5491157390064@c.us" } Read last 100 msgs: { id: "5491157390064@c.us", limit: 100 } Search globally: { query: "meeting" } Search in chat: { id: "5491157390064@c.us", query: "meeting" } Unread conversations: { unread: true } Search contacts: { query: "Alice", entity: "contacts" } List labels: { entity: "labels" } Filter by label: { label: "Important", entity: "conversations" } List communities: { entity: "communities" } Filter by community: { community: "My Community", entity: "conversations" } Find which groups a contact is in: { id: "5491157390064@c.us", entity: "contacts", include_participants: true } List members of a group: { entity: "contacts", group: "120363421729019499@g.us" }
| Name | Required | Description | Default |
|---|---|---|---|
| id | No | Look up a specific conversation, contact, or label by ID | |
| list | No | Filter conversations by list name or ID (Personal accounts) | |
| after | No | Return messages after this ISO 8601 datetime (e.g. "2026-03-01T12:00:00.000Z") for incremental sync | |
| group | No | Filter contacts by group ID — only return contacts that are members of this group | |
| label | No | Filter conversations by label name or ID (Business accounts) | |
| limit | No | Max results (default 25, max 5000) | |
| query | No | Text to search for (names, messages, transcriptions) | |
| before | No | Return messages before this ISO 8601 datetime (e.g. "2026-03-01T12:00:00.000Z") for cursor-based pagination backward | |
| entity | No | Entity type to query. Defaults to "conversations" when listing, or all when searching. Use "session" to list all connected WhatsApp accounts. | |
| unread | No | Only return conversations with unread messages | |
| community | No | Filter conversations by community name or ID | |
| exclude_muted | No | Exclude muted conversations from listings (default false) | |
| target_session | No | Session ID to target a specific WhatsApp account. Get session IDs from entity="session". If omitted, routes to the most recently active account. | |
| exclude_archived | No | Exclude archived conversations from listings (default true) | |
| include_participants | No | Include group participants in results. Useful when looking up a contact by ID to see which groups they belong to, or when querying a group to see its members. |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | No | The JSON-compatible result returned by the Kaption extension |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly/idempotent/non-destructive, so the safety profile is covered. The description adds genuinely non-structured behavior: multiple accounts may be connected, queries default to the most recently active account when target_session is omitted, and id-lookup returns conversation info WITH its messages. It stops short of describing pagination/return shape, but that is largely handled elsewhere.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Well front-loaded — purpose, then the critical session-routing caveat, then usage sections. The example block is long and slightly redundant (several read-message variants), but each example demonstrates a distinct call shape and the structure is scannable.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With an output schema present, return values need not be explained. The description covers the multi-account complexity, entity selection, ID-vs-search distinction, and transcription access — everything an agent needs to invoke this broad 15-parameter tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so baseline is 3, but the description adds real meaning beyond the schema: entity="session" is positioned as the mandatory first step, target_session's default routing behavior is explained, and the id+limit interaction ('use limit to control how many') clarifies semantics not spelled out per-parameter.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Query) and enumerates the exact resources (conversations, contacts, messages, transcriptions, labels, communities) plus the supported operations (listing, searching, filtering, lookup by ID). An agent can distinguish this from siblings like get_contact, list_contacts, and summarize_conversation without opening a schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly routes the agent: query entity="session" FIRST for multi-account routing, then target_session. It names the alternative and the exclusion ('Do NOT use entity="messages" for this — that is for global text search only'), and the example block maps intents to calls.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
summarize_conversationSummarize conversationCIdempotentInspect
Get or generate a summary of a conversation
| Name | Required | Description | Default |
|---|---|---|---|
| message_count | No | Number of messages to use for summary generation (default 50, max 500) | |
| target_session | No | Session ID to target a specific WhatsApp account | |
| conversation_id | Yes | The conversation ID |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | No | The JSON-compatible result returned by the Kaption extension |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations supply the safety profile (readOnlyHint=false, idempotentHint=true, destructiveHint=false), so the agent knows calling it may have side effects but is repeatable. The description's 'get or generate' is broadly consistent with a non-read-only tool, but it never clarifies whether a summary is persisted, how costly generation is, or when generation is triggered.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single short sentence with no filler and the action is front-loaded. It is efficiently sized, though its brevity comes partly from underspecification rather than tight editing.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be explained. However, the central ambiguity of get-vs-generate and the side effects implied by readOnlyHint=false remain unaddressed, leaving the description only minimally adequate for this tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, including the default/max for message_count and the purpose of target_session, so the schema already carries parameter meaning. The description adds nothing beyond that, which is the baseline 3 when the schema does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a verb and resource (summarize/get a summary of a conversation), but the dual 'get or generate' phrasing is vague about what actually happens on a call. No differentiation from siblings like query or manage_chat is offered, so an agent can't tell when this is the right tool.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No when-to-use guidance, no prerequisites, and no mention of alternatives. The description gives no signal about whether this is for fetching an existing summary versus triggering generation, which is exactly the decision an agent must make.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
17 tool updates
- First observed
call_recordings - First observed
download_media - First observed
export_contacts - First observed
get_analytics - First observed
get_contact - First observed
get_contact_groups - First observed
get_group - First observed
list_contacts - First observed
list_groups - First observed
manage_chat - First observed
manage_labels - First observed
manage_lists - First observed
manage_notes - First observed
manage_reminders - First observed
manage_scheduled_messages - First observed
query - First observed
summarize_conversation
Related MCP Servers
- AlicenseAqualityCmaintenanceEnables brand visibility monitoring across major AI platforms like ChatGPT, Claude, Gemini, and Perplexity. It allows users to track visibility scores, analyze competitor data, and receive actionable insights to improve AI-generated brand recommendations.1622 npm1MIT
- AlicenseCqualityBmaintenanceCompetitor Monitor AI - MCP server providing AI-powered tools and automation by MEOK AI Labs1114 npm40 PyPIMIT
- AlicenseAqualityCmaintenanceRevnuvo Company Intelligence tells AI agents what changed at a company, with evidence. It observes company websites, technologies, and DNS over time and returns timestamped, confidence-aware changes, signals, and monitoring.9MIT

industrylens-mcpofficial
AlicenseNot gradedqualityBmaintenanceBrowse IndustryLens's published competitive-intelligence reports and head-to-head competitor comparisons from any AI agent — real, source-backed data.MIT
Glama MCP Gateway
Add one secure layer between your agents and this server.