kaption-whatsapp-mcp
Provides tools for reading and managing WhatsApp conversations via WhatsApp Web, including querying conversations, contacts, messages, transcriptions, labels and communities; summarizing conversations; managing Business labels and personal chat lists; managing chat state (archive, mark read/unread, pin, mute, set drafts); creating personal reminders; and scheduling messages to be sent automatically either via Kaption's bot or from the user's own number.
@kaptionai/mcp-extension
MCP server that lets AI assistants read and manage your WhatsApp conversations through the KaptionAI Chrome extension. Also supports WebMCP for zero-config browser-native AI tool discovery.
Claude / Cursor ──stdio──> mcp-whatsapp ──ws://localhost:7865──> KaptionAI Extension ──> WhatsApp Web
Browser AI Agent ──navigator.modelContext──> KaptionAI Extension ──> WhatsApp Web (WebMCP)Setup
1. Install the extension
Install the KaptionAI Chrome Extension and enable the MCP bridge in settings.
2. Configure your AI tool
Claude Desktop
Add to ~/Library/Application Support/Claude/claude_desktop_config.json:
{
"mcpServers": {
"whatsapp": {
"command": "npx",
"args": ["-y", "@kaptionai/mcp-extension@latest"]
}
}
}Claude Code
claude mcp add whatsapp -- npx -y @kaptionai/mcp-extension@latestCursor
Add to .cursor/mcp.json in your project or go to Settings > MCP Servers:
{
"mcpServers": {
"whatsapp": {
"command": "npx",
"args": ["-y", "@kaptionai/mcp-extension@latest"]
}
}
}Windsurf
Add to ~/.codeium/windsurf/mcp_config.json:
{
"mcpServers": {
"whatsapp": {
"command": "npx",
"args": ["-y", "@kaptionai/mcp-extension@latest"]
}
}
}VS Code (Copilot)
Add to .vscode/mcp.json in your project:
{
"servers": {
"whatsapp": {
"type": "stdio",
"command": "npx",
"args": ["-y", "@kaptionai/mcp-extension@latest"]
}
}
}Or add to your VS Code settings.json:
{
"mcp": {
"servers": {
"whatsapp": {
"type": "stdio",
"command": "npx",
"args": ["-y", "@kaptionai/mcp-extension@latest"]
}
}
}
}Zed
Add to Zed settings (~/.config/zed/settings.json):
{
"context_servers": {
"whatsapp": {
"command": {
"path": "npx",
"args": ["-y", "@kaptionai/mcp-extension@latest"]
}
}
}
}OpenAI Agents SDK (Python)
from agents import Agent
from agents.mcp import MCPServerStdio
whatsapp = MCPServerStdio(
name="whatsapp",
command="npx",
args=["-y", "@kaptionai/mcp-extension@latest"],
)
agent = Agent(
name="assistant",
instructions="You can access WhatsApp conversations.",
mcp_servers=[whatsapp],
)Any MCP-compatible client
This package runs as a standard MCP server over stdio. To connect from any client:
npx -y @kaptionai/mcp-extension@latestThe server communicates via stdin/stdout using the MCP protocol. Point your client's MCP configuration to this command.
3. Open WhatsApp Web
Open web.whatsapp.com in Chrome or Edge. The extension will auto-connect to the MCP server.
Related MCP server: gaviwhatsapp-mcp
Tools
query
Query WhatsApp data — conversations, contacts, messages, transcriptions, labels, and communities.
# List conversations
query {}
# Search everything
query { query: "meeting" }
# Unread only
query { unread: true }
# Look up a conversation with messages
query { id: "5511999887766@c.us" }
# Search contacts
query { query: "Alice", entity: "contacts" }
# List labels (Business accounts)
query { entity: "labels" }
# Filter by label
query { label: "Important" }
# List communities
query { entity: "communities" }
# Filter by community
query { community: "My Community" }
# Get session info
query { entity: "session" }Parameter | Type | Description |
| string | Search text (names, messages, transcriptions) |
| string | Look up a specific conversation, contact, or label |
| string |
|
| number | Max results (default 25, max 5000) |
| boolean | Only unread conversations |
| string | Filter by label name or ID |
| string | Filter by community name or ID |
| string | Messages before this ISO 8601 timestamp (pagination) |
| string | Messages after this ISO 8601 timestamp (incremental sync) |
summarize_conversation
Generate a summary of a conversation.
Parameter | Type | Description |
| string | The conversation ID |
manage_labels
Manage WhatsApp Business labels — add, remove, create, or delete labels.
Parameter | Type | Description |
| string |
|
| string | Label name |
| string | Label ID (alternative to name) |
| string | Required for add/remove |
manage_lists
Manage personal chat lists (custom lists). The personal account equivalent of Business labels — organize chats into custom categories.
# List all custom lists
manage_lists { action: "list" }
# Get a list with its chats
manage_lists { action: "get", name: "Family" }
# Create a new list
manage_lists { action: "create", name: "Work", conversation_id: "5511999887766@c.us" }
# Add a chat to a list
manage_lists { action: "add_chat", name: "Family", conversation_id: "5511999887766@c.us" }
# Remove a chat from a list
manage_lists { action: "remove_chat", name: "Family", conversation_id: "5511999887766@c.us" }
# Delete a list
manage_lists { action: "delete", name: "Old List" }Parameter | Type | Description |
| string |
|
| string | List ID |
| string | List name (for create/edit, or to resolve by name) |
| string | Chat ID(s) to add/remove |
manage_chat
Manage chat state — archive, unarchive, mark as read/unread, pin, unpin, mute, unmute, set/clear draft messages.
# Archive a chat
manage_chat { action: "archive", conversation_id: "5511999887766@c.us" }
# Mark as read
manage_chat { action: "mark_read", conversation_id: "5511999887766@c.us" }
# Pin a chat (max 3)
manage_chat { action: "pin", conversation_id: "5511999887766@c.us" }
# Mute for 1 week
manage_chat { action: "mute", conversation_id: "5511999887766@c.us", mute_duration: "1w" }
# Set a draft message
manage_chat { action: "set_draft", conversation_id: "5511999887766@c.us", text: "Hey, I'll call you back" }Parameter | Type | Description |
| string |
|
| string | The conversation ID |
| string |
|
| string | Draft text. Required for |
manage_reminders
Create and manage personal reminders. Stored in the cloud and delivered via the Kaption extension.
# List active reminders
manage_reminders { action: "list" }
# Create a reminder
manage_reminders { action: "create", title: "Follow up with client", datetime: "2026-03-07T14:00:00Z" }
# Complete a reminder
manage_reminders { action: "complete", id: "rem_abc123" }
# List all including completed
manage_reminders { action: "list", filter: "all" }Parameter | Type | Description |
| string |
|
| string | For list: |
| string | Reminder ID (for get/update/delete/complete/uncomplete) |
| string | Reminder text (max 800 chars, no newlines) |
| string | ISO 8601 datetime |
| string |
|
manage_scheduled_messages
Schedule messages to be sent automatically at a specific time, in one of two modes:
bot(default) — sent from Kaption's WhatsApp number, not yours. Stored in Kaption's cloud and sent even when your computer is off. One-to-one chats only; one line, up to 800 characters.local("From this computer") — sent from your own WhatsApp number, as you, while this computer and WhatsApp are open. Nothing leaves your device. Kaption keeps it within safe limits automatically (a few messages an hour and a day, minutes apart, only to chats where the other person has written, at most 3 a day to groups) and refuses what doesn't fit, with the reason. The only mode that can send to groups (ones you can post in and posted in within the last 30 days). Text only. A message whose time passes while the computer is off is marked missed, never sent late on its own.
The local mode has to be turned on by you in Kaption first (the "From this computer" option in the scheduling picker, after reading the limits and risks). An AI assistant can't turn it on; until you do, local requests are refused with (reason: no-consent). It is also rolled out gradually: where it isn't available yet, local requests are refused with (reason: flag-off) and the bot works as before.
# List pending messages of both modes (each one has a "mode")
manage_scheduled_messages { action: "list" }
# Schedule with the Kaption bot
manage_scheduled_messages { action: "create", message: "Hey, just following up!", datetime: "2026-03-07T09:00:00Z", conversation_id: "5511999887766@c.us" }
# Schedule from your own number, or to a group
manage_scheduled_messages { action: "create", mode: "local", message: "Running 10 min late", datetime: "2026-03-07T09:00:00-03:00", conversation_id: "5511999887766@c.us" }
manage_scheduled_messages { action: "create", mode: "local", message: "Standup moved to 10", datetime: "2026-03-07T09:00:00Z", conversation_id: "120363000000000000@g.us" }
# Cancel a scheduled message
manage_scheduled_messages { action: "delete", id: "msg_abc123" }
manage_scheduled_messages { action: "cancel", mode: "local", id: "3f2c…" }Parameter | Type | Description |
| string |
|
| string |
|
| string | For list: |
| string | Scheduled message ID (for get/update/delete/cancel/remove/send_now) |
| string | Chat to send to (for create): a person, or a group with |
| string | Message text. Bot: max 800 chars, no newlines. Local: up to 2000 chars |
| string | ISO 8601 datetime when the message should be sent |
| string | Bot only: |
download_media
Download and decrypt media from a message (images, videos, audio, documents).
Parameter | Type | Description |
| string | The message ID containing media |
| string | The conversation the message belongs to |
Returns base64-encoded media data with mimetype, size, duration, and caption.
manage_notes
Read and write contact notes (WhatsApp Business).
Parameter | Type | Description |
| string |
|
| string | The contact ID |
| string | Note text (required for |
call_recordings
Read WhatsApp call recordings made by the Kaption extension and their transcripts, with every word timed. Each recording names its conversation (the other person, or the group for a group call). Recordings of locked chats are never returned.
Parameter | Type | Description |
| string |
|
| string | Recording ID (required for |
| string | Text to find in names and transcripts (required for |
| string | Only recordings of this contact or group |
| string | ISO 8601 date range of when the call started |
| number | Max recordings (default 20, max 100) |
| boolean | For |
get_api_info
Get HTTP REST API connection info for programmatic access without MCP overhead. Returns URL, auth token, and available endpoints.
Multi-account support
Multiple WhatsApp accounts can be connected simultaneously (e.g. personal + business). Use target_session on any tool to route to a specific account. Query entity: "session" to see all connected accounts and their session IDs.
Multi-instance support
Multiple AI tools can share the same extension connection. The first instance starts a WebSocket hub; subsequent instances auto-detect the existing hub and relay through it. If the hub stops, a relay automatically promotes itself.
WebMCP support
Kaption is WebMCP-ready. On browsers that support the W3C WebMCP draft (navigator.modelContext, Chrome 146+), the extension automatically registers all tools with the browser's native AI tool registry. This means browser-based AI agents can discover and invoke Kaption tools without any MCP server or WebSocket connection — zero configuration.
When WebMCP is available, the extension registers tools prefixed with kaption_ (e.g. kaption_query, kaption_manage_chat) complete with JSON Schema input definitions and readOnlyHint annotations. The tools use the same handlers as the MCP server, so behavior is identical across both paths.
Security
Localhost only — no cloud relay, no external connections
No messages sent — AI assistants can read, organize, schedule, and draft, but never send messages directly
Locked chats hidden — WhatsApp-locked conversations are excluded from all queries
Feature-gated — MCP bridge must be explicitly enabled in the extension
Business features gated — labels and notes require a WhatsApp Business account
Rate limited — draft messages limited to 10 conversations per 5-minute window; write operations include random delays
License
BSL 1.1 — free to use, converts to MIT after 4 years.
Available Tools
18 toolscall_recordingsARead-only
Read WhatsApp call recordings made by the Kaption extension and their transcripts, with every word timed. Each recording names its conversation (the other person, or the group for a group call), so it links to query, get_contact and get_group. Recordings of locked chats are never returned.
Actions: list - Recordings, newest first (optional conversation_id, date_from, date_to, limit) get - One recording and its transcript: turns by speaker ("you" or "contact") with start/end seconds (requires id; include_words adds each word's timing) search - Recordings whose name or transcript matches the search text, with the matching lines (requires search; same filters as list)
Examples: Recent calls: { action: "list", limit: 10 } Calls with one person: { action: "list", conversation_id: "5491157390064@c.us" } Read a transcript: { action: "get", id: "1790000000000-a1b2c3d4" } What was said about the budget: { action: "search", search: "budget" }
| Name | Required | Description | Default |
|---|---|---|---|
| id | No | Recording ID (required for get) | |
| limit | No | Max recordings to return (default 20, max 100) | |
| action | Yes | Call recordings action to perform | |
| search | No | Text to find in names and transcripts (required for search) | |
| date_to | No | Only recordings started on or before this date (ISO 8601) | |
| date_from | No | Only recordings started on or after this date (ISO 8601) | |
| include_words | No | For get: include each word with its start and end in seconds (default false) | |
| target_session | No | Session ID for multi-account routing | |
| conversation_id | No | Only recordings of this conversation (a contact or group ID) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and openWorldHint=false, so the safety profile is covered; the description still earns credit by disclosing that locked-chat recordings are suppressed from results and that transcripts are speaker-attributed with per-word timings. It does not cover rate limits, pagination behavior, or how missing transcripts are represented.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded purpose, then a compact action table, then examples – a structure that is easy to scan and every block earns its place. Minor redundancy remains where the action list restates parameter names already documented in the schema.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description carries the return-value burden and does it well: turns by speaker, start/end seconds, optional per-word timing, and matching lines for search. It omits any mention of target_session (multi-account routing) and result ordering/pagination semantics for search, which leaves small gaps for a 9-parameter tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the description adds genuine meaning beyond the schema by scoping parameters to actions (conversation_id/date_from/date_to/limit on list and search, id on get, include_words only on get) and by showing a real conversation_id format ('5491157390064@c.us') and recording id format.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The opening sentence states a specific verb and resource ('Read WhatsApp call recordings ... and their transcripts, with every word timed') and adds the distinguishing scope that each recording links to query, get_contact and get_group. An agent can separate this from siblings like query or summarize_conversation without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The three actions are enumerated with the parameters each one accepts and when each is required ('requires id', 'requires search'), plus four worked examples that model typical calls. There is an explicit exclusion ('Recordings of locked chats are never returned') but no direct 'use X instead of Y' routing against siblings, which keeps this below a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
download_mediaARead-only
Download media content (image, video, audio, document, sticker) from a WhatsApp message. Returns base64-encoded media data with metadata.
Get message_id from query results. The message must be a media message.
Examples: Download an image: { message_id: "true_123@c.us_3EB0...", conversation_id: "123@c.us" }
| Name | Required | Description | Default |
|---|---|---|---|
| message_id | Yes | The message ID (from query results) | |
| target_session | No | Session ID for multi-account routing | |
| conversation_id | Yes | The conversation ID containing the message |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and openWorldHint=true, so safety is covered. The description adds genuinely new context beyond them: the return payload is base64-encoded media plus metadata, and the input must reference a media message. It omits size limits or rate limits, so it stops short of a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The what/return/precondition are front-loaded in three short lines, and the example is compact and directly actionable. The example largely repeats the schema's required fields, so it earns slightly less than full marks.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description usefully discloses the return format (base64 data + metadata) and the media-only precondition. Combined with annotations and a fully described 3-parameter schema, an agent has what it needs, though routing behavior for target_session remains unexplained.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so both required parameters are already documented in the schema. The description restates that message_id comes from query results but says nothing about target_session or the multi-account routing behavior, so it adds little beyond the schema baseline.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description gives a specific verb (Download) and resource (media content) and enumerates the media kinds covered (image, video, audio, document, sticker). No sibling tool competes in this domain, so an agent can route here unambiguously.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It states a clear precondition ('The message must be a media message') and tells the agent where message_id comes from ('from query results'), which links to the query sibling. It does not name an alternative for non-media messages, but the trigger condition is explicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
export_contactsARead-only
Export all WhatsApp contacts as CSV (RFC 4180) or JSON. Deduped, sorted alphabetically by display name. Default format is CSV. JSON projects the requested fields. Available fields: jid, phone, name, pushname, is_my_contact, is_business. WhatsApp "@lid" privacy identifiers (opaque, non-dialable) are always excluded — only real phone-backed contacts are returned.
Examples: CSV all fields: { format: "csv" } JSON name + phone: { format: "json", fields: ["name", "phone"] } Filtered CSV: { format: "csv", query: "Argentina" } Saved contacts only: { format: "csv", is_my_contact: true }
| Name | Required | Description | Default |
|---|---|---|---|
| query | No | Optional case-insensitive substring filter (matches name, pushname, phone, JID) | |
| fields | No | Whitelist of fields to include. Defaults to all six: jid, phone, name, pushname, is_my_contact, is_business | |
| format | No | Output format. Default "csv" | |
| is_my_contact | No | If true, only contacts saved in the user address book. If false, only un-saved. Omit to include both. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only cover readOnly/openWorld, but the description adds real behavioral detail: deduped, alphabetically sorted by display name, CSV is the default, JSON projects requested fields, and @lid privacy identifiers are always excluded. That last point is exactly the kind of non-obvious side effect the agent needs.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core statement and the exclusion rule, then examples. The examples consume a lot of space, but each one demonstrates a distinct parameter combination rather than padding.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description carries the return-value burden and does so: it names the formats, the field set, sorting, dedup, and the @lid exclusion. An agent can call this correctly without any additional information.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the description adds value by enumerating the six available fields and demonstrating format/fields/query/is_my_contact combinations. It goes slightly beyond the schema's own descriptions rather than merely repeating them.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource (export all contacts) plus the output formats (CSV RFC 4180 / JSON), so the agent knows exactly what it produces. It does not, however, distinguish itself from the sibling list_contacts, which an agent could reasonably confuse it with.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Four concrete examples map intent to parameter combinations (all fields, field projection, query filtering, saved-only), which is strong implied guidance. It still never says when to prefer this over list_contacts or when not to use it, so it stops short of explicit routing.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_analyticsARead-only
Get WhatsApp analytics data: KPIs, activity patterns, rankings, response times, call stats, labels, emojis, words, countries, and more. Supports section-based drill-down, date range filtering, chat/label/community filters, pagination, and chat/contact exports.
Start with section="overview" (default) for a compact summary, then drill into specific sections.
Sections: overview — High-level summary with top 5 chats (~5KB) kpis — Core + account KPIs + account overview activity — Daily/hourly/weekday/monthly + sent/received + message types rankings — Top chats/groups/DMs/senders (paginated) response_times — Avg/median/fastest/slowest + by-hour + by-chat calls — Call statistics (total/answered/missed/video/voice) labels — Labels (business) or Lists (personal) with chat counts emojis — Top emojis (paginated) words — Top words (paginated) countries — Contact country distribution silences — Longest-inactive chats channels — Newsletter/channel details + subscriber counts communities — Community details + sub-groups conversation_starters — Who starts conversations, night msgs, unanswered streaks — Current/longest streak + last active date gaps — Conversation gaps (>1 day silence periods) organization — Pinned/archived/muted/unread chat lists chat_detail — Full analytics for ONE specific chat (requires chat_id) export_chat — Export chat messages in format (requires chat_id + format) export_contacts — Export contacts in format (requires format) community_growth — Community member count history over time channel_growth — Channel subscriber count history over time
Examples: Overview: {} Rankings: { section: "rankings", chat_type: "group", limit: 5 } Filter by label: { section: "activity", label: "Family" } Chat detail: { section: "chat_detail", chat_id: "120363406792713578@g.us" } Export CSV: { section: "export_chat", chat_id: "...", format: "csv", limit: 100 } Export contacts: { section: "export_contacts", format: "vcf" }
| Name | Required | Description | Default |
|---|---|---|---|
| label | No | Filter by label/list name or ID | |
| limit | No | Max items for paginated sections (default 20, max 500) | |
| query | No | Search keyword — filter analytics to only messages containing this text. Shows activity patterns for conversations mentioning a topic. | |
| format | No | Export format (required for export_chat/export_contacts) | |
| offset | No | Skip N items for pagination | |
| chat_id | No | Filter to specific chat. Required for chat_detail/export_chat | |
| date_to | No | Custom end date (ISO 8601) — overrides date_range | |
| section | No | Analytics section to retrieve | overview |
| chat_type | No | Filter rankings by chat type | |
| community | No | Filter by community name or ID | |
| date_from | No | Custom start date (ISO 8601) — overrides date_range | |
| date_range | No | Preset date range (default: "30d") | |
| target_session | No | Session ID for multi-account routing | |
| include_transcriptions | No | Include audio transcriptions in exports (default true) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and openWorldHint=true, so safety is covered. The description adds real behavioral context beyond that: per-section prerequisites (chat_id for chat_detail, format for exports), approximate payload size (~5KB for overview), and pagination scope — but says nothing about rate limits or export delivery mechanics.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the one-line scope, then the workflow tip, then the section reference and examples — a sensible order. The 22-entry section list is long but each entry adds semantic detail the enum alone cannot convey, so it largely earns its space.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 22-mode tool with no output schema, it covers prerequisites, pagination, filtering, exports, and approximate result sizes well. It stops short of describing the shape of export output (e.g., whether a file or inline payload is returned), which is the remaining gap given no output schema exists.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the description goes further by mapping parameters to sections (label filter on 'activity', chat_type on 'rankings', limit on paginated sections) via worked examples. This meaningfully clarifies how the 14 optional parameters interact with the section enum.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Get WhatsApp analytics data') and enumerates the exact scope (KPIs, activity, rankings, response times, exports, growth history). An agent can distinguish this from siblings like summarize_conversation or query purely from the description.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit entry-point guidance ('Start with section="overview" (default) ... then drill into specific sections') plus six concrete invocation examples covering filtering, chat_detail, and exports. It does not, however, say when to prefer this tool over the sibling export_contacts or query, leaving one overlap unresolved.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_api_infoARead-only
Get HTTP REST API connection info for programmatic access without MCP overhead. Returns URL, auth token, and available endpoints.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover the safety profile (readOnlyHint=true, openWorldHint=false), so the bar is lower. The description contributes beyond them by disclosing the return payload (URL, auth token, available endpoints) and flagging that a credential is exposed, which is useful for a tool with no output schema.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences, front-loaded with the action and followed by the return contents. Every clause earns its place with no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no input parameters and no output schema, the description carries the burden of describing returns and does so by enumerating URL, token, and endpoints. It is nearly complete; a note on credential handling or endpoint format would close the remaining gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so the baseline is 4. Schema coverage is 100% and there is nothing to document; the description correctly says nothing about inputs.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource (get HTTP REST API connection info) and explains the distinguishing value proposition of bypassing MCP overhead. No sibling tool in the list provides connection info, so an agent can immediately tell this apart from the domain tools like query or get_contact.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
"For programmatic access without MCP overhead" gives a clear condition for when to reach for this tool over staying within MCP. It does not name an explicit alternative or state when-not to use it, but the triggering context is unambiguous.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_contactARead-only
Look up a single WhatsApp contact by JID or phone number. Pass either parameter — both work. If multiple raw contacts share the same phone (label dupes), the saved variant wins.
Examples: By JID: { jid: "5491155550001@c.us" } By phone with +: { phone: "+5491155550001" } By phone bare: { phone: "5491155550001" }
| Name | Required | Description | Default |
|---|---|---|---|
| jid | No | Full WhatsApp JID (e.g. "5491155550001@c.us") | |
| phone | No | Phone number — leading "+" and "00" are stripped during matching |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and openWorldHint=false, so the safety profile is covered. The description adds a genuinely useful behavioral rule beyond the structured data — the tie-break that the saved variant wins when duplicate raw contacts share a phone — which an agent could not infer from schema or annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Purpose is front-loaded, the tie-break caveat follows, and the examples come last. Three examples is slightly redundant (the bare-phone case largely repeats the leading-+ case), but nothing is wasted and the structure is easy to scan.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-required-param read tool with annotations covering safety, the description is largely adequate. With no output schema, however, it never describes what the returned contact contains or what happens when no match is found, which leaves a real gap for a lookup tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3. The description goes slightly beyond the schema by explicitly confirming that either parameter alone suffices, and the dedupe tie-break clarifies how phone matching resolves when duplicates exist, plus concrete call examples.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('look up a single WhatsApp contact') and names both accepted identifiers (JID or phone number). The word 'single' implicitly separates it from the sibling list_contacts, but no sibling is named outright.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied by the nature of a lookup tool, and the description does clarify that either identifier parameter is valid ('both work'). However, it never states when to reach for this instead of list_contacts or export_contacts, nor any exclusions or prerequisites.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_contact_groupsARead-only
List the WhatsApp groups a specific contact participates in. Reads from cached chat metadata — no network.
Examples: Groups for contact: { jid: "5491155550001@c.us" }
| Name | Required | Description | Default |
|---|---|---|---|
| jid | Yes | Contact JID — must be the full @c.us form (phone alone not accepted here) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and openWorldHint=false, so the safety profile is covered. The description adds genuine behavioral context beyond that: it reads from cached chat metadata with no network call, implying data freshness may lag reality — useful for an agent deciding whether to trust the result.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two front-loaded sentences plus a compact example; the scoping constraint and the no-network caveat both appear before the example. No filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Single fully-documented parameter, no output schema needed, and annotations carry the safety profile. The cache/no-network note covers the main residual concern; only the absence of explicit sibling routing keeps it short of a 5.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and the schema already documents the @c.us JID format requirement. The example call reinforces the format but adds no meaning beyond what the schema states, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (list) and resource (WhatsApp groups) scoped to a single contact, which implicitly separates it from the sibling list_groups. However, it never names the sibling it contrasts with, so the agent must infer that distinction.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies the use case (find groups a given contact is in) and gives an example call, but offers no explicit when-to-use/when-not guidance or routing to alternatives like list_groups or get_group. Usage is inferable but not stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_groupARead-only
Fetch a single group with a LIVE participant list. Forces Store.GroupMetadata.update() against the WA backend before reading, so the result reflects current membership including recent joins/leaves.
Use this when accuracy matters; use list_groups for browsing.
Examples: Live group fetch: { jid: "120363421729019499@g.us" }
| Name | Required | Description | Default |
|---|---|---|---|
| jid | Yes | Group JID — must end in "@g.us" |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations cover the safety profile (readOnlyHint, openWorldHint), but the description adds real behavioral context: it forces a Store.GroupMetadata.update() round-trip against the WA backend before reading, implying freshness guarantees and added latency versus a cached read. It does not mention rate limits or failure modes, so not a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short blocks: what it does, when to prefer it, and a concrete invocation example. The freshness guarantee is front-loaded and no sentence is filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a one-parameter single-resource read with no output schema, the definition covers what is returned (current membership including recent joins/leaves), the freshness behavior, and the alternative tool. Nothing needed to call it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and the single jid parameter is fully documented with format constraints ('must end in @g.us'). The description's example instantiation is a minor convenience but adds no syntax or semantics beyond the schema, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Specific verb (fetch) plus resource (a single group) plus a distinguishing scope qualifier: LIVE participant list. It explicitly separates itself from list_groups, so an agent can route correctly without opening either schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
States the selection condition directly: 'Use this when accuracy matters; use `list_groups` for browsing.' The alternative tool is named and the trade-off that selects between them is spelled out.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_contactsARead-only
List WhatsApp contacts from the encrypted DBR3 cache (no network). Results are deduplicated by phone number — the same person across multiple labels collapses to one row. Saved contacts sort before unsaved, then alphabetically by display name. WhatsApp "@lid" privacy identifiers (opaque, non-dialable) are always excluded — only real phone-backed contacts are returned.
Examples: All contacts: {} Saved contacts only: { is_my_contact: true } Search by name: { query: "Maria" } Page 2 of 50: { limit: 50, offset: 50 }
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | Max contacts to return (default 50, max 500) | |
| query | No | Case-insensitive substring matched against name, pushname, phone, and JID | |
| offset | No | Skip N contacts for pagination (default 0) | |
| is_my_contact | No | If true, only contacts saved in the user address book. If false, only un-saved contacts. Omit to include both. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Goes well beyond the readOnlyHint/openWorldHint annotations by disclosing deduplication by phone number, the sort order (saved first, then alphabetical), and the exclusion of non-dialable @lid identifiers. This tells the agent exactly what shape of data comes back and what is silently filtered out.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core operation, then behavior, then a compact example block. Every sentence carries information and the formatting makes it scannable.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description carries the return-value burden and does so: dedup, ordering, and exclusion rules are all stated. An agent has everything needed to call it and interpret the result.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is already 100%, so the baseline is 3, but the examples add real value by showing how is_my_contact, query, and limit/offset are combined in practice (e.g., {limit: 50, offset: 50} for page 2). It stops short of documenting edge behavior such as how limit/offset interact with the deduplication pass.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('List WhatsApp contacts') and immediately scopes it to the local encrypted DBR3 cache with no network. This clearly separates it from get_contact (singular lookup), export_contacts (dump), and query (search tool) among siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The four worked examples cover the main invocation modes (unfiltered, saved-only, name search, pagination), which effectively communicates when to use each parameter combination. However, it never states when NOT to use this tool or points to an alternative such as get_contact for a single known contact.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_groupsARead-only
List all WhatsApp groups the user belongs to. Reads from the cache — no network.
For a live snapshot of a single group, use get_group (it forces a fresh GroupMetadata.update).
Examples: All groups: {} Search by group name: { query: "family" } Top 10: { limit: 10 }
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | Max groups to return (default 50, max 500) | |
| query | No | Case-insensitive substring match against group name |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and openWorldHint=false, so safety is covered. The description adds genuinely useful behavior beyond that: reads come from cache with no network call, implying possible staleness, and get_group forces a fresh update. Minor gap in not describing return shape or ordering.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences plus three compact examples, with the cache/network constraint and the sibling routing front-loaded. Every element earns its place; the examples are short and directly actionable.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple two-param read tool with full schema coverage and read-only annotations, the description covers the essentials plus the cache-vs-live distinction. No output schema exists and the description doesn't sketch the returned group fields, a small remaining gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so both parameters are documented in the schema and baseline would be 3. The concrete examples ('{ query: "family" }', '{ limit: 10 }') demonstrate how the parameters compose in practice, adding practical meaning beyond the field descriptions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('List all WhatsApp groups the user belongs to') and explicitly distinguishes itself from the sibling get_group by contrasting cache vs live reads. An agent can route between the two without opening either schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit when-to-use context (cached, no network) and names the alternative tool with the condition that selects it (live snapshot of a single group via get_group forcing GroupMetadata.update). Nothing is left to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
manage_chatA
Manage chat state: archive, unarchive, mark as read/unread, pin, unpin, mute, unmute, set/clear draft.
Actions: archive - Archive a conversation unarchive - Unarchive a conversation mark_read - Mark a conversation as read mark_unread - Mark a conversation as unread pin - Pin a conversation (max 3 pinned) unpin - Unpin a conversation mute - Mute notifications (use mute_duration for duration) unmute - Unmute notifications set_draft - Set a draft message in the compose box (requires text) clear_draft - Clear the draft message
| Name | Required | Description | Default |
|---|---|---|---|
| text | No | Draft text (required for "set_draft" action) | |
| action | Yes | Chat action to perform | |
| mute_duration | No | Mute duration (only for "mute" action). Default: "forever" | |
| target_session | No | Session ID for multi-account routing | |
| conversation_id | Yes | The conversation ID |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare the safety profile (not read-only, not destructive, not idempotent, open-world), so the description doesn't need to restate it. It adds a couple of genuine behavior facts — the 3-pin cap and mute's default of 'forever' — but says nothing about permissions, reversibility, or partial-failure behavior across the ten distinct actions.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded summary sentence followed by a compact action table; every line maps to a real enum value and earns its place. Slightly longer than strictly necessary given the schema already enumerates the actions, but the pairing of action-to-meaning is efficient.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 5-parameter dispatcher with no output schema and full schema coverage, the description covers all action semantics and the key per-action constraints. Missing only cross-cutting details like auth/session routing for target_session and what happens on invalid action/parameter combinations.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, and the two enums (action, mute_duration) are fully documented in the schema, so the baseline is 3. The description's per-action lines add only marginal meaning beyond the schema, mostly repeating the conditional requirements already stated in parameter descriptions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a clear verb+resource ('Manage chat state') and then enumerates all ten concrete actions, so an agent knows exactly what surface this tool covers. It does not, however, distinguish itself from any sibling tool — though none of the listed siblings appear to overlap with chat-state mutation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Per-action annotations give useful constraints ('max 3 pinned', mute uses mute_duration, set_draft requires text), which is real usage guidance. But there is no statement of when to reach for this tool versus alternatives, nor any preconditions or failure conditions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
manage_labelsADestructive
Manage WhatsApp Business labels. Requires a WhatsApp Business account.
Actions: add - Add a label to a conversation (requires label_name/label_id + conversation_id) remove - Remove a label from a conversation (requires label_name/label_id + conversation_id) create - Create a new label (requires label_name) delete - Delete a label (requires label_name or label_id)
| Name | Required | Description | Default |
|---|---|---|---|
| action | Yes | Label action to perform | |
| label_id | No | Label ID (alternative to label_name for add/remove/delete) | |
| label_name | No | Label name (for add/remove/create/delete) | |
| conversation_id | No | Conversation ID or array of IDs (required for add/remove) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true, readOnlyHint=false, idempotentHint=false, and openWorldHint=true, so the safety profile is covered. The description adds genuine context beyond that: the WhatsApp Business account prerequisite and the per-action behavior (e.g., create vs. delete operating on different objects). It does not state that delete is irreversible, but the annotation carries that signal.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the purpose and prerequisite in one line, followed by a compact action table mapping each action to its required parameters. Every line earns its place; the formatting is scannable rather than padded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With 4 parameters, annotations covering the safety profile, and no output schema, the description supplies the missing prerequisites and action-to-parameter mapping an agent needs to invoke correctly. Only the return/result behavior is unaddressed, a minor gap for a mutation tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3, but the description adds cross-parameter semantics the schema cannot express: which parameters each action requires (add/remove need label_name/label_id + conversation_id; create needs label_name). That mapping is valuable guidance beyond the per-field descriptions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb+resource ('Manage WhatsApp Business labels') and then enumerates the four concrete actions (add, remove, create, delete), so an agent knows exactly what the tool does. It is clearly distinct from siblings like manage_lists, manage_notes, and manage_reminders, though it never explicitly contrasts itself with them.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied through the prerequisite ('Requires a WhatsApp Business account') and the per-action parameter requirements, which tell the agent how to invoke each mode. However, there is no explicit when-to-use vs. alternative guidance and no statement about which contexts call for this tool over siblings.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
manage_listsADestructive
Manage personal chat lists (custom lists). These are the personal account equivalent of Business labels. Lists allow organizing chats into custom categories like "Family", "Work", etc. Not available on all accounts — check with action "list" first to see if lists are enabled.
Actions: list - List all custom lists (also shows if feature is enabled) get - Get a list and its associated chats (requires id or name) create - Create a new list (requires name, optional conversation_id for initial chats) edit - Edit a list name or replace its chats (requires id or name) delete - Delete a list (requires id or name) add_chat - Add conversation(s) to a list (requires id/name + conversation_id) remove_chat - Remove conversation(s) from a list (requires id/name + conversation_id)
Examples: List all: { action: "list" } Create: { action: "create", name: "Family", conversation_id: ["number@c.us"] } Add chat: { action: "add_chat", name: "Family", conversation_id: "number@c.us" }
| Name | Required | Description | Default |
|---|---|---|---|
| id | No | List ID (for get/edit/delete/add_chat/remove_chat) | |
| name | No | List name (for create/edit, or to resolve by name) | |
| action | Yes | List action to perform | |
| target_session | No | Session ID for multi-account routing | |
| conversation_id | No | Chat ID or array of IDs to add/remove |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true, idempotentHint=false, and readOnlyHint=false, so the safety profile is covered. The description adds genuinely non-structured context: the feature is account-gated and requires a probing 'list' call to detect availability. It stops short of describing deletion consequences or permission needs, but it meaningfully exceeds the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with purpose, then a compact action table where each line earns its place, then minimal examples. Slightly long overall but no filler; the action list is the core content an agent needs for a multiplexed tool.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a seven-action multiplexed tool with no output schema, the description documents every action's requirements and provides worked examples covering read, create, and mutation paths. Combined with full schema coverage and annotations, an agent has everything needed to invoke it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so baseline is 3. The description goes further by mapping each action to its required inputs (get/create/edit need name or id; add_chat needs conversation_id), which the flat schema cannot express. It doesn't clarify the single-vs-array conversation_id semantics beyond the schema, keeping it below 5.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource (manage personal chat lists) and distinguishes it precisely from the sibling manage_labels by calling it the personal-account equivalent of Business labels. An agent can tell what this tool does and how it differs without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear context (organizing chats into categories) and a concrete prerequisite: 'Not available on all accounts — check with action "list" first.' Each action's use is implied by its requirement list, but there is no explicit when-not-to-use guidance beyond the labeling distinction.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
manage_notesAIdempotent
Manage contact notes. Requires a WhatsApp Business account with notes enabled.
Actions: get - Read the note for a contact set - Write/update the note for a contact
| Name | Required | Description | Default |
|---|---|---|---|
| note | No | Note text (required for "set" action) | |
| action | Yes | Note action to perform | |
| contact_id | Yes | The contact ID |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=false, idempotentHint=true, destructiveHint=false and openWorldHint=true, so the write-safety profile is covered. The description adds genuinely useful context beyond that: the WhatsApp Business account requirement and that 'set' is a write/update (upsert) operation rather than an append, which matches the idempotent hint.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences plus a two-item action list; the purpose and prerequisite are front-loaded and every line earns its place. No filler or repetition.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a small 3-parameter mutation tool with full annotation coverage and no output schema, the description covers the resource, the actions, the precondition, and the required input for 'set'. It could say what 'get' returns when no note exists or confirm that 'set' overwrites, but nothing critical for correct invocation is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3. The description notes that the note text is required for 'set', which merely restates the schema's own 'required for "set" action' note, and it adds no format, length, or contact_id semantics.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific resource (contact notes) and enumerates both supported actions (get/set) with one-line semantics, so an agent immediately knows what the tool does. It does not name or contrast any sibling (manage_labels, manage_chat) explicitly, but the notes scope is unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied through the action list (read vs write a contact note) and the stated prerequisite of a WhatsApp Business account with notes enabled. There is no explicit guidance on when to prefer this over siblings such as get_contact or manage_labels, and no statement of when each action is inappropriate.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
manage_remindersADestructive
Manage personal reminders. Reminders are stored in the cloud and trigger notifications via the Kaption extension.
Actions: list - List all reminders get - Get a specific reminder by ID create - Create a new reminder (requires title + datetime) update - Update a reminder (requires id, optional title/datetime) delete - Delete a reminder (requires id) complete - Mark a reminder as completed (requires id) uncomplete - Mark a reminder as not completed (requires id)
Examples: List all: { action: "list" } Create: { action: "create", title: "Follow up with client", datetime: "2026-03-07T14:00:00Z" } Complete: { action: "complete", id: "rem_abc123" }
| Name | Required | Description | Default |
|---|---|---|---|
| id | No | Reminder ID (required for get/update/delete/complete/uncomplete) | |
| title | No | Reminder text (required for create, optional for update) | |
| action | Yes | Reminder action to perform | |
| filter | No | Filter for list action. Default: "active" (non-completed only) | |
| datetime | No | ISO 8601 datetime for the reminder (required for create, optional for update) | |
| target_session | No | Session ID for multi-account routing | |
| notification_type | No | How to notify. Default: "automatic" |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true, openWorldHint=true and idempotentHint=false, and the description usefully adds that reminders live in the cloud and fire notifications through the Kaption extension. However, it never distinguishes the safety profile of the read actions (list/get) from the destructive ones (delete), nor mentions permissions, reversibility of delete, or rate limits.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded one-line summary followed by a scannable action table and three concrete examples; every line carries information. It is slightly redundant with the schema's own parameter descriptions, which keeps it short of a 5.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a seven-parameter dispatch tool with no output schema, the description covers the action surface, required inputs, examples, and the notification side effect. It omits what list/get return and how errors surface, but the 100%-covered schema carries the parameter burden well.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents every parameter including the filter and notification_type enums with defaults. The description restates the same per-action requirements and adds only format examples (ISO 8601 datetime, id shape), which is marginal added value.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific resource ('personal reminders') and enumerates all seven operations with their required inputs, so an agent knows exactly what capability surface this tool exposes. No sibling tool touches reminders, so the scope is unambiguous within the toolset.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The action list effectively tells the agent which operation to pick for each intent (list vs get vs create vs complete), and states per-action required parameters such as title+datetime for create. It stops short of any exclusion guidance or advice on when a reminder action is inappropriate relative to other tools.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
manage_scheduled_messagesADestructive
Schedule WhatsApp messages to be sent automatically at a specific time, in one of two modes: bot (default) - sent from Kaption's WhatsApp number, not the user's own. Stored in Kaption's cloud and sent even when this computer is off. One-to-one chats only (no groups); one line, max 800 characters. local ("From this computer") - sent from the user's own WhatsApp number, as them, while this computer and WhatsApp are open. Stays on this device. Kaption keeps it within safe limits automatically (a few messages an hour and a day, minutes apart, only to chats where the other person has written, at most 3 a day to groups) and refuses what does not fit. The only mode that can send to groups (ones the user can post in and posted in within 30 days). Text only. A message whose time passes while this computer is off is missed, never sent late on its own. The local mode must first be turned on by the user in Kaption (its "From this computer" option, after reading the risks); an assistant cannot turn it on. A refused local request says why and ends with a reason code, e.g. "(reason: per-day)".
Actions: list - List scheduled messages of both modes (each has "mode"); pass mode to list only one get - Get a specific scheduled message by ID create - Schedule a new message (requires message + datetime + conversation_id) update - Change the message and/or datetime (requires id) delete - Cancel/delete a scheduled message (requires id); in local mode it cancels a waiting message and removes a finished one cancel - Local mode only: cancel a waiting, missed or failed message remove - Local mode only: remove a finished message from the list send_now - Local mode only: send a missed or failed message now (still within the limits) Pass mode "local" for every action on a local message.
Examples: List all: { action: "list" } Schedule (Kaption bot): { action: "create", message: "Hey, just following up!", datetime: "2026-03-07T09:00:00Z", conversation_id: "5491157390064@c.us" } Schedule from the user's own number: { action: "create", mode: "local", message: "Running 10 min late", datetime: "2026-03-07T09:00:00-03:00", conversation_id: "5491157390064@c.us" } Schedule to a group: { action: "create", mode: "local", message: "Standup moved to 10", datetime: "2026-03-07T09:00:00Z", conversation_id: "120363000000000000@g.us" } Cancel: { action: "delete", id: "msg_abc123" }
| Name | Required | Description | Default |
|---|---|---|---|
| id | No | Scheduled message ID (required for get/update/delete/cancel/remove/send_now) | |
| mode | No | "bot" (default): sent from Kaption's WhatsApp number. "local": sent from the user's own WhatsApp number while this computer and WhatsApp are open, within safe limits kept automatically; needed for groups. For list, omit to get both modes | |
| action | Yes | Scheduled message action to perform (cancel, remove and send_now need mode "local") | |
| filter | No | Filter for list action. Default: "pending" (not sent yet; in local mode also missed or failed ones waiting for the user) | |
| message | No | Message text to send (required for create, optional for update). Bot: one line, max 800 characters. Local: text only, up to 2000 characters | |
| datetime | No | ISO 8601 datetime when the message should be sent (required for create, optional for update) | |
| target_session | No | Session ID for multi-account routing | |
| conversation_id | No | Chat to send the message to (required for create): a person, or a group with mode "local" | |
| notification_type | No | Bot mode only. Notification type. Default: "automatic" |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Despite annotations already flagging destructive/openWorld, the description adds substantial behavioral detail: local-mode rate limits, 'never sent late' semantics, missed/refused request behavior with reason codes, user-only enablement, and per-action cancel/remove distinctions. This is well beyond what annotations or schema convey.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The text is long but justified for an 8-action, dual-mode tool. Purpose and mode trade-offs are front-loaded, and the action list plus examples each carry distinct information. Some overlap with schema parameter descriptions keeps it from a 5.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description still covers behavior thoroughly and hints at list/get return shape ('each has "mode"'), plus reason codes. Given the complexity (9 params, 4 enums, 2 modes), it is nearly complete, though it does not describe the structure of list/get results in detail.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so baseline is 3, but the description contributes extra meaning: which parameters each action requires, the mode-per-action constraint ('Pass mode local for every action on a local message'), and concrete example payloads with real values. It reinforces rather than merely repeats the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The opening sentence gives a specific verb+resource+scope: scheduling WhatsApp messages for automatic send at a specific time, split into two named modes. An agent can immediately distinguish this from siblings like manage_reminders or manage_chat without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives explicit when-to-use guidance for the bot vs local modes (privacy, off-computer delivery, group support) and states the prerequisite that local mode must be enabled by the user first. It does not, however, compare against sibling tools, so it stops short of a full alternative-routing treatment.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
queryARead-only
Query WhatsApp data: conversations, contacts, messages, transcriptions, labels, and communities. Supports listing, searching, filtering, and looking up by ID.
IMPORTANT: Multiple WhatsApp accounts may be connected (e.g. personal + business). Always query entity="session" FIRST to see all connected accounts and their session IDs. Then use target_session to route queries to the correct account. Each account has different conversations, contacts, and messages.
HOW TO READ MESSAGES: To get messages from a specific conversation, pass its id (e.g. "5491157390064@c.us"). This returns the conversation info WITH its messages. Use limit to control how many. Do NOT use entity="messages" for this — that is for global text search only.
AUDIO TRANSCRIPTIONS: To get audio transcriptions, use entity="transcriptions" with an optional query. Or pass a conversation id to see messages (audio messages include transcription text).
Examples: List sessions: { entity: "session" } List conversations: {} Target specific account: { entity: "conversations", target_session: "sess_abc123" } Read messages: { id: "5491157390064@c.us" } Read last 100 msgs: { id: "5491157390064@c.us", limit: 100 } Search globally: { query: "meeting" } Search in chat: { id: "5491157390064@c.us", query: "meeting" } Unread conversations: { unread: true } Search contacts: { query: "Alice", entity: "contacts" } List labels: { entity: "labels" } Filter by label: { label: "Important", entity: "conversations" } List communities: { entity: "communities" } Filter by community: { community: "My Community", entity: "conversations" } Find which groups a contact is in: { id: "5491157390064@c.us", entity: "contacts", include_participants: true } List members of a group: { entity: "contacts", group: "120363421729019499@g.us" }
| Name | Required | Description | Default |
|---|---|---|---|
| id | No | Look up a specific conversation, contact, or label by ID | |
| list | No | Filter conversations by list name or ID (Personal accounts) | |
| after | No | Return messages after this ISO 8601 datetime (e.g. "2026-03-01T12:00:00.000Z") for incremental sync | |
| group | No | Filter contacts by group ID — only return contacts that are members of this group | |
| label | No | Filter conversations by label name or ID (Business accounts) | |
| limit | No | Max results (default 25, max 5000) | |
| query | No | Text to search for (names, messages, transcriptions) | |
| before | No | Return messages before this ISO 8601 datetime (e.g. "2026-03-01T12:00:00.000Z") for cursor-based pagination backward | |
| entity | No | Entity type to query. Defaults to "conversations" when listing, or all when searching. Use "session" to list all connected WhatsApp accounts. | |
| unread | No | Only return conversations with unread messages | |
| community | No | Filter conversations by community name or ID | |
| exclude_muted | No | Exclude muted conversations from listings (default false) | |
| target_session | No | Session ID to target a specific WhatsApp account. Get session IDs from entity="session". If omitted, routes to the most recently active account. | |
| exclude_archived | No | Exclude archived conversations from listings (default true) | |
| include_participants | No | Include group participants in results. Useful when looking up a contact by ID to see which groups they belong to, or when querying a group to see its members. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already establish readOnlyHint=true and openWorldHint=true, so safety is covered. The description adds real behavioral context beyond that: multi-account routing semantics (omitted target_session routes to the most recently active account) and return shape ('returns the conversation info WITH its messages'). It omits rate limits and pagination behavior, keeping it short of a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with purpose, then sectioned under HOW TO READ MESSAGES, AUDIO TRANSCRIPTIONS, and Examples, which suits a 15-parameter polymorphic tool. The length is largely justified, though the example block restates some guidance already given in prose, adding mild redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a complex, zero-required-parameter, multi-entity tool with no output schema, the description supplies the entity model, session-routing prerequisite, and worked call shapes an agent needs. It could be more complete on return structure for non-message entities and on pagination, but nothing critical to invoking it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3. The description exceeds it by demonstrating parameter composition through worked examples that the schema cannot convey — e.g. that passing id alone returns a conversation with its messages, and how label vs list map to business vs personal accounts. This meaningfully reduces the chance of mis-combining the 15 optional parameters.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Query') and enumerates the exact resource set (conversations, contacts, messages, transcriptions, labels, communities) plus the operations supported (listing, searching, filtering, lookup by ID). It clearly distinguishes its internal modes, but it never differentiates itself from overlapping siblings such as list_contacts, get_contact, list_groups, or get_group, which an agent must resolve on its own.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Offers unusually strong operational guidance: query entity="session" FIRST, then route via target_session, and an explicit exclusion ('Do NOT use entity="messages" for this — that is for global text search only'). The extensive example list shows when each mode applies. The gap is that it gives no guidance for choosing this tool over the overlapping sibling tools like list_contacts or get_group.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
summarize_conversationCRead-only
Get or generate a summary of a conversation
| Name | Required | Description | Default |
|---|---|---|---|
| message_count | No | Number of messages to use for summary generation (default 50, max 500) | |
| target_session | No | Session ID to target a specific WhatsApp account | |
| conversation_id | Yes | The conversation ID |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and openWorldHint=true, so the safety profile is covered. The description adds nothing beyond that: it does not say whether summaries are cached or freshly generated, whether generation is expensive/slow, what happens when no summary exists, or whether an existing summary is overwritten when regenerated. For a dual-mode "get or generate" tool this is a meaningful gap, though it does not contradict the readOnly annotation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single short sentence with no padding and no buried information. It is efficient, though the brevity is partly achieved by omitting behavior rather than by tight editing.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With annotations covering the read-only/open-world profile and a fully documented schema, the remaining burden on the description is modest. Still, the get-vs-generate distinction, cost/latency implications, and any caching behavior are unaddressed, which is the one thing an agent most needs to know here.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all three parameters (conversation_id, message_count, target_session) are already fully documented with defaults, bounds, and meaning. The description adds no parameter-level meaning beyond the schema, which is the expected baseline when the schema does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The verb+resource pair (summarize a conversation) is identifiable, and it is not confusable with siblings like get_analytics or query. However, "Get or generate" leaves a real ambiguity: it is unclear whether this is a pure read of a cached summary, a request that triggers generation, or both depending on state. That ambiguity is central to what the tool does, so this lands at minimum-viable rather than clear.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No when-to-use guidance, no when-not-to-use, and no mention of any sibling tool (e.g. get_analytics, query) that might compete for the same intent. The agent must infer that this is the right tool purely from the name.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
18 tool updates
v0.5.6- First observed
call_recordings - First observed
download_media - First observed
export_contacts - First observed
get_analytics - First observed
get_api_info - First observed
get_contact - First observed
get_contact_groups - First observed
get_group - First observed
list_contacts - First observed
list_groups - First observed
manage_chat - First observed
manage_labels - First observed
manage_lists - First observed
manage_notes - First observed
manage_reminders - First observed
manage_scheduled_messages - First observed
query - First observed
summarize_conversation
TDQS
Scored across 18 tools
Most tools have distinct purposes, but `query` is a mega-tool whose entities (contacts, conversations, groups, labels, communities) overlap heavily with specialized tools like `list_contacts`, `get_contact`, `get_contact_groups`, `list_groups`, and `get_group`. Additionally `get_analytics` embeds an `export_contacts` section that duplicates the standalone `export_contacts` tool, and `manage_labels` vs `manage_lists` are parallel business/personal variants.
Strong verb_noun pattern throughout (get_*, list_*, export_*, manage_*, download_*, summarize_*). Two deviations: the bare `query` with no noun, and the noun-only `call_recordings` with no verb, but overall the convention is predictable.
18 tools is on the heavier side but reasonable given the breadth of the domain (messaging, contacts, groups, analytics, media, reminders, scheduling, lists, labels, notes, recordings). Each tool maps to a coherent capability area.
Broad lifecycle coverage across reads, search, media download, chat state, labels/notes/lists, reminders, scheduling, analytics, and call recordings. Notable gaps: no direct/immediate message-send tool (only scheduling) and no group-membership modification (add/remove participants), though most workflows can be worked around.
Maintenance
Related MCP Connectors
Use WhatsApp from AI apps through the Kaption extension. OAuth 2.1 relay that cannot read messages.
11Let Claude or ChatGPT search, read and send your WhatsApp messages over MCP. OAuth sign-in.
WhatsMCP connects Claude and other MCP-compatible AI agents directly to WhatsApp. Send and receive text, images, documents, and voice notes; manage groups (create, add/remove members, promote admins); look up contacts and profiles; follow channels; and read call and message history — all through a standard MCP interface. For voice use cases, WhatsMCP offers SIP-based calling plans (inbound-only, or full inbound/outbound) so AI voice agents can answer and place WhatsApp calls, plus low-latency WebSocket integrations with voice agent providers like ElevenLabs. Multiple WhatsApp accounts can be paired and managed per workspace, with webhook support for real-time inbound message delivery to your own infrastructure.
WhatsApp CRM for AI agents: search contacts, read chats, manage the sales pipeline, send messages.
Related MCP Servers
- FlicenseNot gradedqualityNot gradedmaintenanceEnables seamless integration with WhatsApp through the Model Context Protocol, featuring multi-user support and Supabase cloud storage for persistent message history and media. Users can send messages, search chat records, and manage contacts across platforms like Claude Desktop, Cursor, and OpenClaw.-
- FlicenseNot gradedqualityDmaintenanceSend WhatsApp messages, manage templates, and run broadcast campaigns from AI coding tools. Works with Cursor, Claude Code, and Codex via npx @gaviwhatsapp/mcp.-
- AlicenseNot gradedqualityBmaintenanceTurns Claude Code, Claude Desktop, Cursor, Windsurf or ChatGPT into a WhatsApp operator that knows your customers, your templates, your wallet, and your funnel.15 npm3MIT
- AlicenseNot gradedqualityBmaintenanceEnables Claude to interact with WhatsApp: read chats, search messages, send messages with a mandatory confirmation step, and transcribe voice notes locally, all with encrypted storage and prompt-injection scrubbing.2MIT