chatwoot-mcp
This server exposes Chatwoot analytics and management data to AI assistants via MCP, mostly through read-only tools, with optional local DuckDB storage and Jev lead scoring.
Get account-level summaries: conversations, response times, resolution rate, CSAT, bot performance, and conversation traffic heat maps.
Monitor agents, teams, inboxes, and labels with performance metrics such as open/resolved conversations, response times, and CSAT.
List and search conversations, retrieve full conversation details and message histories.
Access Captain AI assistant metrics: overview, FAQ stats, resolution flow, resolution trends, and AI-generated conversation summaries.
Analyze CSAT responses and aggregated CSAT metrics with filters by period, rating, inbox, and team.
View real-time conversation counts and grouped metrics by team or agent.
Search contacts, get contact details, and view a contact's conversation history.
Generate advanced reports like inbox-label matrix and first-response-time distribution.
Sync a recent window of conversations/contacts/messages into a local DuckDB store and run read-only SQL queries.
Score leads by propensity to enroll using Jev, list hot leads, view lead profiles, and backtest lead scoring accuracy.
Provides read-only access to Chatwoot analytics and data, enabling tools to retrieve conversations, agent and inbox performance, team metrics, labels, CSAT responses, Captain AI insights, real-time metrics, and contact details.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@chatwoot-mcpCan you give me a summary of our support performance over the past week?"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
chatwoot-mcp
An MCP (Model Context Protocol) server that exposes Chatwoot analytics data to AI assistants. All tools are read-only, except chatwoot_get_conversation_summary which triggers a Captain LLM call on the Chatwoot server.
Features
39 MCP tools covering conversations, agents, inboxes, teams, labels, CSAT, Captain AI, real-time metrics, and contacts
Local DuckDB storage: sync a recent window of conversations/contacts/messages for offline analytics
Jev lead scoring: score contacts by propensity to enroll using Jev (TypeSafe System One) via OpenRouter, with built-in calibration
Configurable domain profile: generic defaults; map an account's own attribute keys, labels, stages and keywords via a gitignored JSON profile
Two transport modes: stdio for Claude Desktop, HTTP for claude.ai remote connector
Multi-tenant HTTP mode: each user provides their own Chatwoot credentials per request; each tenant gets an isolated DuckDB file
Dual output format:
markdown(human-readable) orjsonfor all toolsPeriod shortcuts: built-in ranges like
today,7d,30d,90dfor all time-series queries
Related MCP server: Secureframe MCP Server
Requirements
Python 3.11+
uv package manager
A Chatwoot account with API access
Installation
git clone https://github.com/minholi/chatwoot-mcp.git
cd chatwoot-mcp
uv syncConfiguration
Copy .env.example to .env and fill in your credentials:
cp .env.example .envCHATWOOT_URL=https://your-instance.chatwoot.com
CHATWOOT_ACCOUNT_ID=1
CHATWOOT_API_TOKEN=your_api_token_hereYour API token can be found in Chatwoot under Profile Settings → Access Token.
Local storage & lead scoring (optional)
The DuckDB sync and Jev lead scoring features are optional and configured separately:
# Per-tenant DuckDB files live under data/tenants/<hash>.duckdb
CHATWOOT_DB_DIR=./data
# Optional fixed path (overrides CHATWOOT_DB_DIR; useful for stdio single-tenant)
# CHATWOOT_DB_PATH=./data/chatwoot.duckdb
JEV_SYNC_MAX_PAGES=8000
JEV_SYNC_STOP_MARGIN=10
# Throttling — protects the Chatwoot instance during syncs
CHATWOOT_MAX_RPS=3
CHATWOOT_MAX_RETRIES=4
# Jev (TypeSafe System One) via OpenRouter — https://openrouter.ai
OPENROUTER_API_KEY=
OPENROUTER_BASE_URL=https://openrouter.ai
JEV_MODEL=typesafe/jev-1.13Typical flow:
chatwoot_sync_local_data— ingest a recent window into DuckDB.chatwoot_score_leads— compute features and call Jev per lead (billable, cheap).chatwoot_list_hot_leads/chatwoot_lead_scoring_report— consume the ranking.chatwoot_evaluate_lead_scoring— backtest against historical outcomes.
Long backfills: MCP clients time out on multi-minute tool calls, so large first-time syncs are better run as a detached CLI process calling
src.ingest.sync_all(...); the MCP tool is fine for small incremental refreshes. Requests are paced globally byCHATWOOT_MAX_RPSwith retry/backoff, and the DuckDB store is per-tenant and safe to re-run (upserts).
Domain profile (account-specific vocabulary)
Lead scoring relies on an account's own enrollment vocabulary. The tracked defaults are a generic example; supply a JSON profile to map the logical names to a real account's Chatwoot custom-attribute keys, labels, stages and keywords:
{
"institution_name": "Example University",
"attributes": {
"stage": "enrollment_stage",
"situation": "enrollment_status",
"course": "course_code",
"course_name": "course_name",
"modality": "modality"
},
"label_flags": {
"has_enrolled": { "exact": ["enrolled"] },
"has_no_response": { "prefix": ["no-response"] }
},
"keywords": { "kw_enroll": "enroll|register|sign up" }
}Point CHATWOOT_DOMAIN_CONFIG at the file (or set JEV_INSTITUTION_NAME to just
override the deployment name). Only the keys you provide are overridden; the rest
fall back to the defaults. Logical names are stable, so keep the profile out of
version control (the domain/ directory is gitignored).
Usage
Stdio mode (Claude Desktop)
Run the server directly — credentials come from the .env file:
uv run python main.pyTo integrate with Claude Desktop, add this to your claude_desktop_config.json:
{
"mcpServers": {
"chatwoot": {
"command": "uv",
"args": ["run", "python", "main.py"],
"cwd": "/path/to/chatwoot-mcp",
"env": {
"CHATWOOT_URL": "https://your-instance.chatwoot.com",
"CHATWOOT_ACCOUNT_ID": "1",
"CHATWOOT_API_TOKEN": "your_api_token_here"
}
}
}
}HTTP mode (claude.ai remote connector)
Start the server in HTTP mode:
uv run python main.py --transport http --port 8000Each request must include a composite key in the Authorization or X-API-Key header:
Authorization: Bearer https://your-instance.chatwoot.com|account_id|api_tokenThis allows multiple users to connect with their own credentials without server-side configuration.
Available endpoints:
Endpoint | Description |
| Service info |
| Health check |
| MCP protocol (stateless) |
Docker
docker compose up -dThe compose file reads from .env and runs in HTTP mode on http://127.0.0.1:8000 by default. Set MCP_HOST_PORT in .env to change the exposed port.
Available tools
Overview
Tool | Description |
| Account summary: conversation counts, response times, resolution rate, CSAT |
| Bot performance: handoffs, autonomous resolution rate, average response time |
| Heat map of conversation volume by hour and day of week |
Agents
Tool | Description |
| List all agents with ID, email, and availability status |
| Per-agent metrics: open/resolved conversations, response time, CSAT |
Conversations
Tool | Description |
| List conversations with filters (status, assignee, inbox, team, label) |
| Full-text search across conversations and messages |
| Full details of a specific conversation |
| Complete message history of a conversation |
Labels
Tool | Description |
| List all labels with ID, color, and description |
| Performance metrics for a specific label |
Inboxes
Tool | Description |
| List all inboxes (WhatsApp, email, widget, etc.) |
| Performance metrics for a specific inbox |
Teams
Tool | Description |
| List all teams |
| Performance metrics for a specific team |
Captain AI
Tool | Description |
| List AI assistants configured in the account |
| AI assistant performance metrics |
| FAQ usage and effectiveness statistics |
| Period-to-period comparison overview |
| Funnel of resolution steps |
| Resolution rate trend over time |
| AI-generated summary of a conversation (not read-only — triggers a Captain LLM call, billable) |
CSAT
Tool | Description |
| Survey responses with filters (period, rating, inbox, team) |
| Aggregated CSAT summary with breakdown by rating |
Real-time
Tool | Description |
| Real-time counts: open, unattended, unassigned, pending |
| Real-time conversations grouped by team or agent |
Contacts
Tool | Description |
| Search contacts by name, email, phone, or identifier |
| Full contact details |
| Conversation history of a contact |
Advanced reports
Tool | Description |
| Cross-distribution of conversations between inboxes and labels |
| Histogram of first-response-time distribution |
Local storage (DuckDB)
Tool | Description |
| Incrementally sync conversations/contacts/messages into the local DuckDB store |
| Row counts and sync cursors for the current tenant |
| Read-only SQL (SELECT/WITH) against the local store |
Lead scoring (Jev)
Tool | Description |
| Score leads by propensity to enroll with Jev (not read-only — billable OpenRouter calls) |
| Ranked scored leads with filters (band, course, modality, confidence) |
| Local features + latest Jev decision for a contact |
| Ranked lead-scoring report as a tool |
| Backtest persisted scores against historical outcomes (precision/recall/precision@K) |
Development
# Lint and format
uv run ruff check src/ main.py
uv run ruff format src/ main.py
# Run tests
uv run pytestLicense
MIT — see LICENSE.
Available Tools
30 toolschatwoot_get_agent_performanceARead-only
Returns performance metrics for a specific agent.
Includes: open and resolved conversations, average first response time,
average resolution time, and CSAT.
Args:
agent_id: Agent ID (use chatwoot_list_agents to get IDs).
period: Analysis period.
output_format: Output format.
| Name | Required | Description | Default |
|---|---|---|---|
| period | No | 30d | |
| agent_id | Yes | ||
| output_format | No | markdown |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With readOnlyHint=true and destructiveHint=false annotations already establishing a safe read operation, the description adds the metric breakdown but no further behavioral detail such as timezone handling, default-period effects, or empty-result behavior. It is consistent with the annotations and adequate, but not rich.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is short, front-loaded with the tool's purpose, and uses a compact args list; every line contributes something except the two placeholder parameter lines. It earns a small deduction for those tautological arg descriptions.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a three-parameter read-only tool with enums on period and output_format plus an output schema, the key missing piece for correct invocation is agent_id sourcing, which is explicitly covered. The vague period and output_format lines are partly mitigated by schema enums and defaults, leaving only minor ambiguity about what each format produces.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description needed to carry the parameter-explanation burden. It genuinely helps only for agent_id ('use chatwoot_list_agents to get IDs'); 'period: Analysis period' and 'output_format: Output format' are near-tautological placeholders that add no semantic value beyond the schema's enum lists.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb-resource pair ('Returns performance metrics for a specific agent') and itemizes the metrics, which clearly separates it from sibling performance tools such as team, inbox, and label performance. A selecting agent can identify this as the per-agent metrics tool without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives clear context that this tool is for a single agent and points to chatwoot_list_agents as an ID source, but it never says when to prefer this over comparable siblings such as chatwoot_get_team_performance or chatwoot_get_inbox_performance. Usage is implied rather than explicitly scoped against alternatives.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
chatwoot_get_bot_summaryBRead-only
Returns overall bot performance: handoffs, autonomous resolution rate, and average response time.
Args:
period: Analysis period.
output_format: Output format.
| Name | Required | Description | Default |
|---|---|---|---|
| period | No | 30d | |
| output_format | No | markdown |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and destructiveHint=false, covering the safety profile. The description adds useful behavioral context by specifying that this is an aggregate 'overall' summary and naming the computed metrics, which helps the agent set expectations for output shape.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The main sentence is concise and front-loaded, but the Args section is low-value and tautological. The description is compact, but not every line earns its place, and the argument comments could be omitted without losing information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only summary tool with optional parameters, full enum schemas, defaults, and an output schema, the description is reasonably complete. It conveys the tool's purpose and key metrics, though it could clarify how this differs from get_summary or how period boundaries are interpreted.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate for undocumented parameters, but it fails to do so. 'period: Analysis period' and 'output_format: Output format' are essentially restatements of the parameter names and add no semantic value beyond the schema's existing enum values and defaults.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a clear verb and resource: 'Returns overall bot performance' and specifies the exact metrics included: handoffs, autonomous resolution rate, and average response time. This distinguishes it from sibling tools that focus on agents, teams, inboxes, or general summaries.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives no guidance on when to use this tool versus the many sibling performance tools, such as chatwoot_get_summary or chatwoot_get_agent_performance. It only states what the tool returns, leaving the agent to infer appropriate usage from the name and metrics.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
chatwoot_get_captain_faq_statsBRead-only
Returns usage and effectiveness statistics for a Captain assistant's knowledge base (FAQ).
Args:
assistant_id: Assistant ID (use chatwoot_list_captain_assistants).
range: Analysis period in days or month (7, 30, 90, this_month, last_month).
output_format: Output format.
| Name | Required | Description | Default |
|---|---|---|---|
| range | No | 7 | |
| assistant_id | Yes | ||
| output_format | No | markdown |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and destructiveHint=false, so the safety profile is covered and no contradiction exists. The description adds little behavioral context beyond the read-only nature; it does not describe aggregation behavior or response details, though the output schema can carry part of that burden.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Compact and front-loaded: the purpose sentence comes first, fiollowed by a concise Args block. The output_format line is somewhat redundant, but the rest of the description is eifther informative or a useful cross-reference.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only stats tool with an output schema, the invocation parameters are mostly covered and assistant_id sourcing is supplied. However, it does not clarify what counts as 'usage and effectiveness' or how this tool differs from neighboring Captain analytics tools, so the agent's selection confidence remains incomplete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% so the description carries the param semantics burden. It adds useful meaning for assistant_id by pointing to chatwoot_list_captain_assistants, and for range by describing it as an analysis period in days or month. However, output_format is only restated as 'Output format', adding no real meaning beyond the enum.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific action, 'Returns usage and effectiveness statistics', and a specific resource, 'Captain assistant's knowledge base (FAQ)', which is enough to identify its domain among the Captain analytics siblings. It does not explicitly distinguish itself from chatwoot_get_captain_metrics or chatwoot_get_captain_overview, so it misses the top tier of differentiation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Implied usage is clear: call this when you need FAQ usage and effectiveness stats for a Captain assistant, and the assistant_id hint references chatwoot_list_captain_assistants. However, there is no explicit when-not-to-use guidance or alternative routing among the many Captain analytics tools, leaving selection partially to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
chatwoot_get_captain_metricsBRead-only
Returns performance metrics for a Captain (AI) assistant.
Includes autonomous resolution rate, handoffs to humans, and FAQ usage.
Args:
assistant_id: Assistant ID (use chatwoot_list_captain_assistants).
period: Analysis period.
output_format: Output format.
| Name | Required | Description | Default |
|---|---|---|---|
| period | No | 30d | |
| assistant_id | Yes | ||
| output_format | No | markdown |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and destructiveHint=false, so the safety profile is covered. The description adds metric scope but does not disclose deeper behavioral traits such as period semantics, output behavior, or rate limits. It does not contradict the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is concise and front-loaded with the main purpose, followed by a compact Args section. Every sentence is short, though the arg descriptions are terse and some are tautological.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With annotations covering safety and an output schema present, the description does not need to explain return values. However, given the large set of similar captain metric siblings, the lack of any distinction or usage context leaves the overall definition only minimally complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It gives a useful hint for assistant_id by pointing to chatwoot_list_captain_assistants, but period is only described as 'Analysis period' and output_format as 'Output format', which adds little meaning beyond the parameter names. The schema enums are not explained.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool 'Returns performance metrics for a Captain (AI) assistant' and names specific metric categories like autonomous resolution rate, handoffs, and FAQ usage. It is distinct enough from general chatwoot metric tools, but it does not explicitly differentiate itself from sibling Captain tools such as chatwoot_get_captain_overview or chatwoot_get_captain_resolution_trend.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The only usage-related guidance is the assistant_id arg hint to use chatwoot_list_captain_assistants. The description gives no guidance about when to prefer this tool over the many sibling captain-specific tools, nor does it mention exclusions or alternative tools.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
chatwoot_get_captain_overviewARead-only
Returns overview metrics for a Captain assistant with period-to-period comparison.
Args:
assistant_id: Assistant ID (use chatwoot_list_captain_assistants).
range: Analysis period in days or month (7, 30, 90, this_month, last_month).
output_format: Output format.
| Name | Required | Description | Default |
|---|---|---|---|
| range | No | 7 | |
| assistant_id | Yes | ||
| output_format | No | markdown |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and destructiveHint=false, so the description only needs to add behavioral context. It adds period-to-period comparison and period scoping, which is useful, but does not explain what 'overview metrics' actually include or how the comparison is represented. This is acceptable given the read-only annotation and existing output schema.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is short, front-loaded with the purpose, and organized into compact argument bullets. Each line is mostly informative; the only minor redundancy is the output_format bullet, which adds little over the schema's enum and default.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool has read-only annotations and an output schema, so the description does not need to explain return formats. It covers the required assistant_id, the range semantics, and output format choice. The main remaining gap is lack of explicit sibling differentiation, but for a simple read-only overview tool the definition is otherwise complete enough for an agent to call it.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With schema description coverage at 0%, the description must carry parameter meaning. It does well for assistant_id by telling the agent to use chatwoot_list_captain_assistants, and for range by explaining it as an analysis period in days or months. However, output_format is only described as 'Output format,' which merely restates the schema title and adds no semantic value beyond the enum.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a clear verb-resource pair: "Returns overview metrics for a Captain assistant with period-to-period comparison." This differentiates the tool from lower-level captain metrics by emphasizing the overview and comparison behavior, though it does not explicitly name the sibling it should not be confused with.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives context on how to source assistant_id via chatwoot_list_captain_assistants and enumerates valid range periods. However, it provides no guidance on when to choose this tool over siblings such as chatwoot_get_captain_metrics, and there are no explicit exclusions or alternative-selection criteria.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
chatwoot_get_captain_resolution_flowARead-only
Returns the conversation resolution funnel for a Captain assistant (steps to resolution).
Args:
assistant_id: Assistant ID (use chatwoot_list_captain_assistants).
range: Analysis period in days or month (7, 30, 90, this_month, last_month).
output_format: Output format.
| Name | Required | Description | Default |
|---|---|---|---|
| range | No | 7 | |
| assistant_id | Yes | ||
| output_format | No | markdown |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already establish readOnlyHint=true and destructiveHint=false. The description adds that the result is a funnel with 'steps to resolution', which is mild behavioral context, but it does not disclose rate limits, data scope, or edge-case behavior. The safety profile is covered by annotations, so a 3 is appropriate.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and front-loaded: one sentence of purpose followed by clear Arg lines. There is no redundancy or filler, and every line adds necessary information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only, 3-parameter analytics tool with an output schema and schema-provided defaults, the description adequately covers purpose and parameters. It would be stronger with explicit distinction from resolution_trend/overview siblings, but nothing essential to making a basic call is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With schema description coverage at 0%, the description must carry parameter meaning. It usefully explains assistant_id provenance via list_captain_assistants, range as an analysis period with accepted values, and output_format as a format choice. Only output_format is left largely tautological, so a small gap remains.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description identifies a specific resource ('Captain assistant'), a specific verb ('Returns'), and a distinct concept ('conversation resolution funnel / steps to resolution'), which distinguishes it from sibling analytics tools like get_captain_resolution_trend and get_captain_metrics.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no explicit when-to-use or alternative routing. The only usage hint is how to obtain assistant_id via chatwoot_list_captain_assistants, which is parameter acquisition rather than usage guidance. It never says when to prefer this over resolution_trend, overview, or other Captain analytics tools.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
chatwoot_get_captain_resolution_trendARead-only
Returns the resolution rate trend for a Captain assistant over time.
Args:
assistant_id: Assistant ID (use chatwoot_list_captain_assistants).
range: Analysis period in days or month (7, 30, 90, this_month, last_month).
output_format: Output format.
| Name | Required | Description | Default |
|---|---|---|---|
| range | No | 7 | |
| assistant_id | Yes | ||
| output_format | No | markdown |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and destructiveHint=false, so the safety profile is covered. The description adds the time-series context and range semantics, but no deeper behavioral detail such as how the trend is computed or whether this_month/last_month depend on the current calendar date. No contradiction with annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The tool's purpose is front-loaded in one sentence, and the Args block is compact with no filler phrases. The only minor waste is the nearly tautological 'output_format: Output format.' line, but overall the size is appropriately small for a three-parameter read-only tool.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With an output schema present, the description correctly avoids explaining return values, and annotations cover the safety profile. The description covers the tool's purpose and all parameters adequately; the remaining gap is the lack of guidance distinguishing this from the other captain analytics siblings, which the name itself mostly resolves.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description carries the full burden, and it documents all three parameters. assistant_id gets a useful cross-tool pointer, and range is explained with its enum values; only output_format is weakly described as 'Output format', but its markdown/json enum is visible in the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The first sentence names a specific verb ('Returns'), resource ('resolution rate trend'), and scope ('for a Captain assistant over time'). This distinguishes it from closely named siblings like chatwoot_get_captain_resolution_flow and chatwoot_get_captain_metrics without needing to open any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The only routing guidance is the cross-tool hint 'use chatwoot_list_captain_assistants' for obtaining assistant_id, which is genuinely useful. However, there is no explicit statement of when to choose this tool over the many other captain analytics siblings (overview, metrics, resolution_flow, faq_stats); the usage context is implied by 'trend over time' rather than stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
chatwoot_get_contactARead-only
Returns the full details of a specific contact.
Args:
contact_id: Contact ID (use chatwoot_search_contacts to get IDs).
output_format: Output format.
| Name | Required | Description | Default |
|---|---|---|---|
| contact_id | Yes | ||
| output_format | No | markdown |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and destructiveHint=false, so the safety profile is covered. The description adds little behavioral context beyond that, such as no mention of pagination, rate limits, or error behavior; however, it does not contradict the annotations and the simple read-only nature lowers the burden.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is short and front-loaded with the core purpose before the Args block. The only waste is the tautological 'output_format: Output format.', which is a minor flaw in an otherwise lean definition.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple read-only getter, the description plus annotations and output schema are largely sufficient. It includes how to source the required contact_id, and the output schema can cover return details, so nothing critical is missing for a correct call.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It adds real meaning to contact_id by instructing the agent to use chatwoot_search_contacts to get IDs, but output_format is merely called 'Output format', adding no value beyond the schema's enum.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a clear verb and resource: 'Returns the full details of a specific contact.' It also distinguishes itself from the search sibling by telling the agent to use chatwoot_search_contacts for IDs, so there is no ambiguity about which tool fetches a single contact.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly points to chatwoot_search_contacts as the way to obtain contact_id, which is a concrete alternative and usage route. It does not state when not to use this tool versus other contact-related siblings, but the single-contact purpose is clear.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
chatwoot_get_contact_conversationsARead-only
Returns the full conversation history of a specific contact.
Args:
contact_id: Contact ID (use chatwoot_search_contacts to get IDs).
output_format: Output format.
| Name | Required | Description | Default |
|---|---|---|---|
| contact_id | Yes | ||
| output_format | No | markdown |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and destructiveHint=false, and the 'Returns' phrasing is consistent with that safety profile. The description adds the useful scope that the full contact history is returned, but it does not disclose pagination, ordering, size limits, or what happens when a contact has no conversations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The purpose is front-loaded in one clear sentence, and the Args section is compact and scannable. The 'output_format: Output format.' line is tautological and adds no value, which prevents a perfect score.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only tool with two parameters and no output schema, the description covers the core purpose, the required parameter, and the output format options. Missing details such as pagination and exact return structure are not critical given the low complexity, but would push it to a 5.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
contact_id gains actionable meaning through the instruction to use chatwoot_search_contacts to get IDs, which goes beyond the schema's integer type. However, output_format is merely restated as 'Output format,' adding no real semantics beyond the schema's enum and default.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The first sentence uses a specific verb ('Returns') and a specific resource ('full conversation history of a specific contact'), which clearly states what the tool does. This contact-scoped wording distinguishes it from sibling conversation tools like get_conversation_messages or search_conversations.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives an explicit prerequisite by pointing to chatwoot_search_contacts for obtaining contact_id. However, it never explicitly states when to prefer this tool over alternatives such as get_conversation_messages or search_conversations; the choice is only implied by the resource wording.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
chatwoot_get_conversationCRead-only
Returns the full details of a specific conversation.
Args:
conversation_id: Conversation ID.
output_format: Output format.
| Name | Required | Description | Default |
|---|---|---|---|
| output_format | No | markdown | |
| conversation_id | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The annotations already declare readOnlyHint=true and destructiveHint=false, so the safety profile is covered. However, the description adds almost no behavioral context beyond that: it does not mention output_format behavior, whether 'full details' means all fields, any access requirements, or how errors are surfaced. This is minimal additional transparency.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The main sentence is front-loaded and appropriately brief, but the Args block is mostly redundant with the input schema. It is not bloated, yet it is under-specified: the lines add bulk without adding information that helps an agent invoke the tool correctly.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With an output schema present, return-value details are not needed, and the simple two-parameter interface keeps complexity low. However, the description lacks usage context, parameter semantics, and any differentiation from the large sibling list, so an agent would not know when to call this tool versus nearby alternatives.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. The Args section only restates the schema titles ('conversation_id: Conversation ID', 'output_format: Output format') without adding any real meaning. It does not explain what a conversation_id refers to, where to obtain it, or what the markdown/json output formats actually produce.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly identifies the action ('Returns') and the resource ('a specific conversation'), which is enough for an agent to recognize this as a get-by-ID operation. The phrase 'full details' is somewhat vague and does not explicitly distinguish this from siblings like chatwoot_get_conversation_messages, but it is still specific enough to be functional.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is provided on when to choose this tool over its many siblings such as chatwoot_get_conversation_messages, chatwoot_get_conversation_traffic, chatwoot_list_conversations, or chatwoot_search_conversations. The description only states what the tool returns, not the conditions or context in which it should be used.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
chatwoot_get_conversation_messagesCRead-only
Returns the complete message history of a conversation.
Args:
conversation_id: Conversation ID.
output_format: Output format.
| Name | Required | Description | Default |
|---|---|---|---|
| output_format | No | markdown | |
| conversation_id | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and destructiveHint=false, covering the safety profile. The description adds that the result is the complete message history, but it does not disclose pagination, ordering, or output shape; it is minimal but not contradictory.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is appropriately short and front-loads the core behavior. The Args block is redundant with the schema, but it does not meaningfully bloat the definition.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple 2-parameter read tool with annotations, the description is nearly sufficient. However, there is no output schema and no explanation of what markdown vs json output actually changes, leaving ambiguity about the return value.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description needed to compensate. Instead, the Args block merely restates the schema: 'conversation_id: Conversation ID' and 'output_format: Output format' add no semantic value beyond the titles.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Clearly states a specific verb ('Returns') and resource ('complete message history of a conversation'), which differentiates it from metrics/list siblings. However, it does not explicitly call out an alternative tool, so it stops short of a top score.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives no when-to-use guidance or exclusions. It does not indicate when to choose this over chatwoot_get_contact_conversations or other chatwoot_* readers, so the agent must infer usage from the tool name.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
chatwoot_get_conversation_trafficARead-only
Returns a conversation traffic heat map (hour × day of week).
Useful for identifying peak hours and staffing teams appropriately.
Args:
period: Analysis period.
output_format: Output format.
| Name | Required | Description | Default |
|---|---|---|---|
| period | No | 30d | |
| output_format | No | markdown |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The annotations already declare the tool read-only and non-destructive, and the description does not contradict this. It adds the behavioral context that the result is a heatmap, but does not disclose timezone semantics, data granularity, or other operational details. Given the simple read-only nature, annotations carry most of the burden.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and front-loaded with the core purpose. The 'Useful for' sentence adds genuine usage context. The Args block is redundant with the schema but is minimal enough not to significantly bloat the description.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple analytics tool with two optional enum parameters, defaults, an output schema, and read-only annotations, the description covers the core purpose and use-case. It is incomplete mainly in parameter semantics and explicit differentiation from sibling tools, but it remains adequately usable for selection and invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The Args block merely restates the schema titles: 'period: Analysis period' and 'output_format: Output format'. With schema description coverage at 0%, the description was expected to compensate by explaining enum values, units, or format differences, but it adds no meaning beyond what the schema already provides.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'Returns a conversation traffic heat map (hour × day of week)', which clearly identifies the tool's function. This output type is unique among the sibling analytics and conversation tools, so the purpose is unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The sentence 'Useful for identifying peak hours and staffing teams appropriately' gives a clear context for when this tool is relevant. However, it does not mention alternatives or explicitly state when not to use it, stopping short of full routing guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
chatwoot_get_csat_metricsBRead-only
Returns the aggregated CSAT summary: total responses, breakdown by score, and response rate.
Args:
period: Analysis period.
inbox_id: Filter by a specific inbox.
team_id: Filter by a specific team.
output_format: Output format.
| Name | Required | Description | Default |
|---|---|---|---|
| period | No | 30d | |
| team_id | No | ||
| inbox_id | No | ||
| output_format | No | markdown |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover the safety profile with readOnlyHint=true and destructiveHint=false. The description adds useful aggregation context, but it does not disclose how filters combine or any access/rate implications beyond what annotations already imply. This is adequate but not rich.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is short, front-loaded, and uses a clean Args block. Minor redundancy like 'Output format: Output format' prevents a perfect score, but overall it is appropriately sized.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple read-only metrics call with optional parameters, enum defaults, and an output schema, the description is minimally sufficient for invocation. The main gap is lack of usage guidance relative to sibling CSAT tools, but an agent can likely call this tool correctly using schema defaults.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema property descriptions are absent (0% coverage), so the description must compensate. However, it only gives shallow labels such as 'Analysis period' and 'Output format,' which add little beyond the parameter names and enum values; only inbox_id and team_id gain any real meaning as filters.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The opening sentence states a specific action ('Returns') and resource ('aggregated CSAT summary') with concrete components: total responses, breakdown by score, and response rate. This clearly differentiates it from the raw-response sibling chatwoot_get_csat_responses.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives no guidance on when to use this tool versus the many sibling metrics tools, and it does not mention alternatives like chatwoot_get_csat_responses for raw data. The agent must infer selection solely from the tool name and the first sentence.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
chatwoot_get_csat_responsesARead-only
Returns customer satisfaction survey (CSAT) responses.
Enables qualitative analysis of support quality based on customer feedback.
Args:
period: Analysis period.
rating: Filter by score (1 to 5). None = all scores.
inbox_id: Filter by a specific inbox.
team_id: Filter by a specific team.
page: Results page.
output_format: Output format.
| Name | Required | Description | Default |
|---|---|---|---|
| page | No | ||
| period | No | 30d | |
| rating | No | ||
| team_id | No | ||
| inbox_id | No | ||
| output_format | No | markdown |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and destructiveHint=false, so the description does not need to restate safety. It adds no further behavioral context such as pagination semantics, rate limits, or auth requirements, but the output schema and read-only annotation reduce the burden.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description front-loads the purpose in the first sentence and then presents parameters in a compact list. It is appropriately sized for a simple read tool, though the second sentence and several parameter lines add little beyond the first sentence and schema titles.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For an optional-parameter, read-only tool with an output schema, the description covers the product and all six filters at a usable level. It does not explicitly distinguish itself from chatwoot_get_csat_metrics, but the response-vs-metric distinction is inferable and the schema fills in enum/default details.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 0% schema description coverage, the description is the only source of parameter meaning. It adds genuine semantics for rating ('1 to 5', 'None = all scores') and clarifies inbox_id/team_id as filters, but period/page/output_format descriptions remain largely tautological and depend on schema enums/defaults.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Description uses a specific verb ('Returns') with a concrete resource ('customer satisfaction survey (CSAT) responses'), making the core purpose clear. It is implicitly distinct from sibling chatwoot_get_csat_metrics (responses vs metrics) but does not explicitly name that alternative, so it misses the sibling-differentiation bar for a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'Enables qualitative analysis of support quality' implies the tool is for inspecting individual feedback rather than aggregates, giving some use context. However, it never states when to prefer this tool over sibling chatwoot_get_csat_metrics, nor provides any exclusions or alternative conditions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
chatwoot_get_first_response_distributionBRead-only
Returns a histogram of first response time distribution.
Allows identifying which time range most first responses fall into.
Args:
period: Analysis period.
output_format: Output format.
| Name | Required | Description | Default |
|---|---|---|---|
| period | No | 30d | |
| output_format | No | markdown |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already establish readOnlyHint=true and destructiveHint=false, and the description is consistent with a safe read operation. It adds that the output is a histogram, but it does not disclose aggregation behavior, bucketing, or timezone handling. With the safety profile covered by annotations, this is reasonable but not rich.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The definition is compact and front-loads the core behavior in the first sentence. The second sentence adds a small but useful purpose hint, while the redundant Args block is minor noise that does not materially hurt readability.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple read-only analytics tool with two optional enum parameters, defaults, and an output schema, the description covers the essential invocation contract. It does not explain edge cases or exactly what 'first response time' means, but the schema and annotations carry enough of the structural burden that the tool can be selected and called correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, and the description's Args entries—'period: Analysis period' and 'output_format: Output format'—add almost no meaning beyond the parameter names. The schema itself supplies enums and defaults, but the description fails to compensate for the total lack of schema descriptions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'Returns a histogram of first response time distribution,' and adds an analytical intent. It is clear and distinct in function from siblings like get_csat_metrics or get_agent_performance, though it does not name those alternatives.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'Allows identifying which time range most first responses fall into' implies the use case for analyzing response-time concentration, but it never states when to prefer this tool over a sibling or when not to use it. No alternatives or exclusions are given, so the guidance is only implicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
chatwoot_get_inbox_label_matrixARead-only
Returns the conversation distribution between inboxes and labels (inbox × label matrix).
Useful for understanding which channels generate which types of conversations.
Args:
period: Analysis period.
inbox_ids: Filter by specific inboxes (list of IDs; use chatwoot_list_inboxes).
label_ids: Filter by specific labels (list of IDs; use chatwoot_list_labels).
output_format: Output format.
| Name | Required | Description | Default |
|---|---|---|---|
| period | No | 30d | |
| inbox_ids | No | ||
| label_ids | No | ||
| output_format | No | markdown |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and destructiveHint=false, and the description is consistent with a read-only reporting tool. The description adds no further behavioral context such as rate limits, auth requirements, or edge-case behavior, so it does not go beyond the annotation-provided safety profile.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description front-loads the result and use case, then lists all four parameters compactly. Minor redundancy exists in 'period: Analysis period' and 'output_format: Output format', but overall the structure is efficient and scannable.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given that an output schema exists and annotations cover the safety profile, the description is largely complete: it explains the tool's output, its analytical use case, and how to source inbox/label IDs. The only notable gap is not naming sibling tools for comparison, but this is not critical for invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the Args section is essential. It adds real meaning for inbox_ids and label_ids by defining them as filters and pointing to chatwoot_list_inboxes and chatwoot_list_labels. period and output_format are minimally restated, but their allowed values are already provided by schema enums and defaults.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The first sentence states a specific result: 'Returns the conversation distribution between inboxes and labels' and clarifies the shape as an 'inbox × label matrix'. This is distinct from the single-dimension sibling analytics tools like label or inbox performance, so an agent can identify what this tool uniquely provides.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives a clear use case: 'understanding which channels generate which types of conversations.' It doesn't explicitly contrast with sibling analytics tools or state when not to use it, but it provides enough context for an agent to select it for cross-cutting inbox/label analysis.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
chatwoot_get_inbox_performanceBRead-only
Returns performance metrics for a specific inbox.
Args:
inbox_id: Inbox ID (use chatwoot_list_inboxes to get IDs).
period: Analysis period.
output_format: Output format.
| Name | Required | Description | Default |
|---|---|---|---|
| period | No | 30d | |
| inbox_id | Yes | ||
| output_format | No | markdown |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and destructiveHint=false, and the description's 'Returns ... metrics' is consistent with that read-only profile. The description adds the inbox_id acquisition context but no deeper behavioral disclosure (e.g., what metrics are computed, date-window semantics, aggregation). With annotations carrying the safety burden and an output schema present, a 3 is appropriate — no contradiction, but little added value.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The purpose sentence is front-loaded and the Args block is compact. However, two of the three arg explanations ('Analysis period', 'Output format') are tautological restatements of the schema titles, wasting words that could have explained semantics. It is not padded, but the information density is low in the Args section.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Annotations, enums, and the output schema cover the safety profile, valid period/output values, and return structure, so an agent can call it with the required inbox_id. The main gap is that the description never differentiates this from the cluster of sibling performance tools (agent/team/label/inbox) nor explains what 'period' does to the analysis. Adequate for invocation, incomplete for confident tool selection among the family.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description bears the burden of explaining parameters. It helps meaningfully for inbox_id ('use chatwoot_list_inboxes to get IDs') but period ('Analysis period') and output_format ('Output format') are mere restatements of the parameter names. The enums (today/7d/30d; markdown/json) supply structure but not meaning — an agent still doesn't know what '7d' means semantically or how output_format changes the result. Compensation for the 0% coverage is inadequate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb+resource+scope: 'Returns performance metrics for a specific inbox.' This clearly identifies the entity level (inbox) and naturally differentiates it from sibling performance tools (chatwoot_get_agent_performance, chatwoot_get_team_performance, chatwoot_get_label_performance) by the inbox scope. It stops short of naming a sibling explicitly, so it doesn't earn a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage context is implied: an agent can infer this is the tool for inbox-level performance metrics. The description offers one useful pointer — 'use chatwoot_list_inboxes to get IDs' — but this addresses parameter acquisition, not tool selection. There is no explicit guidance on when to choose this over the sibling performance metrics tools or any exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
chatwoot_get_label_performanceARead-only
Returns performance metrics for a specific label.
Allows quantitative analysis of conversations by category/tag.
Args:
label_id: Label ID (use chatwoot_list_labels to get IDs).
period: Analysis period.
output_format: Output format.
| Name | Required | Description | Default |
|---|---|---|---|
| period | No | 30d | |
| label_id | Yes | ||
| output_format | No | markdown |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and destructiveHint=false, and the description's 'Returns' aligns with that. However, the description adds no additional behavioral context such as output format behavior, limitations, or data freshness. No contradiction with annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is short and front-loaded with the core purpose. The second sentence partially restates the first, but overall it is scannable and contains no unnecessary padding.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only tool with one required parameter and enum-constrained optional parameters, the description plus schema is sufficient to make a correct call. It could mention default period/output_format or metric specifics, but output schema and enums cover the remaining details.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It usefully explains label_id by pointing to chatwoot_list_labels, but period is only 'Analysis period' and output_format only 'Output format' – tautological and no better than the property names. It does not describe the meaning of period values or the difference between markdown and json.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: 'Returns performance metrics for a specific label,' and reinforces it with 'quantitative analysis of conversations by category/tag.' This clearly distinguishes it from sibling performance tools that target agents, inboxes, or teams.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage for label-scoped quantitative analysis, but it does not explicitly state when to prefer this over sibling performance tools or provide any exclusions. No alternatives are named, leaving the routing decision mostly to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
chatwoot_get_live_grouped_metricsBRead-only
Returns real-time conversations grouped by team or agent.
Args:
group_by: Group by team (team_id) or agent (assignee_id).
team_id: Filter by a specific team (optional).
output_format: Output format.
| Name | Required | Description | Default |
|---|---|---|---|
| team_id | No | ||
| group_by | No | team_id | |
| output_format | No | markdown |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and destructiveHint=false, so the safety profile is covered. The description adds only 'real-time' and grouping scope; it does not describe time windows, data volume, or parameter interactions. There is no contradiction with the annotations, but also little behavioral context beyond what the name and schema already imply.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The purpose is front-loaded and the three Args lines are compact. The output_format line is near-tautological and could be replaced with default/choice information, but overall the definition avoids unnecessary prose and stays focused.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the output schema and read-only annotations, this is a serviceable skeleton for a read-only metrics tool. Gaps remain: no usage guidance, no parameter relationships, and no clarification of what 'metrics' or 'real-time' means in practice. It is adequate for a simple call, but not fully complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 0% schema description coverage, the description must carry parameter meaning. It does usefully map group_by values to team_id/assignee_id and explains team_id as an optional filter. However, output_format is only described as 'Output format,' which adds no meaning beyond the schema enum, and the interaction between team_id and group_by=assignee_id is left unspecified.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The first sentence identifies the action (returns), the object (real-time conversations), and the distinguishing grouping dimensions (team or agent). It clearly separates this from the ungrouped get_live_metrics sibling, though it does not explicitly name that alternative. This is clear and specific, but slightly short of a best-practice differentiation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use this tool instead of chatwoot_get_live_metrics, chatwoot_get_summary, or other metric-focused siblings. It does not mention exclusions, alternatives, or preferred contexts. The agent must infer applicability from the tool name and the first sentence alone.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
chatwoot_get_live_metricsBRead-only
Returns real-time conversation counts: open, unattended, unassigned, and pending.
Args:
output_format: Output format.
| Name | Required | Description | Default |
|---|---|---|---|
| output_format | No | markdown |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and destructiveHint=false, and the description's 'Returns' wording aligns with that. The description adds 'real-time' as a behavioral trait but does not disclose any other operational nuances such as refresh behavior or limitations. No contradiction with annotations exists.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The main sentence is concise and front-loaded with the tool's purpose. The Args block is slightly redundant because it repeats the parameter title without adding useful detail, but overall the description is compact and readable.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple read-only tool with one optional parameter, an output schema, and strong annotations, this description is mostly sufficient. It lists the exact metric categories and notes real-time freshness, though it could be more complete by clarifying how this differs from grouped or historical metrics.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description needed to add meaning for the output_format parameter. It only repeats 'output_format: Output format,' which adds nothing beyond the schema's existing title, enum, and default. The schema provides the real value here, not the description.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description starts with a specific verb ('Returns') and identifies a concrete resource: real-time conversation counts for open, unattended, unassigned, and pending conversations. It is clear and not a tautology, though it doesn't explicitly distinguish itself from closely related siblings like chatwoot_get_live_grouped_metrics.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The term 'real-time' implies the tool is appropriate when current conversation counts are needed, which is a mild usage cue. However, there is no explicit guidance about when to choose this tool over alternatives, especially the similarly named live metrics sibling.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
chatwoot_get_summaryARead-only
Returns the account summary: conversations, response times, resolution rate, and CSAT.
Args:
period: Analysis period (today, yesterday, 7d, 15d, 30d, 90d).
output_format: Output format — markdown (default) or json.
| Name | Required | Description | Default |
|---|---|---|---|
| period | No | 30d | |
| output_format | No | markdown |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and destructiveHint=false, so the safety profile is covered. The description adds useful context about the output formats and default, and names the metrics returned, but does not explain aggregation behavior, time boundaries, or response structure beyond the output schema.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and front-loaded: a one-sentence purpose followed by concise parameter documentation. Every sentence adds value, and there is no redundancy with the schema beyond what is needed for clarity.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple read-only summary tool with an output schema, the description provides the key behavioral context: the metric categories, period options, and output format. It could be slightly more complete with explicit guidance on when an account summary is appropriate versus per-agent or per-inbox reports, but that is a minor gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description carries the full burden for parameter meaning. It documents both period and output_format with their allowed values and defaults, adding semantic labels that the schema itself lacks.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states a specific verb and resource: 'Returns the account summary,' and lists the included metrics (conversations, response times, resolution rate, CSAT). This distinguishes it from more specialized siblings like bot summaries or agent performance reports, though it does not name alternatives explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is given on when to prefer this tool over the many sibling analytics tools. The description implies it is for high-level account reporting, but it never states conditions, exclusions, or alternative tools.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
chatwoot_get_team_performanceBRead-only
Returns performance metrics for a specific team.
Args:
team_id: Team ID (use chatwoot_list_teams to get IDs).
period: Analysis period.
output_format: Output format.
| Name | Required | Description | Default |
|---|---|---|---|
| period | No | 30d | |
| team_id | Yes | ||
| output_format | No | markdown |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and destructiveHint=false, so the safe read-only nature is covered. The description adds minimal behavioral context beyond that; it does not mention rate limits, auth requirements, error behavior, or what the returned metrics represent.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is short and front-loaded with the core action, followed by a compact Args block. It contains no real filler, though two of the argument descriptions are low-value restatements.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple read-only metrics tool with an output schema and safety annotations, the description is minimally adequate. It is incomplete on the semantics of period and output_format and does not help the agent choose between the many performance/metric sibling tools.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 0% schema description coverage, the description needed to compensate. team_id gets useful guidance via chatwoot_list_teams, but period and output_format are only restated as 'Analysis period' and 'Output format', adding no meaning beyond the property names and enum values already visible in the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a clear verb and resource: 'Returns performance metrics for a specific team.' It is unambiguous about the team scope, which separates it from agent, label, and inbox performance tools, though it does not explicitly name or contrast alternative siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives actionable guidance: use chatwoot_list_teams to obtain the required team_id. It does not state when to prefer this over get_agent_performance or other performance variants, but the context for invoking it is clear.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
chatwoot_list_agentsARead-only
Lists all agents in the account with ID, email, and availability status.
Args:
output_format: Output format.
| Name | Required | Description | Default |
|---|---|---|---|
| output_format | No | markdown |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already establish this as a safe read-only operation. The description adds the account scope and returned fields, but omits details such as ordering, pagination, or the meaning/range of availability status. With annotations covering the safety profile, this is adequate but not rich.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is very short and front-loaded, which is appropriate for a simple list tool. The 'Args: output_format: Output format' line is largely redundant with the input schema, preventing a perfect score, but it does not add meaningful bloat.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple read-only listing tool with one optional parameter and an output schema, the description is largely complete: it names the resource, scope, and prominent fields. It could mention account context or output format effects more explicitly, but those gaps are minor given the annotations and schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description needed to compensate, but it only repeats 'output_format: Output format,' which adds no real meaning beyond the schema's enum and default. It does not explain the difference between markdown and json output or how the format affects the response.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states a specific action and resource: 'Lists all agents in the account' and names the returned fields (ID, email, availability status). The sibling tool list confirms there is no competing agent-listing tool, so the purpose is unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies when to use this tool—whenever agent details are needed—but provides no explicit guidance, exclusions, or alternatives. Because no sibling tool lists agents, the intended use is fairly obvious, but the guidance is only implicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
chatwoot_list_captain_assistantsARead-only
Lists all Captain (AI) assistants configured in the account.
Args:
output_format: Output format.
| Name | Required | Description | Default |
|---|---|---|---|
| output_format | No | markdown |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The annotations already declare readOnlyHint=true and destructiveHint=false, and the description adds account-scoped context and 'all' helpers. It does not disclose additional behavioral details such as ordering, pagination, or permissions, but for a simple read-only list this is acceptable.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The first sentence is clear and front-loaded with the action and resource. The 'Args' section is redundant with the schema and adds slight noise, but the overall definition remains compact.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool is a simple read-only list operation with one optional parameter, and the schema's enum/default plus annotations cover most invocation needs. The only remaining gap is the lack of meaningful paramter description, which is already reflected in the parameter_semantics score.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The description's paramter entry 'output_format: Output format.' only restates the parameter name and adds no new meaning. With schema description coverage at 0%, the description fails to compensate; the enum and default in the input schema are the only real guidance for valid values.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific action ('Lists'), a specific resource ('Captain (AI) assistants'), and scope ('configured in the account'). This clearly distinguishes it from sibling tools like chatwoot_list_agents or chatwoot_list_inboxes.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies the tool should be used when an agent needs to enumerate Captain assistants, but it does not explicitly discuss when to prefer it over related listing tools or mention exclusions. Usage intent is inferred from the resource name rather than stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
chatwoot_list_conversationsBRead-only
Lists conversations with filters by status, assignee, inbox, team, and label.
Args:
status: Conversation status (open, resolved, pending, snoozed, all).
assignee_type: Filter by assignment type (me, assigned, unassigned, all).
inbox_id: Inbox ID to filter by (use chatwoot_list_inboxes).
team_id: Team ID to filter by (use chatwoot_list_teams).
label: Label to filter by (use chatwoot_list_labels to see available labels).
page: Results page (25 per page).
output_format: Output format.
| Name | Required | Description | Default |
|---|---|---|---|
| page | No | ||
| label | No | ||
| status | No | open | |
| team_id | No | ||
| inbox_id | No | ||
| assignee_type | No | ||
| output_format | No | markdown |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and destructiveHint=false, so the safe-read nature is covered. The description adds modest behavioral context, such as pagination ('25 per page') and the available filter dimensions, but does not disclose response ordering, defaults, or edge cases beyond what the schema and output schema already imply.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is appropriately sized: a one-line purpose followed by a structured Args list. It is front-loaded and each argument line earns its place by adding enum values or helper-tool references. Some redundancy with the schema exists, but it is not excessive.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With an output schema present and annotations covering the read-only safety profile, the description provides enough operational detail for most listing calls. It explains filters, pagination, and helper tools. The main gap is the lack of explicit routing against chatwoot_search_conversations, which could cause an agent to pick the wrong tool for text-based search.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description carries the burden of explaining parameters. It compensates well by listing all status and assignee_type enum values, explaining page size, and pointing to sibling tools for resolving inbox_id, team_id, and label. The only weak spot is output_format, which is described only as 'Output format' without adding meaning beyond the schema enum.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a clear verb and resource: 'Lists conversations with filters by status, assignee, inbox, team, and label.' This is specific and immediately conveys what the tool does. However, it does not explicitly differentiate itself from the sibling chatwoot_search_conversations, so it stops short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to use this tool versus chatwoot_search_conversations or chatwoot_get_conversation. The references to chatwoot_list_inboxes, chatwoot_list_teams, and chatwoot_list_labels help resolve parameter values, but they do not explain when this listing tool is preferable to alternatives.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
chatwoot_list_inboxesBRead-only
Lists all inboxes (channels) in the account: WhatsApp, email, widget, etc.
Args:
output_format: Output format.
| Name | Required | Description | Default |
|---|---|---|---|
| output_format | No | markdown |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and destructiveHint=false, so the safety profile is covered. The description adds useful scoping context by stating that it returns all inboxes in the account, but it does not describe output format behavior or any other operational detail beyond the annotation baseline.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The main sentence is short and front-loaded, clearly stating the core behavior. The Args line is mostly redundant, but the overall description is small and free of unnecessary elaboration.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple read-only list operation with one optional enumerated parameter and an output schema, the description is nearly sufficient. It covers the core function and scope, though it could be improved by noting that only the listed inboxes are returned and clarifying the output_format option in a non-tautological way.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description should compensate, but it only repeats 'Output format' without explaining the meaning of the available values or when to choose one. The enum and default in the schema do most of the work, and the description adds no semantic value beyond a tautology.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the verb ('Lists') and resource ('all inboxes (channels) in the account'), with concrete examples of channel types. It is easy for an agent to understand what the tool does, though it does not explicitly distinguish it from similar sibling tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives no guidance on when to use this tool versus alternatives, and no exclusions or prerequisites are mentioned. An agent is left to infer that this is the right tool whenever a full list of inboxes is needed.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
chatwoot_list_labelsARead-only
Lists all labels in the account with ID, color, and description.
Labels are used to tag and categorize conversations.
Args:
output_format: Output format.
| Name | Required | Description | Default |
|---|---|---|---|
| output_format | No | markdown |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and destructiveHint=false, so the safe-read profile is covered. The description adds useful scope context ('all labels in the account') and return fields, but does not disclose behaviors like pagination, limits, or authentication requirements.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is short and front-loaded with the core purpose. The label-purpose sentence is a small useful addition, though the 'Args: output_format: Output format.' line duplicates the schema without earning its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple list tool with one optional enum parameter, an output schema, and safety annotations, the description is reasonably complete: it gives the resource scope and output fields. It could be improved by naming the sibling tool for label performance metrics, but nothing critical is missing for a correct call.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The only parameter is output_format, but the description merely repeats 'Output format' without adding meaning beyond the schema's enum. With 0% schema description coverage, the description was expected to compensate but instead stays tautological.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: 'Lists all labels in the account' and even names the result fields (ID, color, description). This clearly distinguishes it from sibling tools like chatwoot_list_agents or chatwoot_get_label_performance.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The line 'Labels are used to tag and categorize conversations' gives some contextual value, implying use for understanding label purposes. However, it does not explicitly say when to choose this tool over alternatives, nor provide any exclusions or when-not-to-use guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
chatwoot_list_teamsCRead-only
Lists all teams in the account.
Args:
output_format: Output format.
| Name | Required | Description | Default |
|---|---|---|---|
| output_format | No | markdown |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already provide readOnlyHint=true and destructiveHint=false, so the safety profile is clear and not contradicted. The description adds the account-level scope but does not disclose other behavioral traits such as pagination, ordering, limits, or how the output_format parameter changes the result.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is short and front-loaded, but the 'Args: output_format: Output format.' line is redundant with the input schema. It is concise but not efficiently informative, and the space could have been used to clarify output format behavior or usage context.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple read-only list tool with one optional schema-defined parameter and an output schema, the core operation is adequately covered. However, missing usage guidance and the lack of meaningful parameter semantics leave notable gaps for an agent deciding whether and how to invoke it.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The only parameter is described as 'Output format,' which adds no meaning beyond the schema's enum and default values. Schema description coverage is 0%, and the description fails to compensate by explaining what each output format yields or why an agent would choose one over the other.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a clear operation: 'Lists all teams in the account,' identifying both the resource (teams) and the account scope. It does not explicitly distinguish itself from sibling tools like chatwoot_get_team_performance, but the verb 'lists all' makes its purpose unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance about when to use this tool versus alternatives, no mention of related read-only tools, and no exclusionary context. The description simply states what it does without helping an agent decide when it is the right choice.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
chatwoot_search_contactsARead-only
Searches contacts by name, email, phone, or identifier.
Args:
query: Text to search (name, email, phone, or identifier).
page: Results page.
output_format: Output format.
| Name | Required | Description | Default |
|---|---|---|---|
| page | No | ||
| query | Yes | ||
| output_format | No | markdown |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and destructiveHint=false, so safety is covered. The description adds useful context by listing which contact fields can be searched, but it does not disclose matching behavior (e.g., partial vs. exact, case sensitivity) or pagination behavior beyond a 'page' parameter.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and front-loaded with the main purpose, followed by a brief argument list. The only minor issue is slight redundancy: the first sentence repeats what the 'query' argument line says.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple search tool with a small parameter set, output schema present, and read-only annotations, the description gives enough information to understand what the tool searches and what the key argument means. It lacks alternative-routing guidance and behavioral specifics, but nothing critical is missing for basic invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It does well for 'query' by explaining it matches name, email, phone, or identifier. However, 'page' is only described as 'Results page' and 'output_format' is simply 'Output format', adding little beyond the schema titles and enum.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'Searches contacts'. It also enumerates the searchable attributes (name, email, phone, or identifier), which makes the tool's scope clear and distinguishes it from sibling tools like chatwoot_get_contact or chatwoot_search_conversations.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The use case is implied by the name and description: use this tool to search contacts. However, there is no explicit guidance about when to prefer it over chatwoot_get_contact or other contact-related siblings, and no mention of alternatives or exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
chatwoot_search_conversationsARead-only
Searches conversations and messages by text.
Args:
query: Text to search across conversations and messages.
page: Results page.
output_format: Output format.
| Name | Required | Description | Default |
|---|---|---|---|
| page | No | ||
| query | Yes | ||
| output_format | No | markdown |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and destructiveHint=false, so the safety profile is covered and the description need not restate it. However, the description adds no further behavioral context — no pagination limits, match semantics, or result-scope caveats — and the 'by text' clause is really purpose rather than behavior. It does not contradict the annotations, which keeps it at a viable level.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The one-sentence purpose is front-loaded and contains no filler. The Args block is compact, though 'page: Results page' and 'output_format: Output format' are near-tautological with the schema titles and could be dropped without losing information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists and the annotations cover the safety profile, so return values and non-destructiveness are already handled elsewhere. What is missing is guidance on result behavior (pagination size, match semantics) and how this search relates to adjacent siblings. The core call is fully specified for a 3-param tool, but the definition is thin around the edges.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must carry the parameter-meaning burden. The query arg receives genuine semantics ('Text to search across conversations and messages'), but page ('Results page') and output_format ('Output format') merely echo the schema titles, with the enum providing the only real format information. This is partial compensation at best.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The opening sentence, 'Searches conversations and messages by text,' pairs a specific verb (searches) with a clear resource scope (conversations and messages) and a mechanism (by text). This scope naturally differentiates it from chatwoot_search_contacts and from the many get/list siblings, whose object or operation differs.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The verb 'searches' and the explicit text-scope imply the tool is for finding matching conversational content, but the description never states when to prefer it over chatwoot_get_conversation_messages, chatwoot_list_conversations, or chatwoot_search_contacts. No exclusions, prerequisites, or alternatives are named, so selection is left to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
30 tool updates
v0.1.0- First observed
chatwoot_get_agent_performance - First observed
chatwoot_get_bot_summary - First observed
chatwoot_get_captain_faq_stats - First observed
chatwoot_get_captain_metrics - First observed
chatwoot_get_captain_overview - First observed
chatwoot_get_captain_resolution_flow - First observed
chatwoot_get_captain_resolution_trend - First observed
chatwoot_get_contact - First observed
chatwoot_get_contact_conversations - First observed
chatwoot_get_conversation - First observed
chatwoot_get_conversation_messages - First observed
chatwoot_get_conversation_traffic - First observed
chatwoot_get_csat_metrics - First observed
chatwoot_get_csat_responses - First observed
chatwoot_get_first_response_distribution - First observed
chatwoot_get_inbox_label_matrix - First observed
chatwoot_get_inbox_performance - First observed
chatwoot_get_label_performance - First observed
chatwoot_get_live_grouped_metrics - First observed
chatwoot_get_live_metrics - First observed
chatwoot_get_summary - First observed
chatwoot_get_team_performance - First observed
chatwoot_list_agents - First observed
chatwoot_list_captain_assistants - First observed
chatwoot_list_conversations - First observed
chatwoot_list_inboxes - First observed
chatwoot_list_labels - First observed
chatwoot_list_teams - First observed
chatwoot_search_contacts - First observed
chatwoot_search_conversations
TDQS
Scored across 30 tools
Each tool has a clearly distinct purpose: list vs. get vs. search, and per-entity performance metrics are separated by resource (agent, inbox, team, label, captain). Even the Captain analytics tools (FAQ stats, overview, resolution flow, resolution trend) are differentiated by their focus, reducing selection ambiguity.
All tools follow a strict chatwoot_<verb>_<noun> pattern with consistent verbs (get, list, search). The naming is uniform across all 30 tools, making the API predictable and easy to navigate.
At 30 tools, the server is on the heavier side, but the breadth of Chatwoot's feature set (agents, inboxes, teams, labels, contacts, conversations, CSAT, Captain AI, live metrics, analytics) justifies the count. Each tool covers a distinct aspect, so the number feels appropriate for a comprehensive analytics-focused MCP.
The server is strong for read-only analytics and reporting, covering the main entities and metrics. Minor gaps exist: no single-entity detail tools beyond list (e.g., get_inbox, get_team), and no write operations, but that aligns with its apparent reporting purpose. The core lifecycle of querying and analyzing data is well covered.
Maintenance
Related MCP Connectors
Query your Betterlytics web analytics from AI agents: traffic, funnels, journeys, errors, uptime.
- AgentioOAuthcom.agentio
Ask about your Agentio campaigns, Creator deals, and performance on YouTube and Meta. Read-only.
AI access to Hitsteps analytics, live visitors, uptime, goals, alerts, and chats.
- RulebaseOAuthco.rulebase
CX ops: read conversations, calls and QA evaluations from Zendesk, Freshdesk, Five9 and more.
Related MCP Servers
- AlicenseAqualityCmaintenanceProvides AI assistants with read access to Clamp analytics data including pageviews, visitors, referrers, and custom events. Enables traffic analysis, conversion funnel evaluation, and metric alerts through natural language queries.3327 npm2MIT
- AlicenseNot gradedqualityFmaintenanceProvides AI assistants with read-only access to Secureframe's compliance data, enabling querying of security controls, tests, users, vendors, and more across frameworks like SOC 2 and ISO 27001.8MIT
- FlicenseNot gradedqualityBmaintenanceEnables AI assistants to query internal business data for insights into customers, revenue, subscriptions, sales, and churn through controlled, read-only MCP tools.-
- FlicenseNot gradedqualityBmaintenanceEnables AI assistants to look up and update WorkForceAI platform data, including voice agents, campaigns, phone numbers, call logs, and dashboard stats, through natural language.-