callsense-mcp
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@callsense-mcpscore call_007 against the latest playbook rules"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
callsense-mcp
An MCP server that turns Claude (Claude Desktop, Claude Code, or an agent built on the Agent SDK) into a sales-call analyst for a team's own playbook. Claude reads the call and scores it. The server decides what gets saved: a score is accepted only with a quote copied from the call, said by the right side of the conversation.
It is built on call-scoring-harness, the evaluation pattern behind a call-intelligence pipeline that has been running for a manufacturer's sales team since June 2026. The rules, the synthetic calls and the quote check come from there.
Why the evidence guard is the product
The first production version was killed in its first week over a single score the reps saw as unfair. A score the reps cannot trace to a sentence in their own call gets argued with and then ignored. The pipeline came back when every score had its source quote next to it.
A model in a chat can write a fluent score whether or not the call supports it, so the server never takes the model's word:
A quote or no score. A true score needs text copied word for word from one numbered turn. Spaces and case are ignored, nothing else is.
The right speaker. Churn risk has to be the customer's words. Handling an objection has to be the rep's. Each rule names its speaker, and the server checks who said the turn.
All or nothing. Every rule of the version has to be scored, and one bad score blocks the set. Nothing half-checked reaches the store or the deal card.
Rejections the model can act on. A rejected score comes back with a reason and a hint: the turn where the quote really is, or the closest sentence in the call. Claude fixes it and submits again.
Verdicts are code. At risk, follow up or healthy is computed from the accepted scores by a fixed table, so the same scores always give the same verdict.
Related MCP server: Prompt Lab MCP Server
A rejection
Real output from tests/test_server.py. The submission cleaned a filler out of the customer's words and counted the customer's own remark as the rep's qualification question:
{
"accepted": false,
"saved": false,
"call_id": "call_007",
"rules_version": 2,
"rejected": [
{
"rule_id": "qualification_asked",
"reason": "turn 6 is the customer speaking; qualification_asked needs the rep's own words",
"hint": "rep turns in this call: 1, 3, 5, 7"
},
{
"rule_id": "churn_risk",
"reason": "quote not found anywhere in the call",
"hint": "closest sentence is in turn 2 (customer): \"So we are, uh, thinking about moving the October volume to them.\" Copy it exactly, or score churn_risk false."
}
],
"missing": [],
"next": "Nothing was saved. Fix the rejected rules, add any missing ones, and submit the full set again."
}A quote placed in the wrong turn gets "hint": "found in turn 6".
Tools
Tool | What it does |
| Calls with date, rep, customer, number of turns, whether scored, verdict. |
| Numbered turns |
| A rule set: id, question, evidence, speaker, and whether quotes are required. Latest by default. |
| The evidence guard. Saves the set only if every score passes, otherwise returns |
| Scoring on the server for runs nobody watches: |
| Rate per rule, a table per rep, calls at risk with their quotes, follow-ups, unscored calls. Every count lists the call ids behind it. |
| Agreement with the hand-graded labels and quote validity, pass or fail. |
| The deal-card note and CRM field payload in the shape of the harness write-back. Builds text only, sends nothing. |
Resources: callsense://rules/{version} (the YAML file) and callsense://calls/{call_id}.
Prompts: coach_rep(rep) writes a weekly coaching note from the rep's scored calls, with their quotes. score_unscored() scores every unscored call through the guard and ends with the team report.
Verdict table
callsense_mcp/verdict.py. If an at-risk row fires, the call is at risk. If not and the follow-up row fires, it is follow up. Otherwise it is healthy.
Verdict | When | Reason shown, with the customer's quote |
at risk | churn_risk yes and objection_handled no | customer signalled leaving and the objection was not answered |
at risk | churn_risk yes and next_step_committed no | customer signalled leaving and no next step was agreed |
follow up | next_step_committed no | no next step agreed |
healthy | none of the above |
Flags are shown whatever the verdict: churn signal (with its quote), no next step, not qualified.
Connect
Python 3.10 or newer, MCP Python SDK 2.x.
git clone https://github.com/dsichz/callsense-mcp && cd callsense-mcp
python3 -m venv .venv
.venv/bin/pip install -e ".[dev]" # ".[dev,claude]" to enable the server-side Claude scorerClaude Code:
claude mcp add callsense -- "$PWD/.venv/bin/python" -m callsense_mcpThen ask it to score the unscored calls, or run the prompt as /mcp__callsense__score_unscored.
Claude Desktop, in ~/Library/Application Support/Claude/claude_desktop_config.json on macOS:
{
"mcpServers": {
"callsense": {
"command": "/absolute/path/to/callsense-mcp/.venv/bin/python",
"args": ["-m", "callsense_mcp"]
}
}
}The server reads rules/ and data/ from the repository. To point it at another folder with the same layout, set CALLSENSE_HOME (-e CALLSENSE_HOME=... for claude mcp add, an env block in the Desktop config).
Tests and evals
.venv/bin/pytest -q # 48 tests, no network, no API key
.venv/bin/python -m callsense_mcp eval --rules-version 2 # exits 1 below the gate
.venv/bin/python -m callsense_mcp eval --rules-file draft.yamlThe tests cover the guard (exact quote, paraphrase, quote from another turn, wrong speaker, true score without a quote, missing rules, unknown rules version), the server end to end through an in-memory MCP client and over stdio, the eval gate, and the Claude scorer's retry loop against a fake client. The paid API is never called.
The gate applies to every rule and to the total. A draft that moves next_step_committed to the rep's words keeps the total at 88% and still fails, because that one rule drops to 50%. There is a test for exactly that. A rule with no human labels cannot pass at all.
The baseline scores 100% on the six graded calls. As in the harness, that says something about the set, not about keyword scoring.
A live run
examples/claude_code_session.md is one recorded Claude Code session (model claude-opus-5-5) that scored the two unscored calls and summarised the team report. The guard had nothing to reject in that run. The log says so, and notes what the run did show: Claude's chat summary shortened a quote the guard had kept whole, and Claude guessed a verdict that the verdict table contradicts.
Data
All calls are synthetic. call_001 to call_006 and the labels in data/graded/ are copied from call-scoring-harness: B2B calls shaped after a Russian-speaking manufacturer's, written in English and labelled by hand. call_007 and call_008 were written for this repository and are not in the graded set. Rep and company names are made up. No client data is in this repository.
The shipped store has the six graded calls scored from their human labels (python -m callsense_mcp seed rebuilds that) and call_007 and call_008 unscored.
Layout
Path | What |
| The evidence guard. |
| Verdict and flag table. |
| MCP tools, resources and prompts. |
| The logic behind them. |
| Baseline and Claude scorers. |
| The eval gate. |
| Team report, deal-card note. |
| v1 and v2 from the harness, plus |
| Calls with metadata, hand labels, accepted scores. |
| The recorded Claude Code session. |
Limits
The guard proves that a quote is real and comes from the right side. It does not prove the quote supports the rule; that is what
run_evalmeasures against human labels. On call_008 the baseline counts "Good. The price stays as in the September quote." as objection handling, a keyword hit with no objection in the call, and the guard lets it through because the rep did say it.The graded set is six short calls. The numbers here show the checks work, not how a model scores a real sales floor.
The store is JSON files for one team, with no auth and no multi-user locking.
Artem Baranov · theaigency.space · linkedin.com/in/artem-es
Available Tools
8 toolsdeal_card_noteARead-onlyIdempotent
The note the CRM write-back would put on the deal card: verdict, flags, every score with its quote, and the deal-field payload. Builds the text only; nothing is sent.
| Name | Required | Description | Default |
|---|---|---|---|
| call_id | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, and destructiveHint=false, but the description adds important behavioral context: despite the 'write-back' framing, it clarifies that only text is built and nothing is sent. It also describes the exact note contents, helping the agent understand the output beyond the safety profile.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and front-loads the note's contents, with a short second sentence clarifying the no-send behavior. The first sentence is slightly indirect ('The note the CRM write-back would put'), but overall there is no wasted text.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the rich annotations and the existence of an output schema, the description is largely complete: it explains what is generated, what it contains, and that no side effect occurs. The only notable gap is the absence of any parameter explanation for call_id, though the parameter name makes the intent obvious.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema has one required parameter (call_id) with 0% description coverage, and the description does not mention call_id or explain its meaning. While the parameter name is self-explanatory, the description adds no semantic value beyond what the schema already provides.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific action ('Builds the text only') and resource (the deal card note), and enumerates the note contents (verdict, flags, every score with quote, deal-field payload). It also implicitly distinguishes itself from the actual CRM write-back by clarifying that nothing is sent.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description does not state when to use this tool versus alternatives such as submit_scores or score_call. It only implies preview usage through 'Builds the text only; nothing is sent,' but provides no explicit when-to-use guidance or alternatives.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_callBRead-onlyIdempotent
One call with numbered turns {n, speaker, text}. Quote from these turns when scoring.
| Name | Required | Description | Default |
|---|---|---|---|
| call_id | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, and destructiveHint=false, so the safety profile is covered. The description usefully adds the shape of the returned turns (n, speaker, text), but says nothing about auth, rate limits, or transcript completeness beyond what the output schema carries.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short, front-loaded clauses with no filler; the return shape and the usage hint both earn their place. It is terse rather than padded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With an output schema present, return-value explanation is not strictly required, and the annotations cover safety, so the core need is met. Still, a retrieval tool with a 0%-documented key parameter and no routing guidance against list_calls or score_call leaves gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% for the single required call_id parameter, so the schema supplies no semantics. The description never mentions the identifier or its accepted form (ID vs name), leaving the agent to infer that the call's ID is the key, which is plausible but unconfirmed.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names the resource (a single call) and its contents (numbered turns with n, speaker, text), which implicitly distinguishes it from the sibling list_calls. The retrieval verb itself is only implied by "get_call" and "One call," so it stops short of a full verb+resource statement.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
"Quote from these turns when scoring" gives an implied workflow context, linking this tool to scoring tasks. However, it never states when to use this versus list_calls or score_call, and offers no prerequisites or exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_rulesBRead-onlyIdempotent
Scoring rules of a version (latest by default): id, question, what counts as evidence, and speaker, the side whose own words the evidence has to be. Also says whether quotes are required.
| Name | Required | Description | Default |
|---|---|---|---|
| version | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The annotations already declare that this is a read-only, idempotent, non-destructive, closed-world operation, covering the safety profile. The description adds the default-version behavior and describes what the rules contain, but it does not add deeper behavioral traits such as error handling, pagination, or authentication requirements.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and front-loads the main purpose before listing the rule fields. It is efficient, though the field list could be slightly more structured for scanning.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple read-only tool with rich annotations and an output schema, the description covers the core purpose, the default version behavior, and the key returned content. It is largely complete, with only minor gaps around version input semantics and explicit usage context.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It explains that the version parameter defaults to the latest version, which adds useful meaning, but it does not clarify the expected integer format or what happens when an invalid version is supplied.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states that the tool retrieves scoring rules for a version, and it enumerates the key fields returned (id, question, evidence, speaker, quote requirement). It does not explicitly differentiate from sibling tools, but none of the listed siblings overlap with this resource.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It mentions that the latest version is used by default, which gives some context for the optional version parameter, but it does not say when to use this tool versus alternatives or under what conditions a specific version should be requested. There is no explicit when/when-not guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_callsARead-onlyIdempotent
List calls with date, rep, customer, number of turns, whether the call is scored, and its verdict. Filter by rep name and by scored=true/false.
| Name | Required | Description | Default |
|---|---|---|---|
| rep | No | ||
| scored | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true and destructiveHint=false, so safety is fully covered. The description adds the useful fact that results carry scoring/verdict fields, but says nothing about ordering, pagination, or limits on a list operation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two compact sentences: the first declares what is returned, the second declares the filters. No filler and well front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be explained, yet the description lists them anyway and covers both filters despite 0% schema coverage. The only remaining gap is pagination/ordering behavior for a list tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0% – the parameters are bare 'Rep' and 'Scored' with no descriptions. The description compensates by clarifying that 'rep' is a name and 'scored' is a true/false selector, adding real meaning beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('List calls') and enumerates the returned fields (date, rep, customer, turns, scored, verdict), so the agent knows exactly what it gets. It does not explicitly distinguish itself from the sibling get_call, so it stops short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The second sentence tells the agent which filters exist ('Filter by rep name and by scored=true/false'), implying when to reach for this tool versus an unfiltered pull. There is no explicit when-not guidance and no mention of the get_call sibling, so usage is only implied.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
run_evalARead-onlyIdempotent
Measure a rule set against the hand-graded calls before it ships: agreement with human labels per rule, quote validity, and pass or fail. Every rule and the total have to reach min_agreement. Pass rules_yaml to test a draft rule set without saving it. A rule with no human labels cannot pass.
| Name | Required | Description | Default |
|---|---|---|---|
| scorer | No | baseline | |
| rules_yaml | No | ||
| min_agreement | No | ||
| rules_version | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already establish the read-only, idempotent, non-destructive profile, so the bar is lower; the description adds real behavioral rules beyond that: every rule and the total must reach min_agreement, and a rule without human labels cannot pass. It does not explain scoring cost, latency, or what a failing result looks like, so a 4 rather than 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three dense clauses with no filler, front-loaded on what is measured before moving to the draft override and the pass constraint. It is slightly packed, but every sentence adds a distinct constraint.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With an output schema present, the description needn't describe return values, and the annotations cover safety. It covers the core evaluation semantics and the draft path; the only real gap is the meaning of scorer and rules_version for a 4-parameter tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must carry the load, but it only covers two of four parameters: rules_yaml (draft, unsaved evaluation) and min_agreement (the pass threshold). The scorer enum values and rules_version are given no meaning at all, leaving half the surface unexplained.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description gives a specific verb and resource ('measure a rule set against the hand-graded calls') and enumerates the outputs (per-rule agreement, quote validity, pass/fail), which clearly separates it from read-only siblings like get_rules. It does not explicitly name an alternative tool, so it falls just short of the 5 bar.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It states the timing context ('before it ships') and a distinct mode of use ('Pass rules_yaml to test a draft rule set without saving it'), which is genuine when-to-use guidance. It stops short of naming alternatives or stating when NOT to call it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
score_callA
Score a call on the server, for runs with nobody in the loop. 'baseline' is a keyword floor that needs no key. 'claude' calls the Anthropic API and needs ANTHROPIC_API_KEY in the server's environment. The result goes through the same evidence guard as submit_scores and is saved only if it passes. In a chat, prefer get_call plus submit_scores.
| Name | Required | Description | Default |
|---|---|---|---|
| scorer | No | baseline | |
| call_id | Yes | ||
| rules_version | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Adds real behavior beyond the annotations: 'claude' requires ANTHROPIC_API_KEY in the server environment, and the result goes through the same evidence guard as submit_scores and 'is saved only if it passes' — so a call may silently not persist. It does not address the idempotentHint=false implication (re-scoring the same call), which keeps it from a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four short sentences, front-loaded with the action and scope, then modes, then side effects, then routing. Dense and waste-free, though the quoted-keyword phrasing is slightly packed.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need no explanation, and the description covers auth, persistence guard, and routing. The only gap is the unexplained rules_version parameter, a minor omission.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must carry the load. It explains the two scorer enum values and their differing auth/behavior requirements, which is the most important parameter, but says nothing about call_id or rules_version.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Score a call on the server') and immediately scopes it to headless/'nobody in the loop' runs, which separates it from the interactive submit_scores path. An agent can distinguish it from all siblings without opening a schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives an explicit usage condition (runs with nobody in the loop) and an explicit alternative for the other case: 'In a chat, prefer get_call plus submit_scores.' This is exactly the when-to-use vs when-not-to-use routing the dimension asks for.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
submit_scoresA
Submit one score per rule for a call. The server checks every score before saving anything: the quote has to be in the turn you name (spaces and case are ignored), that turn has to be spoken by the rule's speaker, and with quote_required a true score needs a quote. Every rule of the version has to be present. If anything fails you get accepted=false with rejected[{rule_id, reason, hint}] and missing[]; nothing is saved, so fix those rules and send the full set again. A false score takes an empty quote. On success the server returns the verdict and flags it computed.
| Name | Required | Description | Default |
|---|---|---|---|
| scores | Yes | ||
| call_id | Yes | ||
| rules_version | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With annotations only covering the generic safety profile (readOnlyHint=false, destructiveHint=false, idempotentHint=false), the description carries rich extra context: validation is all-or-nothing ('the server checks every score before saving anything'), the exact failure contract (accepted=false with rejected[{rule_id, reason, hint}] and missing[]), quote matching semantics (case/spaces ignored, speaker must match, quote_required coupling), and the success payload (verdict flagged computed). This is exactly the beyond-annotations behavior an agent needs.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core action, then dense with genuinely load-bearing contract details; almost no filler. It is longer than average but the length is justified by atomic-validation and error-shape disclosure, and the single paragraph stays readable.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 3-required-param mutation tool with annotations and an output schema, the description covers the failure contract, retry path, quote rules, and success signal, so return values need not be re-explained. The omission of what call_id/rules_version refer to and the lack of contrast with score_call keep it just short of complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Top-level schema coverage is 0% for call_id, rules_version, and scores, but the description compensates heavily for the nested Score fields: quote must be verbatim from the named turn, turn must be spoken by the rule's speaker, true requires a quote and false requires an empty one, and rule_id corresponds to rules from get_rules. call_id and rules_version themselves are only implied (via get_call / 'the version'), leaving a small gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The opening sentence states a specific verb+resource ('Submit one score per rule for a call') and the 'every rule of the version has to be present' requirement implies batch semantics that distinguish it from the near-twin sibling score_call. However, it never names or contrasts with that sibling explicitly, so the agent must infer the distinction from the sibling's name alone.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives strong operational context: all rules of the version must be sent, a false score takes an empty quote, and on failure you 'fix those rules and send the full set again.' What's missing is any explicit routing guidance against the sibling score_call, so the agent is not told when this tool is preferred over a single-score alternative.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
team_reportARead-onlyIdempotent
Team view over scored calls: rate per rule, a table per rep, calls at risk with the quotes behind them, calls to follow up, and unscored calls. Every count lists the call ids it counts. Filter by rep and by date with since=YYYY-MM-DD.
| Name | Required | Description | Default |
|---|---|---|---|
| rep | No | ||
| since | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare the full safety profile (readOnly, idempotent, non-destructive), so the description is free to add behavioral context instead. It does: every count lists the call ids it counts, and the report is scoped to scored calls plus unscored-call visibility. This is meaningful disclosure beyond annotations, though pagination/size behavior is unaddressed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Content summary is front-loaded in the first sentence, with the filter detail last. Two sentences, no filler, dense but readable. Slightly list-heavy, but every clause describes actual report content.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With an output schema present, return structure needn't be spelled out, annotations cover safety, and a 2-param read-only report is low-complexity. The only real gap is the semantics of the 'rep' parameter, so completeness is high but not perfect.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description carries the burden. It compensates for 'since' by giving the format (since=YYYY-MM-DD) and states both params act as filters, but it says nothing about what 'rep' accepts (name, email, id), leaving one of two params underspecified.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific resource (a team-level report over scored calls) and enumerates its exact contents: rate per rule, per-rep table, at-risk calls with quotes, follow-ups, unscored calls. This distinguishes it well from list_calls, get_call, and score_call. It stops short of naming a sibling directly, so it lands at 4 rather than 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied by the enumerated contents (an agent can infer this is the aggregate 'team overview' tool rather than a single-call or write tool). However, no explicit when-to-use/when-not guidance or named alternatives are given, and the sibling tools are never referenced.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
8 tool updates
v0.1.0- First observed
deal_card_note - First observed
get_call - First observed
get_rules - First observed
list_calls - First observed
run_eval - First observed
score_call - First observed
submit_scores - First observed
team_report
TDQS
Scored across 8 tools
Most tools have clearly distinct purposes (list_calls vs get_call vs team_report vs run_eval). The only real overlap is submit_scores vs score_call, but the descriptions explicitly distinguish the human-in-the-loop vs automated path and recommend one, which resolves most confusion.
All names use snake_case and are readable. A few are noun phrases (deal_card_note, team_report) rather than verb_noun, a minor deviation from the otherwise consistent verb_noun pattern.
Eight tools is well-scoped for a call-scoring/QA domain, covering retrieval, scoring, reporting, and evaluation without redundancy.
The read/score/report/eval lifecycle is covered, but rule management is a gap: run_eval can test a draft rules_yaml yet there is no tool to create, save, or version a rule set, and no delete/unscore or call-ingestion operation is exposed.
Maintenance
Related MCP Connectors
Paid sales-call scoring and CRM next-step analysis for AI agents.
Sales coaching data in your AI assistant: calls, Influence Scores, objections and meeting prep.
Turn sales-call transcripts into traced proposals, contracts, and NDAs.
Conversational AI Coaching from calls; permissioned Team Dynamics reports in a limited U.S. pilot.
Related MCP Servers
- AlicenseNot gradedqualityNot gradedmaintenanceEnables intelligent conversation management with 4 AI agents that provide semantic analysis, pattern discovery, automatic documentation, and relationship mapping. Logs and analyzes Claude conversations with 70% token optimization and multi-language support.2-
- AlicenseAqualityDmaintenanceEnables prompt optimization loops and regression test suites for Claude Code, with a companion web UI for real-time visualization of scores and prompt revisions.19MIT
- FlicenseNot gradedqualityBmaintenanceEnables collection and scoring of sales chat conversations via CDP, computing response times and generating evidence-linked LLM quality reviews.-
- AlicenseNot gradedqualityBmaintenanceEnables analyzing call transcripts, scoring dimensions against declarative rubrics, and listing rubrics, for evidence-backed conversation evaluation.MIT