callsense-mcp
Server Configuration
Describes the environment variables required to run the server.
| Name | Required | Description | Default |
|---|---|---|---|
No arguments | |||
Instructions
Guidance the server publishes about itself, which clients place ahead of the tool catalog so the model reads it before choosing anything.
This server publishes no instructions, or was last inspected before Glama recorded them.
Capabilities
Features and capabilities supported by this server
Protocol revision2025-11-25
| Capability | Details |
|---|---|
| tools | {
"listChanged": false
} |
| prompts | {
"listChanged": false
} |
| resources | {
"subscribe": false,
"listChanged": false
} |
| experimental | {} |
Tools
Functions exposed to the LLM to take actions
| Name | Description |
|---|---|
| list_callsA | List calls with date, rep, customer, number of turns, whether the call is scored, and its verdict. Filter by rep name and by scored=true/false. |
| get_callB | One call with numbered turns {n, speaker, text}. Quote from these turns when scoring. |
| get_rulesB | Scoring rules of a version (latest by default): id, question, what counts as evidence, and speaker, the side whose own words the evidence has to be. Also says whether quotes are required. |
| submit_scoresA | Submit one score per rule for a call. The server checks every score before saving anything: the quote has to be in the turn you name (spaces and case are ignored), that turn has to be spoken by the rule's speaker, and with quote_required a true score needs a quote. Every rule of the version has to be present. If anything fails you get accepted=false with rejected[{rule_id, reason, hint}] and missing[]; nothing is saved, so fix those rules and send the full set again. A false score takes an empty quote. On success the server returns the verdict and flags it computed. |
| score_callA | Score a call on the server, for runs with nobody in the loop. 'baseline' is a keyword floor that needs no key. 'claude' calls the Anthropic API and needs ANTHROPIC_API_KEY in the server's environment. The result goes through the same evidence guard as submit_scores and is saved only if it passes. In a chat, prefer get_call plus submit_scores. |
| team_reportA | Team view over scored calls: rate per rule, a table per rep, calls at risk with the quotes behind them, calls to follow up, and unscored calls. Every count lists the call ids it counts. Filter by rep and by date with since=YYYY-MM-DD. |
| run_evalA | Measure a rule set against the hand-graded calls before it ships: agreement with human labels per rule, quote validity, and pass or fail. Every rule and the total have to reach min_agreement. Pass rules_yaml to test a draft rule set without saving it. A rule with no human labels cannot pass. |
| deal_card_noteA | The note the CRM write-back would put on the deal card: verdict, flags, every score with its quote, and the deal-field payload. Builds the text only; nothing is sent. |
Prompts
Interactive templates invoked by user choice
| Name | Description |
|---|---|
| coach_rep | Weekly coaching note for one rep, from their scored calls and the quotes behind each score. |
| score_unscored | Score every unscored call with get_rules, get_call and submit_scores, fixing rejected quotes, then run team_report. |
Resources
Contextual data attached and managed by the client
| Name | Description |
|---|---|
No resources | |
TDQS
Scored across 8 tools
Most tools have clearly distinct purposes (list_calls vs get_call vs team_report vs run_eval). The only real overlap is submit_scores vs score_call, but the descriptions explicitly distinguish the human-in-the-loop vs automated path and recommend one, which resolves most confusion.
All names use snake_case and are readable. A few are noun phrases (deal_card_note, team_report) rather than verb_noun, a minor deviation from the otherwise consistent verb_noun pattern.
Eight tools is well-scoped for a call-scoring/QA domain, covering retrieval, scoring, reporting, and evaluation without redundancy.
The read/score/report/eval lifecycle is covered, but rule management is a gap: run_eval can test a draft rules_yaml yet there is no tool to create, save, or version a rule set, and no delete/unscore or call-ingestion operation is exposed.