Skip to main content
Glama
dsichz

callsense-mcp

by dsichz

submit_scores

Submit one score per playbook rule for a call. The server validates quotes, speakers, and rule completeness, saving all scores only when every check passes; otherwise it returns reasons and hints.

Instructions

Submit one score per rule for a call. The server checks every score before saving anything: the quote has to be in the turn you name (spaces and case are ignored), that turn has to be spoken by the rule's speaker, and with quote_required a true score needs a quote. Every rule of the version has to be present. If anything fails you get accepted=false with rejected[{rule_id, reason, hint}] and missing[]; nothing is saved, so fix those rules and send the full set again. A false score takes an empty quote. On success the server returns the verdict and flags it computed.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
scoresYes
call_idYes
rules_versionYes

Output Schema

TableJSON Schema
NameRequiredDescriptionDefault

No arguments

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observedv0.1.0

TDQS

A4.2/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With annotations only covering the generic safety profile (readOnlyHint=false, destructiveHint=false, idempotentHint=false), the description carries rich extra context: validation is all-or-nothing ('the server checks every score before saving anything'), the exact failure contract (accepted=false with rejected[{rule_id, reason, hint}] and missing[]), quote matching semantics (case/spaces ignored, speaker must match, quote_required coupling), and the success payload (verdict flagged computed). This is exactly the beyond-annotations behavior an agent needs.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core action, then dense with genuinely load-bearing contract details; almost no filler. It is longer than average but the length is justified by atomic-validation and error-shape disclosure, and the single paragraph stays readable.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 3-required-param mutation tool with annotations and an output schema, the description covers the failure contract, retry path, quote rules, and success signal, so return values need not be re-explained. The omission of what call_id/rules_version refer to and the lack of contrast with score_call keep it just short of complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Top-level schema coverage is 0% for call_id, rules_version, and scores, but the description compensates heavily for the nested Score fields: quote must be verbatim from the named turn, turn must be spoken by the rule's speaker, true requires a quote and false requires an empty one, and rule_id corresponds to rules from get_rules. call_id and rules_version themselves are only implied (via get_call / 'the version'), leaving a small gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The opening sentence states a specific verb+resource ('Submit one score per rule for a call') and the 'every rule of the version has to be present' requirement implies batch semantics that distinguish it from the near-twin sibling score_call. However, it never names or contrasts with that sibling explicitly, so the agent must infer the distinction from the sibling's name alone.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives strong operational context: all rules of the version must be sent, a false score takes an empty quote, and on failure you 'fix those rules and send the full set again.' What's missing is any explicit routing guidance against the sibling score_call, so the agent is not told when this tool is preferred over a single-score alternative.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.