tif-score-mcp
Server Quality Checklist
Latest release: v1.0.0
- Disambiguation3/5
get_lead_score and score_conversation both analyze call transcripts and clearly overlap — get_lead_score even notes it runs the same analysis as score_conversation. This creates real ambiguity about which to call. diagnose_workflow_error is completely distinct, but the two scoring tools blur boundaries significantly.
Naming Consistency3/5Tools use a consistent verb_noun pattern (get_lead_score, diagnose_workflow_error, score_conversation) with snake_case throughout. The naming styles are readable and consistent, though the verbs are somewhat varied (get, diagnose, score). No mixing of case conventions or chaotic patterns.
Tool Count2/5Three tools is at the thin end but not necessarily inappropriate. However, two of the three tools (get_lead_score and score_conversation) perform overlapping transcript analysis, making the actual distinct tool surface effectively two tools. The scope is unclear — the server mixes transcript scoring with completely unrelated n8n workflow diagnosis, suggesting no coherent singular purpose.
Completeness2/5The transcript-analysis side lacks balance: it covers extraction and scoring but has no update, correction, or batch-processing capability. The n8n workflow tool appears to be an orphan — a single diagnostic tool with no companion tools for related functionality. The two unrelated domains each feel incomplete, and the overall surface has no clear lifecycle coverage.
Average 4/5 across 3 of 3 tools scored.
See the Tool Scores section below for per-tool breakdowns.
- No community issues in the last 6 months
- 4 commits in the last 12 weeks
- No stable releases found
- No critical vulnerability alerts
- No high-severity vulnerability alerts
- No code scanning findings
- CI status not available
This repository is licensed under MIT License.
This repository includes a README.md file.
No tool usage detected in the last 30 days. Usage tracking helps demonstrate server value.
Tip: use the "Try in Browser" feature on the server page to seed initial usage.
Add a glama.json file to provide metadata about your server.
If you are the author, simply .
If the server belongs to an organization, first add
glama.jsonto the root of your repository:{ "$schema": "https://glama.ai/mcp/schemas/server.json", "maintainers": [ "your-github-username" ] }Then . Browse examples.
Add related servers to improve discoverability.
How to sync the server with GitHub?
Servers are automatically synced at least once per day, but you can also sync manually at any time to instantly update the server profile.
To manually sync the server, click the "Sync Server" button in the MCP server admin interface.
How is the quality score calculated?
The overall quality score combines two components: Tool Definition Quality (70%) and Server Coherence (30%).
Tool Definition Quality measures how well each tool describes itself to AI agents. Every tool is scored 1–5 across six dimensions: Purpose Clarity (25%), Usage Guidelines (20%), Behavioral Transparency (20%), Parameter Semantics (15%), Conciseness & Structure (10%), and Contextual Completeness (10%). The server-level definition quality score is calculated as 60% mean TDQS + 40% minimum TDQS, so a single poorly described tool pulls the score down.
Server Coherence evaluates how well the tools work together as a set, scoring four dimensions equally: Disambiguation (can agents tell tools apart?), Naming Consistency, Tool Count Appropriateness, and Completeness (are there gaps in the tool surface?).
Tiers are derived from the overall score: A (≥3.5), B (≥3.0), C (≥2.0), D (≥1.0), F (<1.0). B and above is considered passing.
Tool Scores
- Behavior3/5
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the disclosure burden. It clearly states the tool analyzes and produces derived metrics but does not reveal whether there are side effects (e.g., persistence, write-back to CRM), whether the raw transcript is retained, or what happens on unprocessable input. For an analysis tool this is moderate transparency, but no side-effect or input-limit behavior is mentioned.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Conciseness4/5Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single dense sentence that front-loads the action and enumerates outputs efficiently. It is complete without being verbose, though it could arguably be split for readability. No wasted words, earns its length.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Completeness3/5Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool has 4 required parameters, no output schema, and no annotations, so the description is the sole guide. It explains the broad outputs well, but with no output schema it does not communicate the return format or structure of the extracted intelligence, and the sibling get_lead_score suggests a related scoring ecosystem whose differentiation is unresolved. Adequate but leaves room for more.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Parameters3/5Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 75%, so parameters are largely self-documenting (raw_outcome gets an example, call_duration gets units, source_system gets examples). The description adds that analysis is structured 'business intelligence' but does not clarify how call_duration or source_system influence the scoring, which would add value beyond the schema. One parameter has no description in schema, though the description partially covers the overall intent.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Purpose5/5Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('Analyze') with a clear resource ('a call transcript') and enumerates the exact structured outputs it produces: intent, sentiment, urgency, objections, outcome, topics, flags, revenue signal, churn risk, and a summary. This clearly distinguishes it from sibling tools like get_lead_score (score retrieval) and diagnose_workflow_error (error diagnostics).
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Usage Guidelines3/5Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage for call transcripts yielding business intelligence but provides no explicit when-to-use vs alternatives guidance. It doesn't exclude scenarios or direct the agent toward when the score_conversation tool is preferable over get_lead_score, which could be relevant given 'revenue signal' and 'churn risk' outputs overlap thematically.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
- Behavior4/5
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description discloses a key behavior: that it internally reuses score_conversation's analysis before applying weighting. This gives the agent insight into the tool's internal mechanics. It also notes the repeat_caller param depends on external lookup, adding context. No annotations exist, so the description carries the burden, and it discloses the reuse behavior that would otherwise be invisible.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Conciseness4/5Is the description appropriately sized, front-loaded, and free of redundancy?
Two concise sentences that state purpose, output, and the key relationship to the sibling tool. Efficient and front-loaded with the outcome. Could potentially clarify the weighting scheme, but no wasted words.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Completeness4/5Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a scoring tool with 5 params (4 required) and no output schema, the description communicates the core output (0-100 score, hot/warm/cold classification) and the internal pipeline. Missing return-format detail (e.g., does it return a JSON object, structured fields?) but given no output schema and moderate complexity, this is reasonably complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Parameters3/5Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 80%, with descriptions on raw_outcome, call_duration, repeat_caller, and source_system. The tool description doesn't elaborate on parameter semantics beyond what the schema provides, but that's acceptable given high coverage. The behavior it adds about repeat_caller being externally determined aligns with the schema's note. Baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Purpose5/5Does the description clearly state what the tool does and how it differs from similar tools?
Clear verb+resource ('Score a call transcript 0-100 and classify it hot/warm/cold'). Describes what it does, what it returns, and explicitly distinguishes from sibling score_conversation by noting it 'Internally runs the same analysis' then applies 'point-based weighting'.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Usage Guidelines4/5Does the description explain when to use this tool, when not to, or what alternatives exist?
The description clearly distinguishes from score_conversation by noting the internal reuse and additional weighting step, implying get_lead_score builds on the base scoring. However, it doesn't explicitly state when a user should choose this tool over score_conversation or diagnose_workflow_error, leaving some ambiguity about the decision.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
- Behavior4/5
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries full burden and handles it well: it explicitly declares statelessness, no live fetching/verification/deployment, and the honest 'not fixable' output possibility. It also warns about applying fixedWorkflow without review. This is strong behavioral disclosure absent any annotation support.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Conciseness5/5Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, zero filler. The first sentence packs purpose, inputs, and outputs. The second clarifies the stateless boundary. The third advises on safe application of the result. Every sentence earns its place with front-loaded core purpose.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Completeness4/5Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 2-param tool with nested objects but no output schema, the description covers purpose, behavioral scope, and honest-limit disclosures well. Missing output schema means the agent can't see the return shape (rootCause/fixedWorkflow/confidence structure), but the description names the return elements in prose, partially compensating. Could add parameter format specifics, but overall adequate.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Parameters3/5Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 50%: errorMessage has only a minLength constraint, and workflowJson has a description ('The full n8n workflow JSON object (nodes + connections)'). The tool description adds input context but doesn't elaborate much beyond the schema. The description implies how parameters are consumed (as inputs to diagnosis) but adds limited format/semantic detail beyond what the schema states.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Purpose5/5Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('diagnose') plus resource (n8n workflow error), states the inputs (workflow JSON + error message), and clearly lists the outputs (root cause, proposed fix, confidence level). It also distinguishes from siblings by being stateless diagnosis while get_lead_score and score_conversation imply scoring/analytics tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Usage Guidelines4/5Does the description explain when to use this tool, when not to, or what alternatives exist?
The description clearly states what it does NOT do ('does not verify, deploy, or fetch anything live') and advises pairing with human or separate review. It doesn't explicitly name alternative tools for when-not-to-use, but the stateless scope and non-verification caveat provide clear contextual boundaries.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
GitHub Badge
Glama performs regular codebase and documentation scans to:
- Confirm that the MCP server is working as expected.
- Confirm that there are no obvious security issues.
- Evaluate tool definition quality.
Our badge communicates server capabilities, safety, and installation instructions.
Card Badge
Copy to your README.md:
Score Badge
Copy to your README.md:
Latest Blog Posts
- Who's Calling? MCP Hosts Are an Identity Blind Spot (And the Spec Knows It)By Om-Shree-0709 on .mcpAgent IdentityOAuth 2.1
- Your AI Chatbot Just Exposed Your CEO's Salary to an InternBy Om-Shree-0709 on .Agent IdentityMCP SecurityOAuth Delegation
- Why MCP Servers Need Execution Sandboxing (And Why Your Current Stack Isn't Enough)By Om-Shree-0709 on .Agentic AiPrompt InjectionWebAssembly
MCP directory API
We provide all the information about MCP servers via our MCP API.
curl -X GET 'https://glama.ai/api/mcp/v1/servers/dsvxmedia/tif-score-mcp'
If you have feedback or need assistance with the MCP directory API, please join our Discord server