Skip to main content
Glama
Kungie

gutfeel-mcp

likely

Read-onlyIdempotent

Judge whether a claim about an email, comment, or diff is true, returning yes, no, or unsure with the model's probability when human fallback is allowed.

Instructions

Judge whether a claim is true of a text. Answers yes or no with the model's probability that the claim is true -- or unsure, when ask_human is set.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
leanNoWhich way to err when unsure is not allowed.
stakesNoHow sure the model must be before an answer counts. Needs ask_human.
subjectYesThe text to judge: an email, a comment, a diff.
questionYesA short claim about the text, e.g. "asks for a refund". Not an open question.
ask_humanNoAllow the outcome `unsure` when the model cannot tell. Off by default.

Output Schema

TableJSON Schema
NameRequiredDescriptionDefault

No arguments

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observedv0.1.0

TDQS

A3.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations cover the safety profile (readOnly, idempotent, closed-world), so the bar is lower. The description adds real value by disclosing the outcome model: yes/no plus a probability, with an 'unsure' escape hatch gated on ask_human. It does not discuss failure modes or confidence thresholds beyond stakes, but the added context is meaningful.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tight sentences that lead with the core action and immediately qualify the return semantics. No filler; only the mildly awkward line break across 'model's / probability' costs it a point.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With an output schema present, return-value detail need not live in the description, and all five parameters are documented with descriptions. The description plus structured fields give an agent enough to call the tool correctly; only cross-tool routing guidance is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents subject, question, lean, stakes, and ask_human. The description only restates the effect of ask_human (unsure outcome), adding no syntax or constraints beyond the schema, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource: 'Judge whether a claim is true of a text,' which is concrete and actionable. It does not explicitly differentiate itself from the sibling tools classify, rate, and each, so it falls short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The clause 'or unsure, when ask_human is set' implies a usage condition, but there is no explicit when-to-use-this-vs-alternatives guidance and no exclusions. The relationship to the sibling judging tools is left entirely to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Deploy Server

Other Tools