Skip to main content
Glama

Server Details

Calibrated judgments for text: yes/no probabilities, picks from your options, or scores.

If you are the author of this connector, you can claim ownership by verifying the domain or GitHub account it belongs to. Claimed connector authors can inspect health checks, view analytics, and manage their listing.
Status
Healthy
Last Tested
Transport
Streamable HTTP · MCP 2025-11-25
URL
Repository
Eddy1919/hunch-mcp
GitHub Stars
0
Server Listing
Hunch MCP

TDQS

A4.6/5.0

Scored across 8 tools

Disambiguation5/5

Each tool has a distinct purpose: hunch_ask (single yes/no), hunch_multi (several questions on one batch), hunch_pick (categorization), and hunch_score (ordered scale) are clearly separated judgment modes, while balance/buy_credits/get_key/quote are distinct account operations. The ask-vs-multi boundary is made explicit (one question vs up to 10), so misselection is unlikely.

Naming Consistency5/5

All eight tools use a uniform hunch_ prefix with a consistent verb_noun pattern (hunch_ask, hunch_buy_credits, hunch_get_key, hunch_quote, hunch_score). No mixed conventions or casing deviations.

Tool Count5/5

Eight tools is well within the ideal range, and each earns its place: four judging modes plus four supporting account/billing helpers. No redundancy or filler.

Completeness4/5

The domain (batch text judging with credit lifecycle) is covered end-to-end: judging, credit check, quoting, purchasing, and free-key acquisition all exist. Minor gaps like viewing past jobs/history or a batched payment-status tool are workable around.

Available Tools

8 tools
hunch_askHunch: yes/no probabilityAInspect

Judge a batch of short texts against one yes/no question and get back a calibrated probability (0 to 1) per text, not generated prose. Use it to score, tag, filter or triage many leads, support tickets, reviews, survey answers or emails at once, for example "Is this lead a decision maker?" or "Is this email urgent?". Prefer it to judging the texts yourself once there are more than about 25: one call returns a number per text and keeps the texts out of your context. Costs 1 credit per answered text (blank texts and texts repeated elsewhere in the same call are free; the same text asked again in a later call is charged again). Limits: the model reads the text only, no math, counting or dates; English works best; put the full definition of what counts as yes inside the question, since the model sees nothing else.

ParametersJSON Schema
NameRequiredDescriptionDefault
textsYesShort texts to judge (leads, tickets, reviews, survey answers, emails, ...), one answer per text. Blank entries and texts repeated elsewhere in the same call cost nothing. Chunked internally into calls of 40.
api_keyNoOnly for the keyless /mcp/try connection: a key from hunch_get_key or a purchase. Leave out to use the free sample. Ignored when the connection is already signed in.
questionYesA yes/no question, e.g. "Is this lead a decision maker who can approve a purchase without asking someone else?". Put the full definition of yes/no in the question text.

Output Schema

ParametersJSON Schema
NameRequiredDescription
sampleNoTrue when the answers came from the keyless free sample, so credits is the sample rows left today.
chargedYesCredits spent on this call.
creditsYesCredits left on the key after this call.
resultsYes
checkoutNoPresent when texts were skipped for lack of credits: a checkout link for the person to open (url, plan, price_usd, credits_added).

TDQS

A4.9/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Goes well beyond the annotations by disclosing the credit model (1 credit per answered text; blanks and in-call duplicates free; re-asking in a later call is charged again), the model's limits (text only, no math/counting/dates, English best), internal chunking into calls of 40, and the requirement to embed the full definition of 'yes' in the question.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core action and output, then use cases, the preference threshold, cost rules, and limits. Every sentence adds a distinct operational fact; nothing is padded or repeated.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, yet the description still characterizes the return value (a number per text), and it covers cost, batching, model limits, and question-authoring requirements. Nothing an agent needs to invoke this correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3, but the description adds meaningful semantics: the question must carry the complete yes/no definition because the model sees nothing else, and per-text cost depends on blank/duplicate handling. It still doesn't add much on api_key beyond the schema, keeping it short of a 5.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a precise verb+resource: judging a batch of short texts against one yes/no question and returning a calibrated probability (0-1) per text, explicitly 'not generated prose'. The yes/no framing plus the output type distinguishes it from a multi-class or generation sibling at a glance.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives explicit use cases (score, tag, filter, triage leads/tickets/reviews/emails) with concrete example questions, and a quantified decision rule: prefer it over judging yourself once there are more than ~25 texts. That is a genuine when-to-use threshold, not just a restatement of purpose.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

hunch_balanceHunch: check creditsA
Read-onlyIdempotent
Inspect

Check how many Hunch credits are left on this key and how many have been used so far. Read-only, costs nothing. Call it before a large batch, or when a judging tool reports texts were skipped for lack of credits. On the keyless /mcp/try connection it reports the free sample rows left today.

ParametersJSON Schema
NameRequiredDescriptionDefault
api_keyNoOnly for the keyless /mcp/try connection: a key from hunch_get_key or a purchase. Leave out to use the free sample. Ignored when the connection is already signed in.

Output Schema

ParametersJSON Schema
NameRequiredDescription
usedYesCredits used on this key so far, lifetime.
creditsYesCredits left on the key.

TDQS

A4.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, and destructiveHint=false. The description adds genuinely new context beyond them: calling it 'costs nothing' (it does not consume credits) and that on the keyless connection it reports free sample rows left today. It does not discuss rate limits or caching, so a 4 rather than 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short sentences, front-loaded with purpose, then cost profile, then call triggers. No filler; every sentence carries a distinct piece of information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists so return values need not be described. The description covers purpose, cost, trigger conditions, and the keyless-mode caveat — everything an agent needs to invoke it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the api_key semantics are already documented (baseline 3). The description adds meaning about connection modes — the keyless /mcp/try path and what the key does/does not apply to — which helps an agent decide whether to pass it.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('check how many Hunch credits are left on this key and how many have been used'), which is unambiguous against siblings like hunch_buy_credits or hunch_ask. An agent can identify the tool's function without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives explicit trigger conditions: 'Call it before a large batch, or when a judging tool reports texts were skipped for lack of credits.' It also clarifies the keyless /mcp/try behavior, so the agent knows which connection mode changes the result.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

hunch_buy_creditsHunch: get a checkout linkAInspect

Returns a link the person can open to buy more credits with a card on Stripe. Nothing is charged until they finish paying on that page; you cannot pay for them. Starter is $29 for 5,000 rows (a key's first purchase), Top-up is $19 for 25,000 rows on a key that already paid. Credits never expire and there is no subscription. Call it when a judging tool says the key is out of credits, or when hunch_quote says the job needs more than the key holds. Show the link to the person and let them decide.

ParametersJSON Schema
NameRequiredDescriptionDefault
planNoOptional. Defaults to Top-up when the key already paid, Starter otherwise.
api_keyNoOnly for the keyless /mcp/try connection: a key from hunch_get_key or a purchase. Leave out to use the free sample. Ignored when the connection is already signed in.

Output Schema

ParametersJSON Schema
NameRequiredDescription
urlYesCheckout link for the person to open.
planYes
price_usdYes
applies_toNo
credits_addedYes
expires_in_secondsNo

TDQS

A4.9/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only declare the mutation/idempotency profile; the description goes further by stating nothing is charged until payment completes, the agent cannot pay on the user's behalf, credits never expire, and there is no subscription. These are material behavioral facts not derivable from annotations or schema.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The core action (returns a link, nothing charged) is front-loaded, followed by plan economics and finally the call trigger. Every sentence carries information the agent needs; there is no padding.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With an output schema covering the return value and annotations covering the write profile, the description only needed to explain side effects and trigger conditions, and it does both. Nothing required to invoke it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3. The description adds real meaning beyond the schema by attaching prices and credit amounts to each enum value ($29/5,000 rows for Starter first purchase, $19/25,000 rows for Top-up) and explaining the keyless api_key case, which helps the agent pick a plan.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a concrete verb and resource: returns a checkout link for buying credits via Stripe. It is clearly distinguishable from siblings like hunch_balance or hunch_quote, which only read or estimate rather than produce a payment link.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicit trigger conditions are given: call it when a judging tool reports the key is out of credits, or when hunch_quote shows the job needs more than the key holds. It also adds a handling instruction (show the link, let them decide), which is exactly the routing guidance an agent needs.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

hunch_get_keyHunch: get a free 100-row keyAInspect

Returns a Hunch key with 100 free rows, in this response, with no email and no card. For the keyless /mcp/try connection: pass the key as api_key to the other tools (or reconnect with it). Limited to 2 keys per address per day. Treat the key like a password.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription
keyYesA hunch_ key. Secret.
creditsYes

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Goes well beyond the annotations by disclosing a concrete rate limit ('2 keys per address per day'), the credential nature of the output ('Treat the key like a password'), and that no email or card is required. This is exactly the operational context the readOnlyHint/destructiveHint flags cannot convey.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short sentences, front-loaded with the outcome (a free key) before the usage instructions and constraints. Every clause carries distinct information with no repetition of the title.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists so return values need no prose, and the description covers the rate limit, security handling, and next steps. It omits lifetime/expiry of the key and whether the 100 rows are one-time, minor gaps for a credential-issuing tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so the baseline is 4 and there is nothing for the description to disambiguate. The mention of 'api_key' refers to downstream tools, not to this tool's input, so it neither helps nor misleads here.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Returns a Hunch key') with a quantified scope ('100 free rows') and the distinguishing condition ('no email and no card'). This clearly separates it from the sibling tools, which all consume rather than issue credentials.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly states the context of use ('For the keyless /mcp/try connection') and what to do with the result ('pass the key as api_key to the other tools, or reconnect with it'). It does not say when not to call it (e.g., if a key already exists), which keeps it short of a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

hunch_multiHunch: several yes/no questions at onceAInspect

Ask up to 10 yes/no questions about the same batch of texts in one call, one probability per question per text, not generated prose. Use it when several judgments read the same text at once, for example "Can they buy?", "Are they angry?", "Is it urgent?" on the same support ticket, for a fraction of the tokens of separate calls. Prefer it to judging the texts yourself once there are more than about 25: one call returns a number per text and keeps the texts out of your context. Costs 1 credit per answered question per text (blanks and duplicate texts are free). Limits: the model reads the text only, no math, counting or dates, English works best, and each question needs its own definition of yes inside it.

ParametersJSON Schema
NameRequiredDescriptionDefault
textsYesShort texts to judge (leads, tickets, reviews, survey answers, emails, ...), one answer per text. Blank entries and texts repeated elsewhere in the same call cost nothing. Chunked internally into calls of 40.
api_keyNoOnly for the keyless /mcp/try connection: a key from hunch_get_key or a purchase. Leave out to use the free sample. Ignored when the connection is already signed in.
questionsYesUp to 10 yes/no questions, each answered once per text, e.g. ["Can they buy?", "Are they angry?", "Is it urgent?"].

Output Schema

ParametersJSON Schema
NameRequiredDescription
sampleNoTrue when the answers came from the keyless free sample, so credits is the sample rows left today.
chargedYes
creditsYes
resultsYes
checkoutNoPresent when texts were skipped for lack of credits: a checkout link for the person to open (url, plan, price_usd, credits_added).

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only mark non-readOnly/non-idempotent/non-destructive, which the description supersedes with real behavior: 1 credit per answered question per text, blanks and duplicate texts free, chunking into internal calls of 40, and limits (no math/counting/dates, English best). That is far beyond what the annotations provide.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core action and then the when-to-use rule, cost, and limits in descending priority. Dense but largely earned; the cost-and-limits tail is slightly packed into long sentences, costing a point.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With an output schema present, return values need no explanation, and the description still covers the caller's real unknowns: credit cost, free cases, chunking, and the model's blind spots. An agent has everything needed to call this correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is already 100%, so baseline is 3; the description adds meaning the schema lacks by constraining questions ('each question needs its own definition of yes inside it') and noting the 10-question cap ties to per-text cost. It does not explain the api_key parameter's behavior beyond what the schema says.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (Ask) and resource (up to 10 yes/no questions over a batch of texts) with the exact output form: one probability per question per text, not generated prose. This cleanly separates it from hunch_ask (single question) and hunch_score/pick siblings.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives an explicit selection rule with a threshold ('Prefer it to judging the texts yourself once there are more than about 25') plus alternatives it beats ('a fraction of the tokens of separate calls'). Concrete examples of batched judgments on one ticket make the intended use unambiguous.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

hunch_pickHunch: pick one optionAInspect

Sort a batch of short texts into one of your own categories and get back the chosen option plus how confident the model is, not generated prose. Use it to route support tickets, classify feedback, or tag leads by type, for example options ["billing: invoices and charges", "refund", "bug", "other"]. Prefer it to judging the texts yourself once there are more than about 25: one call returns a number per text and keeps the texts out of your context. Costs 1 credit per answered text (blanks and duplicates in the same call are free). Limits: 2 to 255 options, each "label" or "label: description" to disambiguate a short label; the model reads the text only, no math, counting or dates, English works best.

ParametersJSON Schema
NameRequiredDescriptionDefault
textsYesShort texts to judge (leads, tickets, reviews, survey answers, emails, ...), one answer per text. Blank entries and texts repeated elsewhere in the same call cost nothing. Chunked internally into calls of 40.
api_keyNoOnly for the keyless /mcp/try connection: a key from hunch_get_key or a purchase. Leave out to use the free sample. Ignored when the connection is already signed in.
optionsYesThe options to choose from, 2 to 255 of them. Each is "label" or "label: description" when the label alone is ambiguous, e.g. "billing: invoices and charges".
questionNoOptional. What is being decided, e.g. "Which category does this ticket belong to?". Defaults to "Which option best describes this text?".

Output Schema

ParametersJSON Schema
NameRequiredDescription
sampleNoTrue when the answers came from the keyless free sample, so credits is the sample rows left today.
chargedYes
creditsYes
resultsYes
checkoutNoPresent when texts were skipped for lack of credits: a checkout link for the person to open (url, plan, price_usd, credits_added).

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Goes well beyond the annotations (which only cover readOnly/openWorld/idempotent/destructive) by disclosing cost (1 credit per answered text, blanks and duplicates free), the 2-255 option limit, internal chunking at 40 (schema), the constraint that the model reads text only with no math/counting/dates, and that English works best. These are exactly the operational facts an agent needs before calling a paid classification tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Purpose and return value are front-loaded, followed by usage, cost, and limits, so every sentence earns its place. It is delivered as a single dense block rather than broken into scannable clauses, which costs a little readability but nothing is redundant.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With an output schema present, return-format detail is unnecessary, and the description still covers cost, limits, language constraints, and the sample-key parameter. For a paid, chunked batch-classification tool, nothing material for correct invocation is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3, but the description adds real meaning: the worked example options list ['billing: invoices and charges', 'refund', 'bug', 'other'] and the rationale for the 'label: description' form (disambiguating short labels). That is more than a restatement of the schema's type/enum info.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('sort a batch of short texts into one of your own categories') and pins the return shape ('the chosen option plus how confident the model is, not generated prose'), which separates it from prose-producing siblings like hunch_ask. An agent can tell what this tool produces without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives concrete use cases (route tickets, classify feedback, tag leads) and an explicit alternative with a trigger threshold: 'Prefer it to judging the texts yourself once there are more than about 25.' It never names a sibling tool (e.g. hunch_ask or hunch_multi) as the competing option, so routing among the hunch_* family still requires inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

hunch_quoteHunch: price a jobA
Read-onlyIdempotent
Inspect

Free, no key needed. Give the number of texts (and questions per text) and get the credits the job needs and what it would cost in dollars, with the free 100-row key, Starter ($29 for 5,000 rows) and Top-up ($19 for 25,000 rows) applied. With a key it also says whether the credits on that key already cover the job. Call it before a large batch so you can tell the person the price first.

ParametersJSON Schema
NameRequiredDescriptionDefault
rowsYesHow many texts the job has.
api_keyNoOnly for the keyless /mcp/try connection: a key from hunch_get_key or a purchase. Leave out to use the free sample. Ignored when the connection is already signed in.
questions_per_textNoYes/no questions per text for hunch_multi (one credit each). Leave out for the other tools.

Output Schema

ParametersJSON Schema
NameRequiredDescription
rowsYes
enoughYesTrue when the credits on the key (or the free 100 rows, with no key) already cover the job.
summaryYes
purchaseNoCheapest purchase that covers the gap: items (plan, quantity, credits, usd) and total_usd.
credits_leftNoCredits on the key. Only with a key.
credits_neededYesRows times questions per row.
to_buy_creditsNo

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnly/idempotent/non-destructive, and the description adds value beyond them: free, no key needed, keyless-try behavior, that api_key is ignored when signed in, and that with a key the response also reports whether existing credits cover the job. It stops short of describing output shape, which the output schema covers.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the key fact ('Free, no key needed') and keeps to three purposeful sentences. The pricing-tier enumeration is dense but directly relevant to the quote the tool returns; nothing is wasted, though it borders on verbose.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a read-only pricing tool with a full output schema and 100% schema coverage, the description covers cost model, key handling and the recommended call timing. Little an agent needs in order to invoke it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so rows, api_key and questions_per_text are already documented in the schema (including the 'leave out' semantics). The description adds useful pricing-tier context but no new per-parameter syntax or format beyond what the schema carries, so baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (price/quote) and resource (a job defined by texts and questions), and precisely says what comes back: credits needed and dollar cost. An agent can distinguish this from hunch_ask, hunch_multi and hunch_balance from the description alone.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives an explicit usage directive: 'Call it before a large batch so you can tell the person the price first.' It also explains the keyed vs keyless case, but names no alternative sibling or when-not-to-use condition.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

hunch_scoreHunch: score on a scaleAInspect

Place a batch of short texts on your own ordered scale (2 to 10 levels, low to high) and get back a probability-weighted position, the most likely level, and confidence, not generated prose. Use it for sentiment ("angry|disappointed|neutral|happy|delighted"), fit scoring ("no fit|weak|good|perfect"), or any low-to-high rating. Prefer it to judging the texts yourself once there are more than about 25: one call returns a number per text and keeps the texts out of your context. Costs 1 credit per answered text (blanks and duplicates in the same call are free). Limits: the model reads the text only, no math, counting or dates, English works best, and the question should say what is being scored.

ParametersJSON Schema
NameRequiredDescriptionDefault
textsYesShort texts to judge (leads, tickets, reviews, survey answers, emails, ...), one answer per text. Blank entries and texts repeated elsewhere in the same call cost nothing. Chunked internally into calls of 40.
levelsYesThe scale, low to high, 2 to 10 levels, e.g. ["angry", "disappointed", "neutral", "happy", "delighted"]. Each may be "label: description".
api_keyNoOnly for the keyless /mcp/try connection: a key from hunch_get_key or a purchase. Leave out to use the free sample. Ignored when the connection is already signed in.
questionYesWhat is being scored, e.g. "How does the reviewer feel about the product overall?".

Output Schema

ParametersJSON Schema
NameRequiredDescription
sampleNoTrue when the answers came from the keyless free sample, so credits is the sample rows left today.
chargedYes
creditsYes
resultsYes
checkoutNoPresent when texts were skipped for lack of credits: a checkout link for the person to open (url, plan, price_usd, credits_added).

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Adds substantial context beyond the annotations: credit cost of 1 per answered text, free blanks and duplicates, internal chunking into calls of 40, and explicit model limits (text only, no math/counting/dates, English works best). The cost disclosure is especially valuable given readOnlyHint=false, since the operation consumes credits.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single dense paragraph that front-loads the purpose and return shape, then layers examples, preference guidance, cost, and limits with no redundant sentences. Every clause carries information an agent would otherwise have to guess.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Complete for a batch scoring tool: purpose, scale constraints, cost model, capability limits, and usage threshold are all stated, and the output schema exists so return-value structure need not be explained. Nothing needed to call it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3, but the description adds cross-parameter semantics: levels are low-to-high and shown in concrete form, and the question should say what is being scored. This clarifies how texts, levels, and question fit together rather than merely restating field docs.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource: place a batch of short texts on an ordered scale and return a probability-weighted position, most likely level, and confidence. Explicitly says it returns numbers, not generated prose, which separates it from prose-generating siblings like hunch_ask and hunch_quote.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives concrete use cases (sentiment, fit scoring, any low-to-high rating) and an explicit prefer-this-over-doing-it-yourself threshold ('more than about 25'). It does not, however, route the agent among siblings such as hunch_pick or hunch_multi, which handle adjacent selection-style judgments.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 8 tool updates
    • First observedhunch_ask
    • First observedhunch_balance
    • First observedhunch_buy_credits
    • First observedhunch_get_key
    • First observedhunch_multi
    • First observedhunch_pick
    • First observedhunch_quote
    • First observedhunch_score

Related MCP Connectors

Related MCP Servers

  • A
    license
    A
    quality
    B
    maintenance
    Enables typed, calibrated judgment calls through classify, score, check, and batched ask tools, each returning full probability distributions for programmatic decisions.
    5
    1,141 npm
    8
    MIT
  • A
    license
    A
    quality
    A
    maintenance
    Enables agents to verify claims against cited evidence, screen content for prompt injection and relevance before reading it, and rank candidates by meaning, all with calibrated probability verdicts.
    11
    8,716 npm
    415
    MIT
Try in Browser

Glama MCP Gateway

Add one secure layer between your agents and this server.