Skip to main content
Glama
Kungie

gutfeel-mcp

gut

PyPI CI Downloads

Judgment calls as one line of Python — built for Jev, and running on any small model.

Try it in your browser → The site runs gut's local model in the page: no key, no server.

Your code keeps running into questions that aren't logic: Is this comment spam? Which team owns this ticket? How urgent is it? Is the agent's task done? Until now there were three answers:

  • Regex and keyword rules — free and instant, and wrong the moment someone phrases it differently.

  • A frontier LLM — understands anything, at seconds and cents a call, with prose to parse.

  • Train a classifier — cheap to run, once you have the labelled data, the pipeline and the week.

There is a fourth: a small model made for exactly these questions. TypeSafe AI's Jev answers typed questions directly — a probability for yes, a distribution over options, a score on a scale — with nothing to generate or parse, billed on input only. gut is built around it, and makes it a line of code:

import gut

gut.configure(backend=gut.JevBackend())   # or just set TYPESAFE_API_KEY

if gut.likely(comment, "is spam"):
    hide(comment)

No prompt, no parsing, no threshold — and no model named at the call site.

Three questions

gut.likely(ticket, "is a bug report")                      # yes / no
gut.classify(ticket, Team)                                 # which one — an Enum
gut.rate(ticket, ["can wait", "this week", "right now"])   # how much

Related MCP server: jev-flash-router

It knows when it doesn't know

A regex never hesitates, and neither does an LLM. gut can:

match gut.likely(email, "the customer threatens to cancel", ask_human=True):
    case gut.YES:    escalate(email)
    case gut.NO:     auto_reply(email)
    case gut.UNSURE: send_to_a_person(email)

Say how careful to be in words — lean="yes", stakes="high" — and gut works out the thresholds.

A thousand subjects, one line

spam = gut.each(comments).likely("is spam")     # one decision per comment, in order
teams = gut.each(tickets).classify(Team)

Jev gets concurrent requests, a local model batched passes, and nothing already cached is asked twice. @gut.semantic does the same for several questions about one subject.

Jev first, any model

Jev is the model gut is designed around. It is not the only one: the model is configuration, and the same line runs unchanged on any of these.

gut.configure(backend=gut.JevBackend())                            # TypeSafe AI's Jev
gut.configure(backend=gut.JevBackend.openrouter())                 # Jev, through OpenRouter
gut.configure(backend=gut.JevBackend.ollaya("winnow:e4b"))         # open decision model, Ollaya
gut.configure(backend=gut.ZeroShotBackend())                       # NLI model, on your CPU
gut.configure(backend=gut.TransformersBackend("Qwen/Qwen3-0.6B"))  # small LLM, on your machine
gut.configure(backend=gut.OpenAICompatibleBackend(                 # Ollama, vLLM, llama.cpp
    "qwen2.5:1.5b", base_url="http://localhost:11434/v1"))
gut.configure(backend=gut.OpenAICompatibleBackend("gpt-4.1-nano")) # OpenAI

Or several at once. Cascade asks the cheapest model first and passes on only what it is unsure of:

gut.configure(backend=gut.Cascade(
    gut.ZeroShotBackend(),   # free and local: settles the obvious
    gut.JevBackend(),        # sees only what the first could not
))

Every answer is a model's own probabilities, never parsed from text, and decision.model names the model that gave it. Your own model can be a backend too: here is how.

Install

pip install "gutfeel[jev]"           # + JevBackend: TypeSafe, OpenRouter or Ollaya
pip install "gutfeel[local]"         # + ZeroShotBackend and TransformersBackend (PyTorch)
pip install gutfeel                  # core: any OpenAI-compatible server; FakeBackend for tests
pip install gutfeel-mcp             # + the MCP server, gutfeel-mcp

The package on PyPI is gutfeel (gut was taken); the import is plain import gut. No model at hand? gut.FakeBackend(answers={"is spam": 0.97}) answers from fixtures, for tests.

From a shell, and for agents

The big model thinks; the small one decides, fast. An agent with gut judges a thousand files, commits or search results in one command instead of reading each one itself:

git ls-files | gut filter "retries failed requests" --read-files --max-cost 0.50
claude mcp add gut --env TYPESAFE_API_KEY=your-key -- uvx gutfeel-mcp
npx skills add Kungie/gut --skill gut     # teaches a coding agent when to reach for it

Docs

The documentation, one page per idea: Getting started · Backends · Knowing when it doesn't know · Asking everything at once · Async · Exact costs · Caching and observability · Command line · MCP server · Honest limitations. examples/ runs the same code on every backend.

Status and license

Pre-1.0, Apache-2.0. Every code block in these docs runs in the test suite · contributing

Available Tools

4 tools
classifyA
Read-onlyIdempotent

Pick the option that fits a text best: a team, a topic, a language, an intent. Answers with the chosen label, every option's probability and a confidence.

ParametersJSON Schema
NameRequiredDescriptionDefault
stakesNoHow sure the model must be before an answer counts. Needs ask_human.
optionsYesThe categories: {label: description}, or a list of labels. Include a catch-all such as "other", or every text is forced into one of them.
subjectYesThe text to judge: an email, a comment, a diff.
questionNoWhat to decide, if the options do not say it.
ask_humanNoAllow the outcome `unsure` when the model cannot tell. Off by default.

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.9/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnly, idempotent, and closed-world behavior, so the bar is lower. The description adds genuine behavioral value by disclosing the return shape (chosen label, per-option probabilities, confidence), which the annotations do not cover.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tightly written sentences with no filler; the core action is front-loaded and the output behavior follows. Every clause earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be re-explained, and all 5 parameters carry full schema descriptions. The description is complete enough to invoke the tool correctly, though it could note the stakes/ask_human interaction more explicitly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3, but the description goes further by illustrating what 'options' actually look like in practice (team/topic/language/intent), giving the agent a concrete model of the parameter beyond the schema's structural typing.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (pick/classify) and resource (a text against options), with concrete examples of option kinds (team, topic, language, intent). It is clear what the tool does, but it makes no attempt to distinguish itself from siblings like likely, rate, or each.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is implied by the description — classify a text into the provided option set — but there is no explicit when-to-use, when-not-to-use, or routing away from sibling tools such as rate or likely. The agent must infer the selection criteria.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

eachA
Read-onlyIdempotent

Ask the same question about many texts at once: likely with just a question, classify with options, rate with levels. Far faster than one call per text.

ParametersJSON Schema
NameRequiredDescriptionDefault
levelsNoA scale, as for rate.
stakesNoHow sure the model must be before an answer counts. Needs ask_human.
optionsNoCategories, as for classify.
questionNoA claim, as for likely. Alone, it asks likely; with options or levels, it says what to decide.
subjectsYesThe texts to judge, all with the same question.
ask_humanNoAllow the outcome `unsure` when the model cannot tell. Off by default.

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.8/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnly, idempotent, and closed-world, so the safety profile is covered. The description adds a performance claim ('far faster'), which is useful context, but says nothing about how batch failures or per-item uncertainty behave — the stakes/ask_human interaction is only in the schema.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, zero filler. The mode mapping comes first and the efficiency argument closes it out; every clause carries information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With an output schema present, return values need no explanation, and all six params are documented in the schema. The only real gap is whether results are returned in subjects order and how per-item errors surface, which the description doesn't touch.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already explains subjects, question, options, levels, stakes, and ask_human. The description restates the question/options/levels relationship in prose without adding syntax or format detail beyond the schema, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific operation (ask the same question about many texts) and maps it onto the three sibling modes (likely/classify/rate) via 'likely with just a question, classify with options, rate with levels.' The name 'each' alone is opaque, but the text rescues it and an agent can distinguish this batch tool from the single-item siblings.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives clear selection logic via the question/options/levels mapping and a reason to prefer it ('Far faster than one call per text'). It does not explicitly say to fall back to likely/classify/rate for a single text, so the when-not is left to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

likelyA
Read-onlyIdempotent

Judge whether a claim is true of a text. Answers yes or no with the model's probability that the claim is true -- or unsure, when ask_human is set.

ParametersJSON Schema
NameRequiredDescriptionDefault
leanNoWhich way to err when unsure is not allowed.
stakesNoHow sure the model must be before an answer counts. Needs ask_human.
subjectYesThe text to judge: an email, a comment, a diff.
questionYesA short claim about the text, e.g. "asks for a refund". Not an open question.
ask_humanNoAllow the outcome `unsure` when the model cannot tell. Off by default.

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations cover the safety profile (readOnly, idempotent, closed-world), so the bar is lower. The description adds real value by disclosing the outcome model: yes/no plus a probability, with an 'unsure' escape hatch gated on ask_human. It does not discuss failure modes or confidence thresholds beyond stakes, but the added context is meaningful.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tight sentences that lead with the core action and immediately qualify the return semantics. No filler; only the mildly awkward line break across 'model's / probability' costs it a point.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With an output schema present, return-value detail need not live in the description, and all five parameters are documented with descriptions. The description plus structured fields give an agent enough to call the tool correctly; only cross-tool routing guidance is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents subject, question, lean, stakes, and ask_human. The description only restates the effect of ask_human (unsure outcome), adding no syntax or constraints beyond the schema, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource: 'Judge whether a claim is true of a text,' which is concrete and actionable. It does not explicitly differentiate itself from the sibling tools classify, rate, and each, so it falls short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The clause 'or unsure, when ask_human is set' implies a usage condition, but there is no explicit when-to-use-this-vs-alternatives guidance and no exclusions. The relationship to the sibling judging tools is left entirely to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

rateA
Read-onlyIdempotent

Rate a text on an ordered scale such as urgency or severity. Answers with a score, which can fall between levels, the nearest level and a confidence.

ParametersJSON Schema
NameRequiredDescriptionDefault
levelsYesBetween 2 and 10 level descriptions, in order; the first is level 0.
stakesNoHow sure the model must be before an answer counts. Needs ask_human.
subjectYesThe text to judge: an email, a comment, a diff.
questionNoWhat to rate, if the levels do not say it.
ask_humanNoAllow the outcome `unsure` when the model cannot tell. Off by default.

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, and openWorldHint=false, so the safety and side-effect profile is covered. The description adds output behavior by saying answers can fall between levels and include a nearest level and confidence, but it does not mention the schema-level unsure outcome enabled by ask_human or the effect of stakes.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two efficient sentences with no filler. The purpose is front-loaded, followed immediately by the output shape, making it easy for an agent to scan.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the rich schema, output schema, and read-only/idempotent annotations, the description is nearly complete for invocation. The main remaining gap is sibling-relative guidance for choosing rate over classify, likely, or each.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all five parameters are already documented in the schema. The description adds only the general notion of an ordered scale and does not provide additional parameter-level meaning beyond what the schema already supplies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: 'Rate a text on an ordered scale such as urgency or severity.' It also characterizes the distinctive output ('a score, which can fall between levels, the nearest level and a confidence'), which helps separate it from categorical siblings like classify, though it does not name an alternative directly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The phrase 'on an ordered scale such as urgency or severity' implies the use case and distinguishes it somewhat from discrete classification, but there is no explicit when-to-use or when-not-to-use guidance relative to siblings such as classify, likely, or each.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 4 tool updatesv0.1.0
    • First observedclassify
    • First observedeach
    • First observedlikely
    • First observedrate

TDQS

A3.7/5.0

Scored across 4 tools

Disambiguation4/5

likely, classify, and rate have clearly distinct output types (binary truth, categorical label, ordinal score). 'each' overlaps as a batch version of all three, which could cause some confusion if the description is not read, but its purpose is clarified as handling many texts at once.

Naming Consistency2/5

Names are all lowercase single words but mix parts of speech: classify and rate are verbs, while likely is an adjective and each is a determiner. There is no consistent verb_noun or action-oriented pattern, making the set less predictable.

Tool Count5/5

Four tools is a lean but appropriate set for a focused text-judgment service. Each tool covers a distinct primitive or batch variant, with no obvious redundancy.

Completeness4/5

The surface covers binary judgment, multi-class classification, ordinal rating, and batch versions of all three, forming a complete core for quick text judgments. Minor gaps exist (e.g., no explicit pairwise comparison or human escalation for classify/rate), but agents can work around them.

Maintenance

ActivityNo data
ResponsivenessNo issues

Related MCP Connectors

  • Jev-powered decisions, web search, PDF/web to Markdown, summarize. From $0.001, no API key.

  • Deterministic AI agent microtools, no accounts/API keys. fetch_extract: 98% token cut. 38 tools.

  • Search Fragments — two tools for the queries an agent can't place, both built to decline rather than guess. resolve_fragment takes a half-remembered, cross-source query ("a musician who became famous for stopping performing") and returns a grounded answer, ranked web sources to confirm by eye, or an explicit no-resolution. verify_claim takes a specific factual assertion and returns supported, partially_supported, insufficient_evidence, or unsupported, with cited evidence and a stated_limits field that is always present. There is no confidence score — insufficient_evidence fires freely, and unsupported requires a source that explicitly contradicts, never mere absence of confirmation. Every verdict is decide-by-eye: "supported" means current web sources confirm it, not that the claim is true. Calibrated against 18 known claims before release. Free, no signup. Streamable HTTP (MCP 2025-11-25). Read-only.

  • Deterministic decision layer for autonomous agents: reproducible PROCEED/REVIEW/SKIP verdicts.

Related MCP Servers

  • A
    license
    A
    quality
    A
    maintenance
    Enables agents to get fast, calibrated probabilistic answers from Jev (Typesafe AI) to yes/no, scale, or choice questions about provided material, without using a generative model.
    1
    208 npm
    2
    MIT
  • A
    license
    A
    quality
    B
    maintenance
    Enables AI coding agents to make fast, zero-output-token decisions by evaluating context, diffs, logs, or options through the OpenRouter Decisions API using TypeSafe Jev, returning calibrated probabilities for binary, categorical, or scoring questions.
    1
    552 npm
    4
    MIT
  • A
    license
    B
    quality
    C
    maintenance
    Enables AI agents to obtain typed judgments from TypeSafe's Jev System One models, including yes/no probabilities, multiple-choice selections with distributions, and rubric-based scores, directly usable in code.
    5
    2
    AGPL 3.0
  • A
    license
    A
    quality
    B
    maintenance
    Connect AI agents to Jev AI (jev-ai.pro) for classification, scoring, yes/no checks, advisory action assessment, batched typed decisions and saved judges. Returns structured answers and probabilities using your Jev AI API key.
    6
    MIT