humor-mcp
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@humor-mcpsearch for jokes about technology"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
humor-mcp
An MCP server over a local humor corpus. Every line it returns says who wrote it.
Point a model at a body of jokes and it will happily launder them: you get the material back with no idea whose it was, whether you may publish it, or whether you may sell it. This server treats that as the primary problem. Every result carries the credit for the line it came from, material that may not be redistributed is withheld unless you ask for it by name, and the exporter refuses to hand out anything whose licence does not permit it.
No database server, no network, no required dependencies — Python stdlib and one SQLite file. It runs the same whether or not the machine that built the corpus is awake.
(That holds for the server and every text path. Audio ingest is the one
exception: it needs numpy, scipy and soundfile, and is only imported when
you actually use it.)
packs/<id>/pack.json who made this material, under what licence
packs/<id>/lines.jsonl the lines
packs/<id>/pairs.jsonl chosen/rejected preference pairs (optional)
humor_mcp/cli.py the one entry point; subcommands below
server.py the MCP server (stdio, read-only)
build.py packs -> humor.db (SQLite + FTS5)
import_audio.py audio + timed transcript (measured reactions)
import_transcript.py transcript (text reaction markers)
import_corpus.py a jokes file
audio_reactions.py the laughter/applause detector
transcripts.py shared .srt/.vtt/.json parsing
paths.py where the bundled and user corpora live
server.py shim, so a checkout runs with nothing installedInstall
Python 3.9+, no required dependencies. One line:
claude mcp add humor -- uvx humor-mcpThat is the whole install. A corpus ships inside the package — 1,248 lines with 975 human ratings and the best-of-N picks — so there is nothing to download, no build step, and nothing to wire together. The first launch compiles the database itself; your client just starts working. No account, no checkout.
To run the unreleased tip instead:
uvx --from git+https://github.com/zoombulous/humor-mcp humor-mcpCloned — adds the third-party packs (CUP, ExPUNations, r/Jokes, New Yorker), which are too large and too non-commercial to bundle:
git clone https://github.com/zoombulous/humor-mcp
cd humor-mcp
python -m humor_mcp.cli build
claude mcp add humor -- python "$(pwd)/server.py"On Windows use the full path in place of $(pwd). Any MCP client works; this
speaks plain stdio JSON-RPC over stdin/stdout, and running the command with no
arguments is what starts the server.
Audio ingest is the one optional extra:
pip install "humor-mcp[audio]" # or: pip install numpy scipy soundfileCommands
humor-mcp serve on stdio <- what an MCP client runs
humor-mcp where which corpus directories are in use
humor-mcp build compile packs -> humor.db
humor-mcp build --export DIR copy out only what may be redistributed
humor-mcp import-corpus a .txt/.csv/.jsonl of jokes
humor-mcp import-transcript a transcript
humor-mcp import-audio audio + a timed transcript
humor-mcp reactions FILE what the detector hears in a recordingFrom a clone with nothing installed, python -m humor_mcp.cli ... is the same
thing.
Check it:
python test_server.py # 61 checks over real stdio
python test_import.py # 52 checks over the importers and the pack model
python test_audio.py # 27 checks against synthesized audio (needs numpy)The suites build their own temporary corpus, so they pass on a fresh clone regardless of which packs you have.
Related MCP server: mcp-copilotcli-history
Your corpus lives in your own folder
There are two pack directories, and the second one is yours:
<checkout>/packs the packs that ship with this repo (clone only)
~/.humor-mcp/packs YOUR corpus <- imports go here by default
~/.humor-mcp/humor.db the built databaseIf you pip-installed rather than cloned, the first of those does not exist and your corpus is the only one — the wheel deliberately carries no packs, since they are ~9 MB of other people's material and two of those packs are non-commercial.
Both are read and merged at build time, so your material is searchable next to
the bundled packs without ever being mixed into the checkout. It survives
git pull, nothing in this repo touches it, and you will never resolve a merge
conflict over a joke file.
humor-mcp import-corpus --id my-sets --input jokes.txt \
--title "My tight five" --authors "Your Name" --license CC-BY-4.0
# -> ~/.humor-mcp/packs/my-sets
humor-mcp buildIf a pack of yours claims the same id as a bundled one, yours wins and the
build says so — that is how you replace a shipped pack rather than working
around it.
Every importer takes --packs-dir if you want a pack somewhere specific.
Environment
variable | what it does |
| move your corpus somewhere other than |
| read packs from exactly these directories instead ( |
| build or serve a database somewhere other than |
Together these let one checkout serve several corpora, or keep everything outside the repo entirely.
Tools
tool | what it gives you |
| full-text search, filtered by pack / kind / minimum human score |
| the lines a human actually scored highly |
| liked vs disliked with the score distribution — calibration, not just examples |
| the structural analysis of a joke: mechanism, setup/turn, technique |
| what won head-to-head and what lost |
| a paste-ready brief: exemplars + calibration + the attribution block |
| every pack: authors, licence, whether it may be shared or sold, how to cite |
| what is loaded |
Every result carries a credit object. Packs whose licence bars redistribution
are hidden unless a call passes include_restricted: true. limit is clamped to
200 — tool results land straight in the calling model's context, and a response
should not be able to flood it.
Bring your own corpus
From audio + a transcript (best)
humor-mcp reactions set.mp3 # what does it hear? look first
humor-mcp import-audio --id tight-five --audio set.mp3 --transcript set.srt \
--performer "Your Name" --title "Comedy Cellar, March" --i-own-this
humor-mcp buildThe audio supplies the reactions, the transcript supplies the words, and each
reaction is anchored to the line that earned it. Needs a timestamped transcript
(.srt, .vtt, Whisper .json); a plain .txt carries no timings and cannot
be aligned.
⚠️ The bundled detector does not work on real audience audio
Measured against 14 minutes of real stand-up with per-line ground truth from an AudioSet classifier: F1 0.24 — precision 0.15, recall 0.54. It fired 93 times for 26 real laughs.
The reason is not a bad threshold, it is that the features carry no signal here. On that recording, median spectral flatness is 0.0582 for laughter and 0.0549 for speech; 4–8 Hz modulation is 0.3255 against 0.3246. The best single threshold on any of the three features scores exactly the majority-class baseline — no threshold beats always guessing "speech".
It scored well on synthetic signals only because those were built with the property being tested: pure harmonic speech at flatness 0.004 against noise-based laughter at 0.47. Real mic'd, compressed, reverberant room audio collapses that two-order-of-magnitude gap to nothing.
Use a trained classifier instead.
MIT/ast-finetuned-audioset-10-10-0.4593has laughter and applause classes and is what produced the ground truth above. Feed its output in with--reactions, below.
It does not transcribe. Nothing downloads a model or calls a service — bring
a transcript from whatever you already use. Audio needs numpy, scipy and
soundfile; nothing else in this repo has dependencies. wav/flac/ogg/mp3 read
directly, m4a and aac need converting with ffmpeg first.
How the detector works, and where it breaks. Two features do the work: spectral flatness separates the room from the voice (a voice has pitch so its spectrum is peaky; audience noise is broadband — measured on synthetic signals, speech 0.004 against laughter 0.47 and applause 0.56), then 4–8 Hz modulation separates laughter from applause, since laughter has a syllable rhythm and applause is a wash. Loudness is used only to rule out silence: the performer is mic'd and the room is not, so real applause runs about 3 dB above speech and any volume threshold finds either everything or nothing.
It is a heuristic detector, not a trained model. Music, a noisy room, or a comic
laughing into their own mic will fool it. Run humor-mcp reactions on your file
and read the flatness values before trusting it; --flat-speech-max is the knob.
From a transcript
The usual case — you have a recording of a set, not a tidy file of jokes:
humor-mcp import-transcript --id tight-five --input set.srt \
--performer "Your Name" --title "Comedy Cellar, March" --i-own-this
humor-mcp build.srt, .vtt, Whisper .json and plain .txt all work, with or without
SPEAKER: labels. Timestamps, cue numbers, music cues and stage directions are
stripped. Pass several files at once, or --append to add another set later;
--performer is recorded per line, so one pack can hold several people.
How it finds the jokes. If the transcript keeps its reaction cues —
[laughter], (laughs), [applause] — they are used the way an audio pipeline
uses the waveform: the sentence before the cue is the line, what came before it
is the context, and the cue's strength becomes a laugh weight (a chuckle
ranks below laughter, which ranks below laughter-and-applause). That weight is
ordering, not measurement.
If there are no cues, the transcript is segmented into sentences with a rolling context window and nothing is marked as a punchline, because nothing in the text says which parts landed. You get material, not judgements, and the importer tells you so rather than quietly presenting utterances as jokes.
Transcripts of other people's performances default to not redistributable —
that is what a transcript of a copyrighted set is. --i-own-this flips it to
CC-BY-4.0, or set --license explicitly.
From a structured file
If you already have jokes in a file:
humor-mcp import-corpus --id my-sets --input jokes.txt \
--title "My tight five" --authors "Your Name" --license CC-BY-4.0
humor-mcp build.txt (blank-line separated), .csv (a text/line/joke column, optionally
context, score, note) and .jsonl all work. --authors and --license
are required. If you don't know the licence, pass --license UNKNOWN: the pack
still loads and is still searchable locally, but it is flagged everywhere it
surfaces and refused by the exporter.
Or skip the importer and write packs/<id>/pack.json plus a lines.jsonl by
hand — that is the whole contract:
{"text": "...", "context": "...", "score": 3, "note": "...", "attribution": "..."}The pack model
packs/*/pack.json is a declarative model and humor-mcp build is its compiler, so the
model is checked before anything is built — every field has a declared type with
a value predicate, and a bad manifest fails at author time with all its
problems listed at once, not one per rebuild.
$ humor-mcp build
6 problem(s) in the corpus model — nothing was built:
- packs/zz/pack.json: unknown field 'redistributible' — did you mean 'redistributable'?
- packs/zz/pack.json: 'redistributable' must be a real boolean — true/false, not "true"/"false", got 'false'
- packs/zz/pack.json: licence 'CC-BY-NC-4.O' is not a recognised identifier. Fix the
typo, or set "custom_license": true to state that you meant it.
- packs/zz/pack.json: default_hidden is true but hidden_reason is empty — record why,
or the reason lives only in someone's memoryThis matters more than ordinary input validation, because the fields are the
licence policy. "redistributable": "false" is a non-empty string, so it used to
coerce to True and the exporter would ship material whose manifest said three
separate times not to. Typed fields make that unrepresentable.
Cross-field rules are checked too: ALL-RIGHTS-RESERVED cannot be
redistributable, an -NC- licence cannot allow commercial_use, a placeholder
licence cannot be license_verified, and default_hidden requires a
hidden_reason.
The tool schemas the MCP advertises are derived from the Python function signatures rather than maintained alongside them, so a parameter cannot exist in one and not the other; an undocumented parameter raises at import.
Sharing a corpus
humor-mcp build --export ./shareCopies out only the packs whose licence actually permits it, and writes a
CREDITS.md covering exactly what went. It refuses anything non-redistributable,
anything with an unverified licence, and anything whose licence field is still
UNSET — so an unlabelled pack can't leak by accident.
What ships with this repo
Cloning gets you a working corpus, not an empty shell — but read the licence column before you build anything on top of it. Two of these packs are non-commercial, and the credit attached to every result will tell you so.
pack | lines / pairs | licence | |
| 1,248 + 463 | CC-BY-4.0 | © James Barker — engine output plus his own ratings, h2h picks and eval verdicts. Ships inside the wheel too, so |
| 2,753 | CC-BY-NC-4.0 | Context-Situated Pun dataset, Sun et al., EMNLP 2022 |
| 1,899 | CC-BY-NC-4.0 | ExPUNations, Sun et al., EMNLP 2022 |
| 4,000 pairs | CC-BY-4.0 | Reddit r/Jokes via SocialGrep. Off-rubric — hidden by default. |
| 1,500 pairs | CC-BY-4.0 | New Yorker Caption Contest, Hessel et al., ACL 2023. Text only, no cartoons. Off-rubric — hidden by default. |
The author's own local corpus also holds a standup pack of transcribed
crowd work, which is not in this repo: its licence does not permit
redistribution, so .gitignore keeps it out and humor-mcp build --export refuses it.
That asymmetry is the point of the tool — the same rules that hide it from you
hide your restricted material from everyone else.
Two gates, deliberately separate
A pack can be withheld from default results for two unrelated reasons, and collapsing them into one switch loses information:
redistributable: false— a licence question. You may not pass this material on.standupis the only such pack. Reachable withinclude_restricted: true, never exported.default_hidden: true— a relevance question. The licence is fine and the exporter will happily ship it; the corpus owner judged it unrepresentative.rjokesandnyccare 92% of all pairs and were excluded from DPO v6 for pulling the model off Mallard's voice, so they are held back rather than drowning the 463 on-rubric pairs 12:1. Reachable withinclude_hidden: true, or by naming the pack — asking for it by name is answer enough.
preference_pairs reports what it withheld in a withheld field rather than
quietly returning less than you asked for.
A note on de-duplication
The ingest merges duplicate lines within a pack, because some source tables overlap — ExPUN arrives once carrying its structural breakdown and again carrying its explanation, so every pun would otherwise appear twice with half its fields each.
The merge key is (text, attribution), never text alone. Two people can deliver
the same line, and collapsing those would credit one of them for the other's
material; 25 texts in the standup pack already appear under more than one
credit and are kept separate. Identical lines in different packs are never
merged at all, since they carry different credit by definition.
Licences
Two, as is normal for a repo that ships both software and data:
The code — the
humor_mcppackage and the test suites — is MIT, © 2026 James Barker. See LICENSE.The corpora are licensed individually, per pack, in
packs/<id>/pack.json. They are not covered by the MIT licence, two of them forbid commercial use, and some may not be redistributed at all. See NOTICE for the summary; thesourcestool prints the live state of all of them.
Owning something and licensing it are different: the authors field in each
pack records who owns the material, and the license field records what
everyone else is permitted to do with it.
Available Tools
8 toolsbreakdownA
Structural breakdowns (mechanism, setup/turn, technique) for lines matching a query — how the joke is built, not just the text.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | how many results (capped at 200) | |
| query | Yes | words to search for; empty = browse by score | |
| include_hidden | No | include packs the corpus owner marked off-rubric — their licence is fine, they were judged unrepresentative | |
| include_restricted | No | include packs whose LICENCE bars redistribution (local reference only) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden of behavioral disclosure. The description explains the conceptual output but doesn't disclose return format, pagination behavior, how results are ordered (beyond the query-empty browse mode), or whether results include the raw text. For a tool that returns structural analysis, the agent is left guessing about response shape and ordering semantics.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a tight single sentence that front-loads the core purpose. It's efficient with no wasted words. It could arguably add a touch more behavioral detail, but for its length it earns its place. Not verbose, not under-specified to the point of confusion.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema and no annotations, plus 4 parameters and conceptual complexity (categories of structural breakdown), the description is reasonably complete on the 'what' but thin on the 'how results come back' — no ordering default beyond score-browse, no return shape, no inclusion/exclusion semantics beyond what's in the parameter docs. Adequate but with clear gaps for a moderately complex tool with zero annotation coverage.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, and the description adds meaningful conceptual context beyond the schema: it frames 'query' as words to search for, connects empty-query to score-based browsing (aligning with the top_rated sibling), and distinguishes include_hidden packs as 'on-rubric but judged unrepresentative' and include_restricted as 'licence bars redistribution'. The description enriches each parameter with real behavioral meaning.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description identifies a specific verb+resource ('structural breakdowns ... for lines matching a query') and clarifies the category ('mechanism, setup/turn, technique'). It contrasts with siblings like search_humor and sources by emphasizing it's about joke construction, not text retrieval. However, it doesn't explicitly name which sibling to distinguish from, so it misses the top marks for differentiation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies the tool is a search-and-browse mechanism ('lines matching a query', 'empty = browse by score'), and the contrast with top_rated (score browsing) is hinted at. But it doesn't give explicit guidance on when to choose this over search_humor, top_rated, or sources, nor does it state exclusions or prerequisites. Usage context is implied rather than stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
corpus_statsC
What is loaded: counts by pack and by kind, size on disk.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden of behavioral disclosure. It reveals the output nature (counts by pack/kind, disk size) but doesn't disclose whether this is read-only, how it handles an empty corpus, execution cost, or what 'pack' and 'kind' refer to. For an analytical tool this is thin coverage.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single short sentence that efficiently conveys the tool's informational scope. No wasted words, though it could be slightly more explicit about the verb (report, summarize).
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter reporting tool, the description is reasonably complete about what information is returned (counts by pack and kind, disk size). However, lacking an output schema and any guidance on how 'pack' and 'kind' are defined relative to siblings like 'breakdown' or 'sources' leaves ambiguity in a toolset with several aggregation options.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 0 parameters (100% schema coverage trivially) and no params to document, the baseline of 4 applies. The description appropriately describes what the output reveals since there are no input parameters requiring explanation.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description 'What is loaded: counts by pack and by kind, size on disk' states what the tool reports on (loaded corpus, counts, size), giving reasonable purpose clarity. However, it lacks a clear verb-resource structure and doesn't differentiate from siblings like 'breakdown' or 'sources' which could plausibly return similar aggregate data.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is provided on when to use this tool versus siblings. Given siblings like 'breakdown', 'sources', and 'style_pack', the agent has no way to know which aggregation tool to choose. The description gives no context about scenarios where this tool is appropriate.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
preference_pairsA
Chosen vs rejected pairs — what won head-to-head and what lost. Off-rubric packs are withheld by default and reported in withheld.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | how many results (capped at 200) | |
| query | No | words to search for; empty = browse by score | |
| source | No | restrict to one pack id — naming a pack overrides both gates below | |
| include_hidden | No | include packs the corpus owner marked off-rubric — their licence is fine, they were judged unrepresentative | |
| include_restricted | No | include packs whose LICENCE bars redistribution (local reference only) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the burden of behavioral disclosure. It discloses that off-rubric packs are withheld by default and reported in `withheld`, and notes the licence fine/relevance caveat for hidden packs. However, it doesn't describe the output structure beyond the `withheld` field or clarify what 'restricted' packs mean for returned data.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single tight sentence that conveys core purpose plus an important default behavior (withholding off-rubric packs). No wasted words. The parameter descriptions in schema carry the detail. Could arguably be slightly richer but earns a high score for efficiency.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool has 5 parameters, no enums, no output schema, and no annotations - so the description needs to cover a fair amount of behavioral ground. It explains the default withholding behavior and the `withheld` reporting field. The sibling context (search_humor, top_rated, breakdown) suggests this is a specialized comparison tool, and the description sets it apart reasonably well given the constraints.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3. The description adds value by explaining that naming a source pack 'overrides both gates below,' and clarifies the 'empty query = browse by score' behavior. It also explains the hidden/restricted filters in plain language, going beyond the raw schema descriptions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool returns chosen vs rejected pairs from head-to-head comparisons, with a mechanism for off-rubric packs being withheld. It uses a specific verb and resource, and the 'what won head-to-head and what lost' phrasing distinguishes it from siblings like top_rated or search_humor. However it doesn't explicitly name sibling tools for differentiation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies browsing by score when no query, and mentions off-rubric packs are 'withheld by default.' The filter semantics of the parameters (source overriding gates, include_hidden, include_restricted) are present in the schema descriptions. It gives clear context for when to use, though it doesn't name explicit alternatives among the siblings.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
search_humorA
Full-text search the humor corpus. Returns lines with their credit attached. Filter by source pack, kind or minimum human score.
| Name | Required | Description | Default |
|---|---|---|---|
| kind | No | joke / pun / candidate / slate_winner / eval / utterance / word | |
| limit | No | how many results (capped at 200) | |
| query | No | words to search for; empty = browse by score | |
| source | No | restrict to one pack id — naming a pack overrides both gates below | |
| min_score | No | only lines a human scored at least this highly | |
| whole_lines | No | drop transcript offcuts that start or stop mid-sentence; heuristic, and strict about a final full stop | |
| include_hidden | No | include packs the corpus owner marked off-rubric — their licence is fine, they were judged unrepresentative | |
| include_restricted | No | include packs whose LICENCE bars redistribution (local reference only) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
There are no annotations to rely on, so the description carries the transparency burden. It discloses return content (lines with credit) and a heuristic behavior on whole_lines. However, it doesn't explain pagination behavior, how the sort/ordering works, or what happens with special inputs — moderate disclosure but with gaps for a read/search tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences that front-load the core purpose and immediately list the key filters. Zero waste — every clause carries meaning. The return description ('lines with their credit attached') is folded neatly into the first sentence.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
This is a moderately complex tool with 8 parameters but no output schema. The description covers the search-and-filter story well but doesn't explain the default ordering when query is empty, the interaction between source and the gating flags, or how results are ordered/ranked. For a search tool with no output schema, slightly more about result shape and ordering would improve completeness, though the parameters are well-covered by the schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the schema documents all 8 parameters. The description adds semantic value by grouping parameters into meaningful categories (source pack, kind, minimum human score), which helps the agent reason about which filters apply together. It also labels whole_lines as 'heuristic' and 'strict about a final full stop,' adding behavioral nuance beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states it performs 'full-text search the humor corpus' and returns 'lines with their credit attached.' It also names the filter dimensions (source pack, kind, minimum human score), distinguishing it from siblings like top_rated, corpus_stats, and sources which serve different purposes.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies search/filtering use cases via the filtering verbs, but doesn't explicitly state when to choose this over alternatives like top_rated, taste_profile, or breakdown. There's no when-not-to-use guidance, though the 'browse by score' hint in the query param schema partially indicates an alternative to top_rated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
sourcesA
Every corpus pack loaded: who wrote it, under what license, whether it may be redistributed or used commercially, and how to cite it.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided at all (no readOnlyHint, no destructiveHint), the description carries the full burden of behavioral disclosure. This is clearly a read-only informational tool with zero parameters, so the risk profile is low and inherently safe. The description accurately conveys it's a passive reporting tool, which is sufficient given the tool's trivial side-effect surface. It doesn't hide any destructive or mutating behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, dense sentence that packs in exactly what's needed: coverage scope ('every corpus pack loaded') and the four dimensions reported (authorship, license, redistribution/commercial use, citation). Zero wasted words, appropriately sized for a parameterless reporting tool.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter tool with no output schema, the description is reasonably complete. It tells the agent the scope ('every corpus pack loaded') and the content dimensions. It doesn't describe the return format or structure, but with no output schema and a simple informational purpose, the description covers the essential usage context well enough.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters, and the schema description coverage is 100% (trivially, since there are no properties to document). The no-parameter nature is consistent with the description's framing as a holistic report of all loaded corpus packs. With no parameters to explain, there's nothing missing in the description.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states what the tool does: it reports information about every loaded corpus pack covering authorship, license, redistribution/commercial use rights, and citation format. The verb is implied (it 'reports' or 'shows' this information), and it names a specific resource (corpus packs). It's distinguishable from siblings like corpus_stats since it focuses on provenance/legal metadata rather than usage statistics, though it doesn't explicitly differentiate itself.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The context is implied: one would use this to learn about licenses, citation, and redistribution rights for loaded corpus packs. However, there's no explicit guidance on when to use this vs. alternatives, no stated exclusions, and no mention of how this differs from corpus_stats or breakdown. The usage context is understandable but not explicitly framed.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
style_packC
A compact, paste-ready style brief: exemplars, liked/disliked calibration, and the attribution block that must travel with any output.
| Name | Required | Description | Default |
|---|---|---|---|
| n | No | how many examples per side (capped at 200) | |
| topic | No | focus the exemplars on a subject; omitted = highest-rated overall | |
| include_hidden | No | include packs the corpus owner marked off-rubric — their licence is fine, they were judged unrepresentative | |
| include_restricted | No | include packs whose LICENCE bars redistribution (local reference only) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden of behavioral disclosure. It mentions the tool includes packs that may be 'off-rubric' or 'restricted' based on parameter docs, but the description itself doesn't disclose what gets generated, whether any mutation occurs, license/reuse implications, or output format. The 'attribution block that must travel' hints at a licensing expectation but doesn't explain its mechanics or consequences.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single sentence and compact, which is good. However, it packs in several undefined jargon terms ('exemplars', 'liked/disliked calibration', 'attribution block') that force the reader to parse marketing-style language for meaning. It's concise but not maximally clear; front-loading a plain-verb statement of the operation would serve the agent better.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-style tool with full schema coverage and no output schema, the description is adequate but leaves gaps. It doesn't clarify the return format (is it text, structured data?), how restricted packs are handled in output, or why/when to enable the hidden and restricted flags. Given four boolean/integer parameters influencing output content and licensing-sensitive behavior, the description could meaningfully expand on when those flags matter, which it does not.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so each parameter (n, topic, include_hidden, include_restricted) is documented in the schema itself. The description adds no parameter-specific meaning beyond what the schema provides; it references 'exemplars' and 'calibration' generically but doesn't clarify how 'n' maps to sides, or what 'liked/disliked calibration' means operationally. Per baseline for full schema coverage, a 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description reads like a marketing blurb ('compact, paste-ready style brief') using jargon ('exemplars', 'liked/disliked calibration', 'attribution block') that isn't precisely defined. It conveys a general sense that this produces a style summary with examples, but doesn't clearly state the tool's verb-resource operation (e.g., 'Generate a style brief'). The jargon and stylistic phrasing obscure rather than clarify the core purpose, distinguishing it only weakly from siblings like taste_profile and preference_pairs.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is minimal guidance on when to use this tool. The description references an 'attribution block that must travel with any output,' implying a downstream requirement, but doesn't explicitly state when to prefer this over search_humor, taste_profile, preference_pairs, or breakdown. No exclusions or alternatives are mentioned, and the context of 'style' vs. sibling concepts like 'taste' or 'humor' is left implicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
taste_profileB
What this corpus's human rater actually liked versus disliked, with the score distribution. Use this to calibrate register before writing.
| Name | Required | Description | Default |
|---|---|---|---|
| n | No | how many examples per side (capped at 200) | |
| kind | No | joke / pun / candidate / slate_winner / eval / utterance / word | |
| rater | No | restrict to one rater's judgements | |
| include_hidden | No | include packs the corpus owner marked off-rubric — their licence is fine, they were judged unrepresentative | |
| include_restricted | No | include packs whose LICENCE bars redistribution (local reference only) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden of behavioral disclosure. The description doesn't disclose how results are returned, the meaning of the 'liked vs disliked' structure, sorting, pagination, or what happens with edge cases like an rater with no judgements. For a tool without annotations, this is a thin behavioral disclosure.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences that front-load the primary purpose and then give actionable guidance on when to use it. No wasted words or redundant detail. Could arguably add a note on return format but remains efficiently concise.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool has 5 optional parameters and no output schema, so the description carries the burden for return semantics. The description explains the conceptual purpose well but leaves the actual output structure (what a 'score distribution' looks like, how liked vs disliked items are presented) unspecified. Adequate but with a notable gap given the absence of an output schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% and all five parameters have inline descriptions in the schema, so the baseline is 3. The description adds no parameter-level meaning beyond what the schema already documents; the schema's 'kind' and 'rater' descriptions are serviceable. The description doesn't compensate further but isn't required to at this coverage level.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool shows what the corpus rater liked vs disliked with a score distribution, using a specific verb ('calibrate') plus resource ('register'). It distinguishes itself from siblings by its focus on the rater's taste profile versus other tools like search_humor or breakdown. However, it doesn't explicitly compare against sibling distinctions beyond the implicit calibration framing.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly frames this as a calibration tool before writing ('Use this to calibrate register before writing'), which gives clear context for when to invoke it. It doesn't name specific alternatives or exclusions, but the calibration-when-before-writing guidance plus sibling names like top_rated and breakdown provide implied differentiation.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
top_ratedB
The highest human-rated lines in the corpus. Single-word lexicon entries use a different scale and are excluded unless kind='word'.
| Name | Required | Description | Default |
|---|---|---|---|
| kind | No | joke / pun / candidate / slate_winner / eval / utterance / word | |
| limit | No | how many results (capped at 200) | |
| source | No | restrict to one pack id — naming a pack overrides both gates below | |
| min_score | No | only lines a human scored at least this highly | |
| include_hidden | No | include packs the corpus owner marked off-rubric — their licence is fine, they were judged unrepresentative | |
| include_restricted | No | include packs whose LICENCE bars redistribution (local reference only) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden of behavioral disclosure. It does disclose the value scale nuance (single-word lexicon entries use a different scale and are excluded unless kind='word'), which is useful. However, it doesn't describe default behaviors like the min_score=2 gate, the include_hidden/include_restricted defaults, the limit cap, or the source override behavior — all of which meaningfully affect what gets returned.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, focused sentence that states purpose and the key exception in one breath. Efficient and front-loaded with the essential fact. Minor opportunity: it could note the default gate values, but the sentence is well-constructed and wastes no words.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With 6 parameters, no output schema, and no annotations, the description leaves gaps on default behavior (min_score filter, hidden/restricted exclusions, limit cap, source override priority). The kind='word' exception is the one piece of nuance disclosed. For a listing tool with array of defaults, the description is adequate but doesn't fully orient the agent on what makes results appear or disappear from the listing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema documents all 6 parameters well. The description adds marginal value by explaining the special-case exclusion rule tied to kind='word', which supplements the schema's bare enum list. Since coverage is high, the description doesn't need to re-document parameters; credit for the one meaningful extra it does provide.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a clear purpose: returns the highest human-rated lines in the corpus, with a noteworthy exclusion for single-word lexicon entries unless kind='word'. It identifies the resource (highest-rated lines) and a distinguishing condition. However, it doesn't explicitly differentiate from siblings like search_humor or preference_pairs, relying instead on the inherent uniqueness of a 'top rated' listing.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage context (fetching top-rated content) but provides no explicit guidance on when to use this vs alternatives like search_humor, taste_profile, or preference_pairs. The kind parameter hints at different categories but the description doesn't articulate when this tool is the right choice over sibling tools.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
8 tool updates
v0.1.2- First observed
breakdown - First observed
corpus_stats - First observed
preference_pairs - First observed
search_humor - First observed
sources - First observed
style_pack - First observed
taste_profile - First observed
top_rated
TDQS
Scored across 8 tools
Most tools are clearly distinct: search_humor finds text, top_rated ranks by score, breakdown explains structure, preference_pairs shows pairwise comparisons, and sources/corpus_stats describe corpus metadata. However, taste_profile, style_pack, and preference_pairs have some overlap in that they all serve rater/calibration purposes and could cause misselection when an agent wants calibration guidance.
Names use a consistent snake_case convention throughout, mostly combining a noun subject with a descriptive suffix (search_humor, top_rated, taste_profile, corpus_stats, style_pack). Most are noun-headed which is descriptive though not strictly verb_noun; minor deviation like 'breakdown' and 'sources' being bare nouns breaks the pattern slightly.
8 tools is right in the sweet spot for a specialized domain server. Each tool addresses a distinct need: searching, ranking, calibration, structural analysis, pairwise comparison, attribution, and corpus statistics. Nothing feels redundant or frivolous.
The surface covers the main workflow: search for humor, understand quality via ratings, get structural breakdowns, see pairwise preferences, and obtain attribution/licensing. Minor gaps could include functionality to inspect individual lines by ID or to generate/write new humor, but for an analysis-oriented humor corpus, the coverage is strong.
Maintenance
Related MCP Connectors
Multi-engine scholarly research server for search, traversal, full text, and reading lists.
Resolve, search and verify legal citations against the official sources, with provenance.
Search biomedical papers, inspect publication records, and traverse citation or semantic graphs.
Ingest and search LogsLoom logs from coding agents.
Related MCP Servers
- FlicenseAqualityDmaintenanceEnables AI agents to query a local knowledge graph built from document collections using hybrid search (BM25 + vector fusion) and entity-relationship extraction. Supports privacy-first, offline operation with tools for semantic search, entity graph exploration, and corpus statistics.3-
- AlicenseAqualityCmaintenanceEnables searching and analyzing GitHub Copilot's conversation history stored locally, providing tools for full-text search, session listing, statistics, and file-based retrieval.64MIT
- AlicenseNot gradedqualityDmaintenanceEnables querying and exploring a bundled knowledge base through tools for manifest, table of contents, node retrieval, and full-text search.5 npm1MIT
- AlicenseBqualityAmaintenanceProvides read-only hybrid RAG search and discovery over a local-first AI knowledge corpus, enabling semantic and keyword search, browse, digest, and status tools.4PolyForm Noncommercial 1.0.0