jev-kit
Integrates with Vercel AI Gateway to use Vercel AI credits for Jev access and to transcribe audio recordings via OpenAI Whisper for Laugh Track analysis.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@jev-kitrun my best man speech through Laugh Track and flag the weakest lines"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
β‘ Jev Kit
Your LLM thinks. Jev decides.
A growing kit of tools that plug Jev, TypeSafe AI's System One decision model, into Claude, ChatGPT, your browser and your terminal. Jev answers typed questions (yes/no, pick one, score on a scale) with probabilities in well under a second for fractions of a cent, so you can afford to ask thousands of small questions that would be slow or expensive to send to an LLM. Your LLM does the writing and explaining; Jev does the judging.
Published on npm as @leighefford/jev-kit. It's unrelated to the separate jev-kit package.
π Laugh Track: rehearse in front of a room
Paste a speech, stand-up set or pitch, or bring a recording of yourself giving it. Two hundred simulated people react to every line: laughs, groans, silence, or being moved. Jokes are judged on laughs and heartfelt lines on warmth, and you get the three lines to rework first. With a recording, it's transcribed for you, your pace and pauses are part of the analysis, and you can play it back alongside the room.
npx -y @leighefford/jev-kit laugh-trackYour browser opens, the seats fill up, and each line plays to the room in turn. When it finishes you get a score for every line, averages for the jokes and the sincere lines, and the three lines to rework first:
It also works from Claude or ChatGPT ("run my best man speech through Laugh Track and rewrite the weakest lines") and from the terminal. The 13-line wedding speech above cost $0.00097 in Jev calls.
Related MCP server: Jev MCP Server
More tools in the kit
Each tool works from Claude or ChatGPT (MCP), the terminal and your own code. New tools are added over time; see Adding a tool.
Tool | What it does | Why it needs Jev |
π§ Steering Wheel | Judge every sentence of a draft: off-goal, filler, repetition, invented facts. Rewrite only what's flagged. | A judgment per sentence is cheap enough to run on everything you write |
π Semantic Bisect |
| Jev reads each commit's output and calls it good or bad |
π¦ Anything-vs-Anything Physics | What happens when a rubber duck meets a laser? A common-sense rulebook for sandbox games and interactive stories. | 13 typed judgments per collision, every pair in parallel |
π₯ Collision Engine | Score every pair of your notes for "surprising and useful if combined". | 5,000 notes is 12.5 million pairs; only a model this cheap makes that thinkable |
Plus jev_ask for asking Jev your own typed questions.
Quick start
Pick the way you want to use it. Every way runs on your own Jev key, on your own machine.
If you want to⦠| Run | You get |
Try it in your browser |
| Laugh Track: paste a speech and watch 200 seats react, line by line |
Use it from Claude or ChatGPT | "Run my speech through Laugh Track" in chat | |
Use it in the terminal |
| Results printed line by line |
Build it into your own app |
|
|
1. Get a key
Either works:
Vercel AI Gateway (uses your Vercel AI credits): create a key under AI Gateway β API Keys.
TypeSafe direct: get a key at console.typesafe.ai/keys.
The browser page asks for the key the first time, checks it with Jev, and saves it to ~/.config/jev-kit/config.json (readable only by you). Elsewhere, set AI_GATEWAY_API_KEY or TYPESAFE_API_KEY, or on a Mac keep it in the Keychain:
security add-generic-password -a "$USER" -s ai-gateway -w2. Try it in your browser
npx -y @leighefford/jev-kit laugh-trackYour browser opens at http://127.0.0.1:4747. Paste a speech (or pick an example), say where you're performing, and press Play it to the room. Each line plays in turn while the seats light up in the crowd's reaction colours; at the end you get the three lines to rework first.
Switch to Audio to drop in a recording (MP3, M4A, WAV or WebM, up to 24 MB) or record straight from your microphone. It's transcribed through Vercel AI Gateway (openai/whisper-1, about $0.006 per minute), split into lines, and the transcript appears for you to tidy up. Press play for the line-by-line analysis; pace and pauses from the recording are part of it. Turn on Play my recording while the room reacts (off by default) to hear your take in sync with the room, with a waveform coloured by each line's reaction. Recordings need a Vercel AI Gateway key; a TypeSafe key covers Jev only.
The page only runs on your computer and only accepts requests from itself, so other websites can't use your key through it.
Add it to your AI
Claude Code
claude mcp add jev-kit -e AI_GATEWAY_API_KEY=your_key -- npx -y @leighefford/jev-kit(If the key is in your Keychain, leave out -e AI_GATEWAY_API_KEY=your_key.)
Claude Desktop: Settings β Developer β Edit Config, then add:
{
"mcpServers": {
"jev-kit": {
"command": "npx",
"args": ["-y", "@leighefford/jev-kit"],
"env": { "AI_GATEWAY_API_KEY": "your_key" }
}
}
}Cursor, Windsurf, Zed and other MCP clients use the same command: npx -y @leighefford/jev-kit.
ChatGPT connects to remote MCP servers over HTTPS, so run Jev Kit in HTTP mode and expose it with a tunnel:
npx -y @leighefford/jev-kit serve --http --port 3333It prints a URL with a random secret path, e.g. http://127.0.0.1:3333/mcp/3f9cβ¦. Put a tunnel in front of it (for example cloudflared tunnel --url http://127.0.0.1:3333), then in ChatGPT enable developer mode and add a connector with the tunnel URL plus the secret path. Anyone with that full URL can spend your Jev credits, so keep it private and stop the server when you're done.
Ask for something
"Run my wedding speech through Laugh Track and rewrite the three weakest lines."
"Run ~/Desktop/speech-take-2.m4a through Laugh Track as a best man speech."
"Steer-check this launch post against these release notes, then fix only the flagged sentences."
"Semantic-bisect this repo:
npm run build && node dist/cli.js --helpstarted printing the wrong version somewhere after v2.3.0.""Physics sandbox: a toddler, a vending machine, a bag of flour and a Roomba, in a supermarket."
"Run Collision Engine over ~/notes with the lens 'weekend projects' and turn the top five into proposals."
Terminal use
Everything also runs without an AI client:
npx -y @leighefford/jev-kit doctor # check your key
npx -y @leighefford/jev-kit laugh examples/set.txt
npx -y @leighefford/jev-kit steer examples/pitch.txt --goal "Convince a developer to install Jev Kit" --sources examples/pitch-sources.txt
npx -y @leighefford/jev-kit bisect --good v1.0 --cmd "node app.js" --symptom "the error message is rude"
npx -y @leighefford/jev-kit physics "a rubber duck" "a laser beam" --how "is hit by"
npx -y @leighefford/jev-kit collide examples/notes --top 5Add --json for raw results. Run jev-kit --help for every option.
What it looks like
$ jev-kit laugh examples/best-man-speech.txt --occasion "best man speech at a wedding reception"
π ββββββΒ·Β·Β·Β· 62 joke Tom proposed on a beach in Cornwall. He had the ring, the speech and the perfect sunset. He did not have a plan for the seagull.
π βββββββΒ·Β·Β· 65 joke I won't tell you what the seagull took. I'll just say that Sophie said yes to a man holding half a pasty.
π¬ βΒ·Β·Β·Β·Β·Β·Β·Β·Β· 5 sincere Marriage is an institution that has existed across many cultures for thousands of years.
π βββββββΒ·Β·Β· 68 joke Tom asked me to keep this speech clean, so I've removed the story about the stag do, the story about the other stag do, and most of the verbs.
π₯Ή βββββββΒ·Β·Β· 67 sincere In all seriousness, Tom is the most loyal friend I have. When my dad was ill, he drove four hours every weekend just to sit with me in a hospital car park and eat bad sandwiches.
π βββΒ·Β·Β·Β·Β·Β·Β· 33 sincere So Sophie, you're not just getting a husband. You're getting a man who turns up.
β¦
Landed: 49/100 from 200 people Β· jokes 59 (laughs) Β· sincere lines 35 (warmth)
Rework first:
β’ Marriage is an institution that has existed across many cultures for thousands of years.
β’ So Sophie, you're not just getting a husband. You're getting a man who turns up.
β’ Anyway, I have also prepared some statistics about divorce rates in the UK.
13 Jev calls, 23,102 input tokens, ~$0.00097$ jev-kit steer examples/pitch.txt --goal "Convince a developer to install Jev Kit"
β Jev Kit lets Claude and ChatGPT make thousands of small judgments in milliseconds.
β It plugs into any MCP client with one command.
β Our tools are loved by 87% of Fortune 500 companies. β likely invented fact
β Jev Kit is a kit, and kits are sets of things that go together. β off-goal, filler
β Each tool call typically costs a fraction of a cent.
β You can install it in under a minute and start testing a speech on a virtual audience straight away.$ jev-kit bisect --good <first> --cmd "node app.js" --symptom "the error message is rude, sarcastic or blames the user"
bad 0.85 01f50792d5 tweak copy: Save failed. Try harder.
good 0.08 32b1b5001f tweak copy: Couldn't save your file. Pleas
good 0.33 135f6015d6 tweak copy: Save failed: disk quota. Free
bad 0.97 f841a2a598 tweak copy: Save failed. Obviously. What d
First bad commit: f841a2a598 tweak copy: Save failed. Obviously. What d
4 commits checked out of 6.What it costs
Observed through Vercel AI Gateway while building this (27 September 2026). Your numbers will vary with text length.
Run | Jev calls | Cost |
Laugh Track, 13-line wedding speech | 13 | $0.00097 |
Steering Wheel, 6-sentence pitch | 6 | $0.00013 |
Semantic Bisect, 6 commits | 4 | $0.00006 |
Physics, one collision | 1 | $0.00003 |
Collision Engine, 12 notes (66 pairs) | 66 | $0.00116 |
Collision Engine prints an estimate first and asks for --yes above 500 pairs. --max-pairs (default 2,000) caps spend; above that, pairs are sampled.
How each tool works
Laugh Track (with a recording, first transcribed with word timestamps, split into lines, and given a one-sentence delivery note such as "spoken at 150 words per minute, after a 1.2s pause") sends each line, with the three lines before it, to Jev once. It asks eight audience groups (office worker, student, tired parent, engineer, three-pints punter, retiree, comedy nerd, tourist) to pick a reaction (laugh, chuckle, moved, groan, gasp, applause or silence), then weights them into 200 people. The same call asks whether the line is meant as a joke or as sincere. Each line gets a laugh score (big laughs plus half the chuckles and applause), a warmth score (moved plus half the applause and a quarter of the chuckles) and a landed score that blends the two by how sincere the line is meant to be. "Rework first" ranks by landed score, so a toast isn't marked weak for not being funny. You can pass your own audience.
Steering Wheel asks four questions per sentence: is it on goal, how much it adds, whether it repeats earlier text, and whether it contains an invented specific (a statistic, number, named customer or quote). Pass
sourcesfor the best results; without them it only flags invented specifics it is very confident about.steerStream()in the library steers a live generator sentence by sentence.Semantic Bisect creates a temporary
git worktree(your checkout and stash are never touched), runs your command at each probed commit and asks Jev whether the output shows the symptom. It checks both ends first and warns if they don't look right. Commits that fail for unrelated reasons (build errors) are skipped.Physics asks one choice (overall outcome), one score (intensity) and eleven yes/no effects (breaks, fire, explosion, flees, stuck, wet, noise and so on). Wording matters:
--how "is hit by"gets better answers than the default "collides with".Collision Engine scores each pair on how promising the combination is, whether it's obvious (same topic) and whether it suggests something concrete, then ranks them. Your LLM turns the top pairs into ideas.
Experimental: a local model instead of Jev
Jev Kit can talk to any server that speaks the same POST /v1/systemone protocol as Jev, including self-hosted models such as Laya (an independent, Apache-licensed project, not made or endorsed by TypeSafe). Point it at the server and no Jev key is needed:
JEV_BASE_URL=http://127.0.0.1:8000 npx -y @leighefford/jev-kit laugh-trackAdd JEV_API_KEY if your server requires one, and JEV_MODEL to request a specific model. Audio transcription still needs a Vercel AI Gateway key.
Before relying on it:
We haven't measured Laugh Track's quality on a local model yet. Laya's own documentation says its base checkpoints are meant to be fine-tuned, not used as zero-shot decision engines, so expect different (likely weaker) judgments than Jev.
Confidence numbers aren't comparable between models, and very long option lists are handled differently.
Lock the server to your machine. Laya's server listens on all network interfaces by default; set
LAYA_HOST=127.0.0.1.
Use it as a library
import { Jev, noul, choice } from "@leighefford/jev-kit";
import { laughTrack } from "@leighefford/jev-kit/tools/laugh-track";
const jev = new Jev(); // reads AI_GATEWAY_API_KEY, TYPESAFE_API_KEY or the Keychain
const answers = await jev.ask("Refund request: I was charged twice.", {
refund: noul("The customer wants money back"),
team: choice("Which team handles this?", { billing: "Charges", tech: "Bugs" }),
});
const set = await laughTrack(jev, { script: "β¦", occasion: "best man speech" });
console.log(jev.usageSummary());Development
npm install
npm test # offline, Jev is mocked
npm run smoke # live: spawns the MCP server and makes one real call (needs a key)Limits and honesty
Scores are one model's quick judgment, not ground truth. A laugh score of 46 means "Jev expects a lukewarm room", not a measured audience.
Laugh Track judges words. With a recording it also reads pace and pauses (from word timestamps), but not tone of voice, emphasis or facial expression.
Recordings are split into lines from sentence ends, pauses and long comma runs. Check the transcript before playing: a bad split changes the scores.
Semantic Bisect runs the shell command you give it at old commits. Only use commands you'd be happy to run yourself.
Keys never leave your machine except to call Jev. The only thing stored is the key you save through the browser page, in
~/.config/jev-kit/config.json(owner-only). Nothing is logged.
Credits
Built with Jev by TypeSafe AI and the Model Context Protocol. Not affiliated with TypeSafe AI or Vercel.
MIT licensed. See LICENSE.
Available Tools
7 toolscollision_engineCollision EngineARead-only
Find surprising, useful combinations in a pile of notes by scoring every pair with Jev. Give either a folder/file path or a list of notes. Returns the top collisions; develop each into a concrete idea. maxPairs caps cost (2,000 pairs is typically well under $0.05).
| Name | Required | Description | Default |
|---|---|---|---|
| top | No | ||
| lens | No | Optional focus, e.g. "startup ideas", "short stories" | |
| path | No | Folder of .md/.txt notes, or one file with notes separated by blank lines | |
| notes | No | Notes as plain strings (alternative to path) | |
| maxPairs | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations cover the safety profile (readOnlyHint, openWorldHint), and the description adds genuinely new context: it discloses the return shape ('returns the top collisions') and a concrete cost characteristic for maxPairs (2,000 pairs under $0.05). It does not discuss defaults or failure modes, so not a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight, front-loaded sentences: purpose first, then input forms, then return/cost. Every sentence contributes, with only mild redundancy in 'develop each into a concrete idea'.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema and zero required parameters, the description shoulders return-value explanation reasonably ('returns the top collisions') and covers cost via maxPairs. It remains silent on the 'top' default and what 'lens' does to output, so a small completeness gap remains.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 60%, so the schema already documents path/notes/items well. The description adds real meaning for maxPairs by tying it to cost, and clarifies path vs notes as alternatives, but it omits any mention of 'top' or 'lens', leaving those parameters to the schema alone.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('find ... combinations in a pile of notes') plus the mechanism ('scoring every pair with Jev'), which is distinctive enough to separate it from generic search tools. It does not explicitly name or contrast any sibling (jev_ask, semantic_bisect, etc.), so 4 rather than 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It tells the agent the accepted inputs ('either a folder/file path or a list of notes') and a follow-up step ('develop each into a concrete idea'), which implies intended usage. However, there is no explicit when-to-use vs when-not, and no routing to alternative sibling tools for related tasks.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
jev_askAsk JevARead-only
Low-level access: evaluate a state against your own typed questions. Question types: {"type":"noul","instructions":"..."} (probability 0-1), {"type":"choice","instructions":"...","criteria":{"key":"description"}} and {"type":"score","instructions":"...","criteria":["level 0","level 1"]}.
| Name | Required | Description | Default |
|---|---|---|---|
| state | Yes | The text or situation to judge | |
| questions | Yes | Map of question name to question object |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and openWorldHint=true, covering the safety profile. The description adds the phrase 'low-level access' but does not disclose side effects, rate limits, or result format beyond what annotations provide.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and front-loads the core action before listing question types. It is dense but every sentence earns its place, with only minor jargon in 'Low-level access.'
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given nested object parameters and no output schema, the description thoroughly documents input shapes but says nothing about what the tool returns or how evaluation results are formatted. That omission leaves an agent uncertain about the response format.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Although schema description coverage is 100%, the input schema only describes the questions parameter as a generic object map with additionalProperties: {}. The description compensates by defining three explicit question type structures (noul, choice, score) with their fields and value types, adding substantial meaning not present in the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: 'evaluate a state against your own typed questions.' It distinguishes the tool as a low-level evaluation primitive, but does not differentiate it from siblings, which appear unrelated.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is provided on when to use this tool versus alternatives. There are no exclusions or prerequisites mentioned, leaving usage entirely to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
laugh_trackLaugh TrackARead-only
Test a stand-up set, speech or pitch on a virtual audience of 200 people. Returns each line's crowd reaction (laugh, chuckle, moved, groan, gasp, applause or silence), whether the line is a joke or sincere, a laugh score and a warmth score, and a 0-100 landed score that judges jokes on laughs and heartfelt lines on warmth. Weakest/strongest are ranked by landed score; use them to suggest rewrites. Pass audioPath to judge a recording: it is transcribed, split into timed lines, and pace and pauses are taken into account.
| Name | Required | Description | Default |
|---|---|---|---|
| script | No | The full text, ideally one joke or beat per line | |
| audience | No | Custom audience groups; defaults to a mixed 200-person comedy club crowd | |
| occasion | No | e.g. "stand-up club set", "best man speech", "investor pitch" | |
| audioPath | No | Absolute path to a recording (mp3, m4a, wav, webm) to transcribe instead of script; needs a Vercel AI Gateway key |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Goes well past the annotations (readOnlyHint/openWorldHint) by disclosing the transcription-to-timed-lines pipeline, that pace and pauses affect scoring, the exact reaction taxonomy, and the external dependency that audioPath 'needs a Vercel AI Gateway key'. That is real operational context an agent would otherwise have to discover by failing.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with purpose and return shape, then the audioPath branch, then the rewrite guidance. Dense and mostly waste-free, though the reaction list and dual-scoring explanation make it longer than a single-breath definition.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema, so the description carries the full burden of explaining returns, and it does: per-line reaction, joke-vs-sincere classification, laugh/warmth scores, and the 0-100 landed score with its ranking. An agent has everything needed to call it and interpret the result.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the schema carries most parameter meaning. The description still adds value by explaining what audioPath actually triggers (transcription, line splitting, pace/pause weighting) and by restating the 200-person default crowd, which clarifies intent rather than just syntax.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a concrete verb and resource ('Test a stand-up set, speech or pitch on a virtual audience of 200 people') and immediately names what comes back. Nothing in the sibling list (physics_*, steer_check, semantic_bisect, jev_ask) overlaps, so the agent can route to it unambiguously.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear context for the two input paths: pass script for text, pass audioPath to judge a recording, and it tells the agent to use weakest/strongest to drive rewrites. It never states when not to use the tool or names an alternative, so it stops short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
physics_interactAnything-vs-Anything PhysicsARead-only
Resolve what happens when any two things meet (a rubber duck and a laser, a toddler and a vending machine) using common sense instead of hand-written rules. Returns the outcome, intensity 0-10 and likely effects with probabilities. Narrate the result vividly.
| Name | Required | Description | Default |
|---|---|---|---|
| a | Yes | First thing | |
| b | Yes | Second thing | |
| how | No | How they meet, default "collides with" | |
| setting | No | Where, default "an ordinary room" |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations cover the safety profile (readOnlyHint, openWorldHint), and the description adds real value beyond them by disclosing the return shape: outcome, an intensity 0-10 score, and effects with probabilities. It does not mention auth, limits, or determinism, but for a read-only generative tool the disclosed output structure is the main behavioral gap closed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight sentences, front-loaded with the core action before the return description and the usage instruction. The illustrative parenthetical is short and earns its place; nothing is redundant.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description carries the burden of describing returns and does so (outcome, intensity, probabilistic effects). Optional 'how'/'setting' parameters are schema-documented, so nothing essential is missing; auth/limits remain unaddressed but are minor for a read-only tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% with each of the four parameters described, so the baseline of 3 applies. The description's examples ('a rubber duck and a laser') illustrate the flavor of the a/b pairings but add no syntax or formatting guidance beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific verb ('resolve') and a clear resource (the outcome of two things meeting) with concrete examples that make the intent unambiguous. It distinguishes itself from rule-based siblings via 'using common sense instead of hand-written rules,' though it never names collision_engine or physics_sandbox explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is only implied: the 'common sense instead of hand-written rules' phrase hints this is for ad-hoc, fantastical pairings rather than rule-based collision, and 'narrate the result vividly' tells the agent what to do with the output. No explicit when-to-use/when-not or named alternatives are given, and no sibling (e.g., physics_sandbox, collision_engine) is distinguished.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
physics_sandboxPhysics SandboxCRead-only
Drop 2-8 things into one setting; every pair interacts in parallel and events come back sorted by intensity. Use it to narrate a chain of chaos.
| Name | Required | Description | Default |
|---|---|---|---|
| things | Yes | ||
| setting | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare readOnlyHint=true and openWorldHint=true, covering the safety profile. The description adds real behavioral detail beyond that: pairwise parallel interaction and results sorted by intensity. It still doesn't explain what an 'event' or 'intensity' looks like in the response.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences with the mechanism front-loaded before the usage hint. Slightly poetic phrasing ('chain of chaos') costs a little precision but nothing is wasted.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description should clarify return semantics, and 'events come back sorted by intensity' is the only hint. For a 2-parameter simulation tool, the input semantics and output shape are both under-specified.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must carry the burden. It does explain the '2-8 things' bound (matching minItems/maxItems) and the single 'setting' string, but never clarifies what kinds of values 'things' should be or how 'setting' is used, leaving the core input ambiguous.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description conveys a simulation: drop 2-8 objects into a shared setting where all pairs interact. It is evocative but doesn't specify what a 'thing' is or how this differs from siblings like physics_interact or collision_engine, so an agent can't cleanly discriminate between them.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
'Use it to narrate a chain of chaos' implies a creative-narration context but names no alternatives and gives no when-not conditions. With physics_interact and collision_engine as siblings, routing guidance is absent.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
semantic_bisectSemantic BisectA
Find the commit that introduced a problem a normal test can't detect. Runs command at commits between good and bad in a temporary git worktree (the user's checkout is untouched) and asks Jev whether each output shows symptom. Runs the user's shell command, so confirm it with the user first.
| Name | Required | Description | Default |
|---|---|---|---|
| bad | No | Commit where it is present (default HEAD) | |
| good | Yes | Commit/tag/branch where the problem is absent | |
| repo | Yes | Absolute path to the git repository | |
| command | Yes | Shell command whose output reveals the problem, e.g. 'npm run build && node cli.js --help' | |
| symptom | Yes | What the problem looks like in the output, in plain English | |
| timeoutSec | No | Per-commit timeout, default 120 |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Goes well beyond the annotations: it discloses that execution happens in a temporary git worktree and that 'the user's checkout is untouched,' which explains why destructiveHint=false holds despite running an arbitrary shell command. It also warns that the user's own shell command is executed and must be confirmed, covering the openWorldHint risk profile with concrete context.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, front-loaded with the outcome and followed by mechanism and the safety caveat. Every sentence earns its place and nothing is repeated from the schema.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a six-parameter tool with no output schema, the description conveys the essential model: what it searches for, how it evaluates each commit, and the sandboxing. It omits practical detail an agent might want, such as what the result reports and that the run may be long because the command executes per commit.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents repo, good, bad, command, symptom, and timeoutSec, making 3 the baseline. The description reinforces the semantics of good/bad/command/symptom in prose but adds no syntax, format, or default detail beyond what the schema provides.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a precise verb and outcome: 'Find the commit that introduced a problem a normal test can't detect,' which immediately distinguishes it from ordinary git bisect. It goes on to name the mechanism (git worktree, the `command`, the `good`/`bad` range) so an agent knows exactly what the tool produces. No sibling in the provided list overlaps, so no differentiation is needed.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The qualifier 'a normal test can't detect' tells the agent when this tool is warranted versus a deterministic bisect, and the closing sentence gives an actionable prerequisite (confirm the shell command with the user first). It stops short of stating an explicit when-not or naming an alternative tool, so it is clear context rather than full routing guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
steer_checkSteering WheelARead-only
Judge every sentence of a draft against its goal. Flags sentences that are off-goal, filler, repetitive or likely invented facts (or unsupported by the given sources). Rewrite only the flagged sentences, then call again to confirm.
| Name | Required | Description | Default |
|---|---|---|---|
| goal | Yes | What the piece is for, e.g. 'persuade a busy CTO to try our API' | |
| text | Yes | The draft to check | |
| sources | No | Reference material the draft should stick to | |
| threshold | No | Flag cut-off probability, default 0.6 |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and openWorldHint=true, and the description is consistent with them: 'likely invented facts (or unsupported by the given sources)' explains why an open-world check is being performed. However, it never says what a run returns (per-sentence spans? probabilities? reasons?) and never explains how the threshold parameter changes the outcome, which matters since there is no output schema. The phrase 'Rewrite only the flagged sentences' is ambiguous about whether the tool rewrites or the caller does, but reads as workflow advice to the caller rather than a mutation claim.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, front-loaded with the core action and followed by the flagged categories and the retry loop. Tight overall, though the parenthetical in sentence two slightly interrupts the flow and could be folded into the category list.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 4-parameter tool with no output schema, the description should convey what the agent receives back and how the flag cut-off affects results; it does neither, leaving the agent to infer the return shape. Purpose and iterative use are covered, so it is adequate but incomplete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the schema already documents goal, text, sources, and threshold (including its 0.5-0.95 range and 0.6 default). The description only alludes to 'sources' parenthetically and never mentions the threshold, so it adds no meaning beyond the structured fields; baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
A specific verb+resource: 'Judge every sentence of a draft against its goal,' followed by the exact failure classes it flags (off-goal, filler, repetitive, invented/unsupported facts). No sibling in the provided list overlaps with this function, so no differentiation is needed and none is missing.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives a clear operating loop: check the draft, rewrite only flagged sentences, then re-invoke to confirm. That tells the agent when and how to use it, but there is no explicit exclusion or named alternative for cases where sentence-level judging is the wrong tool (e.g. whole-document structure).
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
7 tool updates
v0.1.0- First observed
collision_engine - First observed
jev_ask - First observed
laugh_track - First observed
physics_interact - First observed
physics_sandbox - First observed
semantic_bisect - First observed
steer_check
TDQS
Scored across 7 tools
Most tools target clearly distinct tasks: comedy testing (laugh_track), goal alignment (steer_check), git bisection (semantic_bisect), pairwise physics (physics_interact/physics_sandbox), idea collision (collision_engine), and generic typed evaluation (jev_ask). The physics pair/sandbox tools share a domain but differ in single-pair versus 2-8 item batch usage, and jev_ask is explicitly low-level, so boundaries are mostly clear.
All names use lowercase snake_case and are readable, with no camelCase or case inconsistency. However the semantic pattern varies (noun_noun, noun_verb, adjective_verb, prefix_verb), so it is not a strict verb_noun convention, but it remains predictable enough.
Seven tools is within the well-scoped 3-15 range and each covers a distinct capability: specialized evaluators plus one low-level primitive. No obvious redundant filler.
The low-level jev_ask primitive allows arbitrary typed evaluations to fill many gaps, while the specialized tools cover audience reaction, goal checking, bisection, physics, and idea collisions. No major lifecycle operation is obviously missing for this judgment/evaluation toolkit, though more domain-specific evaluators could be added later.
Maintenance
Related MCP Connectors
JSON/YAML, regex, diff, JWT, SQL dialects β the keyless millisecond ops an agent needs mid-task.
The agent-native cloud: database, functions, AI, storage, computers. 50 tools, one API key.
AI Reasoning Cache & Consensus Layer with 11 MCP tools via Streamable HTTP.
See, price, and control every tool call your AI agents make: policy checks, cost, and audit tools.
Related MCP Servers
- AlicenseAqualityBmaintenanceRuns TypeSafe Jev System One packs locally, enabling agents to perform typed Choice, Noul, and Score judgments for tasks like PR auditing, intent routing, and locale classification.51MIT
- AlicenseAqualityBmaintenanceEnables AI assistants to perform ultra-fast, calibrated decision tasks such as boolean evaluation, category selection, scoring, and batch decisions through TypeSafe AI's Jev model.4724 npm1MIT
- AlicenseBqualityCmaintenanceEnables AI agents to obtain typed judgments from TypeSafe's Jev System One models, including yes/no probabilities, multiple-choice selections with distributions, and rubric-based scores, directly usable in code.52AGPL 3.0
- AlicenseNot gradedqualityAmaintenanceEnables Claude Code or any MCP client to ask TypeSafe's Jev for calibrated, typed judgments (probabilities, choices, scores) instead of prose, with local caching and cost tracking.1MIT