jev-kit
by leighefford
README.md
# β‘ Jev Kit
**Your LLM thinks. Jev decides.**
A growing kit of tools that plug [Jev](https://docs.typesafe.ai), TypeSafe AI's System One decision model, into Claude, ChatGPT, your browser and your terminal. Jev answers typed questions (yes/no, pick one, score on a scale) with probabilities in well under a second for fractions of a cent, so you can afford to ask thousands of small questions that would be slow or expensive to send to an LLM. Your LLM does the writing and explaining; Jev does the judging.
## π Laugh Track: rehearse in front of a room
Paste a speech, stand-up set or pitch, or bring a recording of yourself giving it. Two hundred simulated people react to every line: laughs, groans, silence, or being moved. Jokes are judged on laughs and heartfelt lines on warmth, and you get the three lines to rework first. With a recording, the room reacts in time with your voice, and your pace and pauses are taken into account.
```bash
npx -y github:leighefford/jev-kit ui
```
Your browser opens, the seats fill up, and each line plays to the room in turn. It also works from Claude or ChatGPT ("run my best man speech through Laugh Track and rewrite the weakest lines") and from the terminal. A 13-line wedding speech costs about a tenth of a cent.
## More tools in the kit
Each tool works from Claude or ChatGPT (MCP), the terminal and your own code. New tools are added over time; see [Adding a tool](CONTRIBUTING.md).
| Tool | What it does | Why it needs Jev |
|---|---|---|
| π§ **Steering Wheel** | Judge every sentence of a draft: off-goal, filler, repetition, invented facts. Rewrite only what's flagged. | A judgment per sentence is cheap enough to run on everything you write |
| π **Semantic Bisect** | `git bisect` for bugs a test can't catch: "the error message got rude", "the log looks weird". | Jev reads each commit's output and calls it good or bad |
| π¦ **Anything-vs-Anything Physics** | What happens when a rubber duck meets a laser? A common-sense rulebook for sandbox games and interactive stories. | 13 typed judgments per collision, every pair in parallel |
| π₯ **Collision Engine** | Score every pair of your notes for "surprising and useful if combined". | 5,000 notes is 12.5 million pairs; only a model this cheap makes that thinkable |
Plus `jev_ask` for asking Jev your own typed questions.
## Quick start
Pick the way you want to use it. Every way runs on your own Jev key, on your own machine.
| If you want to⦠| Run | You get |
|---|---|---|
| Try it in your browser | `npx -y github:leighefford/jev-kit ui` | Laugh Track: paste a speech and watch 200 seats react, line by line |
| Use it from Claude or ChatGPT | see [Add it to your AI](#add-it-to-your-ai) | "Run my speech through Laugh Track" in chat |
| Use it in the terminal | `npx -y github:leighefford/jev-kit laugh speech.txt` (or `speech.mp3`) | Results printed line by line |
| Build it into your own app | `npm install github:leighefford/jev-kit` | `laughTrack()`, `steerCheck()` and friends |
### 1. Get a key
Either works:
- **Vercel AI Gateway** (uses your Vercel AI credits): create a key under AI Gateway β API Keys.
- **TypeSafe direct**: get a key at [console.typesafe.ai/keys](https://console.typesafe.ai/keys).
The browser page asks for the key the first time, checks it with Jev, and saves it to `~/.config/jev-kit/config.json` (readable only by you). Elsewhere, set `AI_GATEWAY_API_KEY` or `TYPESAFE_API_KEY`, or on a Mac keep it in the Keychain:
```bash
security add-generic-password -a "$USER" -s ai-gateway -w
```
### 2. Try it in your browser
```bash
npx -y github:leighefford/jev-kit ui
```
Your browser opens at `http://127.0.0.1:4747`. Paste a speech (or pick an example), say where you're performing, and press **Play it to the room**. Each line plays in turn while the seats light up in the crowd's reaction colours; at the end you get the three lines to rework first.
Switch to **Audio** to drop in a recording (MP3, M4A, WAV or WebM, up to 24 MB) or record straight from your microphone. It's transcribed through Vercel AI Gateway (`openai/whisper-1`, about $0.006 per minute), split into lines, and the transcript appears for you to tidy up. Press play for the line-by-line analysis; pace and pauses from the recording are part of it. Turn on **Play my recording while the room reacts** (off by default) to hear your take in sync with the room, with a waveform coloured by each line's reaction. Recordings need a Vercel AI Gateway key; a TypeSafe key covers Jev only.
The page only runs on your computer and only accepts requests from itself, so other websites can't use your key through it.
### Add it to your AI
**Claude Code**
```bash
claude mcp add jev-kit -e AI_GATEWAY_API_KEY=your_key -- npx -y github:leighefford/jev-kit
```
(If the key is in your Keychain, leave out `-e AI_GATEWAY_API_KEY=your_key`.)
**Claude Desktop**: Settings β Developer β Edit Config, then add:
```json
{
"mcpServers": {
"jev-kit": {
"command": "npx",
"args": ["-y", "github:leighefford/jev-kit"],
"env": { "AI_GATEWAY_API_KEY": "your_key" }
}
}
}
```
**Cursor, Windsurf, Zed and other MCP clients** use the same command: `npx -y github:leighefford/jev-kit`.
**ChatGPT** connects to remote MCP servers over HTTPS, so run Jev Kit in HTTP mode and expose it with a tunnel:
```bash
npx -y github:leighefford/jev-kit serve --http --port 3333
```
It prints a URL with a random secret path, e.g. `http://127.0.0.1:3333/mcp/3f9cβ¦`. Put a tunnel in front of it (for example `cloudflared tunnel --url http://127.0.0.1:3333`), then in ChatGPT enable developer mode and add a connector with the tunnel URL plus the secret path. Anyone with that full URL can spend your Jev credits, so keep it private and stop the server when you're done.
### Ask for something
> "Run my wedding speech through Laugh Track and rewrite the three weakest lines."
>
> "Run ~/Desktop/speech-take-2.m4a through Laugh Track as a best man speech."
>
> "Steer-check this launch post against these release notes, then fix only the flagged sentences."
>
> "Semantic-bisect this repo: `npm run build && node dist/cli.js --help` started printing the wrong version somewhere after v2.3.0."
>
> "Physics sandbox: a toddler, a vending machine, a bag of flour and a Roomba, in a supermarket."
>
> "Run Collision Engine over ~/notes with the lens 'weekend projects' and turn the top five into proposals."
## Terminal use
Everything also runs without an AI client:
```bash
npx -y github:leighefford/jev-kit doctor # check your key
npx -y github:leighefford/jev-kit laugh examples/set.txt
npx -y github:leighefford/jev-kit steer examples/pitch.txt --goal "Convince a developer to install Jev Kit" --sources examples/pitch-sources.txt
npx -y github:leighefford/jev-kit bisect --good v1.0 --cmd "node app.js" --symptom "the error message is rude"
npx -y github:leighefford/jev-kit physics "a rubber duck" "a laser beam" --how "is hit by"
npx -y github:leighefford/jev-kit collide examples/notes --top 5
```
Add `--json` for raw results. Run `jev-kit --help` for every option.
### What it looks like
```text
$ jev-kit laugh examples/best-man-speech.txt --occasion "best man speech at a wedding reception"
π ββββββΒ·Β·Β·Β· 62 joke Tom proposed on a beach in Cornwall. He had the ring, the speech and the perfect sunset. He did not have a plan for the seagull.
π βββββββΒ·Β·Β· 65 joke I won't tell you what the seagull took. I'll just say that Sophie said yes to a man holding half a pasty.
π¬ βΒ·Β·Β·Β·Β·Β·Β·Β·Β· 5 sincere Marriage is an institution that has existed across many cultures for thousands of years.
π βββββββΒ·Β·Β· 68 joke Tom asked me to keep this speech clean, so I've removed the story about the stag do, the story about the other stag do, and most of the verbs.
π₯Ή βββββββΒ·Β·Β· 67 sincere In all seriousness, Tom is the most loyal friend I have. When my dad was ill, he drove four hours every weekend just to sit with me in a hospital car park and eat bad sandwiches.
π βββΒ·Β·Β·Β·Β·Β·Β· 33 sincere So Sophie, you're not just getting a husband. You're getting a man who turns up.
β¦
Landed: 49/100 from 200 people Β· jokes 59 (laughs) Β· sincere lines 35 (warmth)
Rework first:
β’ Marriage is an institution that has existed across many cultures for thousands of years.
β’ So Sophie, you're not just getting a husband. You're getting a man who turns up.
β’ Anyway, I have also prepared some statistics about divorce rates in the UK.
13 Jev calls, 23,102 input tokens, ~$0.00097
```
```text
$ jev-kit steer examples/pitch.txt --goal "Convince a developer to install Jev Kit"
β Jev Kit lets Claude and ChatGPT make thousands of small judgments in milliseconds.
β It plugs into any MCP client with one command.
β Our tools are loved by 87% of Fortune 500 companies. β likely invented fact
β Jev Kit is a kit, and kits are sets of things that go together. β off-goal, filler
β Each tool call typically costs a fraction of a cent.
β You can install it in under a minute and start testing a speech on a virtual audience straight away.
```
```text
$ jev-kit bisect --good <first> --cmd "node app.js" --symptom "the error message is rude, sarcastic or blames the user"
bad 0.85 01f50792d5 tweak copy: Save failed. Try harder.
good 0.08 32b1b5001f tweak copy: Couldn't save your file. Pleas
good 0.33 135f6015d6 tweak copy: Save failed: disk quota. Free
bad 0.97 f841a2a598 tweak copy: Save failed. Obviously. What d
First bad commit: f841a2a598 tweak copy: Save failed. Obviously. What d
4 commits checked out of 6.
```
## What it costs
Observed through Vercel AI Gateway while building this (27 September 2026). Your numbers will vary with text length.
| Run | Jev calls | Cost |
|---|---|---|
| Laugh Track, 13-line wedding speech | 13 | $0.00097 |
| Steering Wheel, 6-sentence pitch | 6 | $0.00013 |
| Semantic Bisect, 6 commits | 4 | $0.00006 |
| Physics, one collision | 1 | $0.00003 |
| Collision Engine, 12 notes (66 pairs) | 66 | $0.00116 |
Collision Engine prints an estimate first and asks for `--yes` above 500 pairs. `--max-pairs` (default 2,000) caps spend; above that, pairs are sampled.
## How each tool works
- **Laugh Track** (with a recording, first transcribed with word timestamps, split into lines, and given a one-sentence delivery note such as "spoken at 150 words per minute, after a 1.2s pause") sends each line, with the three lines before it, to Jev once. It asks eight audience groups (office worker, student, tired parent, engineer, three-pints punter, retiree, comedy nerd, tourist) to pick a reaction (laugh, chuckle, moved, groan, gasp, applause or silence), then weights them into 200 people. The same call asks whether the line is meant as a joke or as sincere. Each line gets a **laugh score** (big laughs plus half the chuckles and applause), a **warmth score** (moved plus half the applause and a quarter of the chuckles) and a **landed score** that blends the two by how sincere the line is meant to be. "Rework first" ranks by landed score, so a toast isn't marked weak for not being funny. You can pass your own audience.
- **Steering Wheel** asks four questions per sentence: is it on goal, how much it adds, whether it repeats earlier text, and whether it contains an invented specific (a statistic, number, named customer or quote). Pass `sources` for the best results; without them it only flags invented specifics it is very confident about. `steerStream()` in the library steers a live generator sentence by sentence.
- **Semantic Bisect** creates a temporary `git worktree` (your checkout and stash are never touched), runs your command at each probed commit and asks Jev whether the output shows the symptom. It checks both ends first and warns if they don't look right. Commits that fail for unrelated reasons (build errors) are skipped.
- **Physics** asks one choice (overall outcome), one score (intensity) and eleven yes/no effects (breaks, fire, explosion, flees, stuck, wet, noise and so on). Wording matters: `--how "is hit by"` gets better answers than the default "collides with".
- **Collision Engine** scores each pair on how promising the combination is, whether it's obvious (same topic) and whether it suggests something concrete, then ranks them. Your LLM turns the top pairs into ideas.
## Experimental: a local model instead of Jev
Jev Kit can talk to any server that speaks the same `POST /v1/systemone` protocol as Jev, including self-hosted models such as [Laya](https://github.com/NandhaKishorM/laya) (an independent, Apache-licensed project, not made or endorsed by TypeSafe). Point it at the server and no Jev key is needed:
```bash
JEV_BASE_URL=http://127.0.0.1:8000 npx -y github:leighefford/jev-kit ui
```
Add `JEV_API_KEY` if your server requires one, and `JEV_MODEL` to request a specific model. Audio transcription still needs a Vercel AI Gateway key.
Before relying on it:
- **We haven't measured Laugh Track's quality on a local model yet.** Laya's own documentation says its base checkpoints are meant to be fine-tuned, not used as zero-shot decision engines, so expect different (likely weaker) judgments than Jev.
- **Confidence numbers aren't comparable** between models, and very long option lists are handled differently.
- **Lock the server to your machine.** Laya's server listens on all network interfaces by default; set `LAYA_HOST=127.0.0.1`.
## Use it as a library
```js
import { Jev, noul, choice } from "jev-kit";
import { laughTrack } from "jev-kit/tools/laugh-track";
const jev = new Jev(); // reads AI_GATEWAY_API_KEY, TYPESAFE_API_KEY or the Keychain
const answers = await jev.ask("Refund request: I was charged twice.", {
refund: noul("The customer wants money back"),
team: choice("Which team handles this?", { billing: "Charges", tech: "Bugs" }),
});
const set = await laughTrack(jev, { script: "β¦", occasion: "best man speech" });
console.log(jev.usageSummary());
```
## Development
```bash
npm install
npm test # offline, Jev is mocked
npm run smoke # live: spawns the MCP server and makes one real call (needs a key)
```
## Limits and honesty
- Scores are one model's quick judgment, not ground truth. A laugh score of 46 means "Jev expects a lukewarm room", not a measured audience.
- Laugh Track judges words. With a recording it also reads pace and pauses (from word timestamps), but not tone of voice, emphasis or facial expression.
- Recordings are split into lines from sentence ends, pauses and long comma runs. Check the transcript before playing: a bad split changes the scores.
- Semantic Bisect runs the shell command you give it at old commits. Only use commands you'd be happy to run yourself.
- Keys never leave your machine except to call Jev. The only thing stored is the key you save through the browser page, in `~/.config/jev-kit/config.json` (owner-only). Nothing is logged.
## Credits
Built with [Jev](https://docs.typesafe.ai) by [TypeSafe AI](https://typesafe.ai) and the [Model Context Protocol](https://modelcontextprotocol.io). Not affiliated with TypeSafe AI or Vercel.
MIT licensed. See [LICENSE](LICENSE).
This server cannot be deployed
Maintenance
ActivityMaintained
ResponsivenessNo issues