BytesJev PlanningChecker
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@BytesJev PlanningCheckerCheck my plan: p1 add Copy link button; repo has Button. Is p1 necessary?"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
BytesJev PlanningChecker
A local MCP server for Claude Code that checks a plan for overengineering after the agent has planned and before it edits anything.
It takes the user's original request, the agent's proposed plan items, and the repository conventions the agent observed, and asks Jev (TypeSafe's System One model) one question per item: is this needed by the request, an existing convention, or a concrete correctness or security need? Each item comes back with a necessity score and an advisory keep / simplify / review. The tool never edits, blocks, or deletes; the agent stays responsible for the plan.
user request ─┐
plan items ─┼─▶ PlanningChecker ─▶ Jev: six yes/no judgments per item ─▶ keep / simplify / review
repo context ─┘ (code decides; Jev judges)Why Jev
Jev returns typed answers with calibrated probabilities, in one request that evaluates several questions at once, for about $0.04 per million input tokens. It does not generate text, so it never talks the agent into or out of anything; it only answers the specific necessity questions the code asks, and the code turns those probabilities into a recommendation with thresholds you can read and tune.
Related MCP server: Agent Fables
Install
You need a TypeSafe API key from https://typesafe.ai. Then, once, from any directory:
claude mcp add --scope user bytesjev -e TYPESAFE_API_KEY=your-key -- npx -y bytesjev
npx -y bytesjev install-skillThe first line registers the MCP server for every Claude Code session on this machine; the key lives in Claude Code's own user config, not in any repository. The second installs the skill to ~/.claude/skills/bytesjev/, which tells Claude Code when to call the tool, what to send, how to treat review, and how to justify keeping a flagged item. Run /mcp in an open session to reconnect, and the PlanningChecker tool is there.
To share the setup with a team, put { "command": "npx", "args": ["-y", "bytesjev"] } under mcpServers.bytesjev in the repository's .mcp.json and the skill under its .claude/skills/bytesjev/; each developer sets TYPESAFE_API_KEY in their own environment.
Watch it judge
npx bytesjev viz # then open http://localhost:4310A local page that shows every PlanningChecker call as it happens: the request, each plan item, Jev's six judgments per item drawn as probability meters with the decision threshold as a tick, and the keep / simplify / review code derives from them, with the reason and the quoted evidence. The MCP server posts each call to the page while it runs, so Claude Code in one window and the page in another give a live picture; if the page is not running, nothing changes for the tool. The page can also run the bundled example itself (it needs TYPESAFE_API_KEY in its environment or in a .env in the current directory). Press ⌘M (or add ?mobile) for a 9:16 phone frame meant for vertical screen recordings; ?dark and ?light force the appearance. Set VIZ_URL if the page is not on http://localhost:4310.
Develop
git clone https://github.com/bennyp11/bytesjev && cd bytesjev
cp .env.example .env # put TYPESAFE_API_KEY in it; .env is gitignored
npm install
npm test # offline: scripted Jev
npm run smoke:plan # one live call: the copy-link example below
npm run smoke:mcp # drives the stdio server the way Claude Code does
npm run viz # the dashboard from source.mcp.json in the checkout runs the server from source (npx tsx src/mcp/index.ts), so opening Claude Code in this directory uses your working copy. The server finds the checkout's own .env wherever it is launched from. Issues and pull requests are welcome at https://github.com/bennyp11/bytesjev.
The tool
Input
user_request: the user's request, verbatim. Never reconstructed from the plan.plan_items: one entry per meaningful proposed addition, each with a stableid, thechange, and the agent's honestrationale. The directly requested work goes in too; it anchors the comparison.repo_context(optional): conventions, utilities, and constraints the agent actually observed. Facts, not guesses.
Output, one result per item in the order sent, plus limitations:
field | meaning |
| Jev's support for "this is needed", in [0, 1]. Not a verified probability; false positives are expected. |
|
|
| one sentence quoting the submitted text that drove the call |
| the submitted sentences it relied on |
| true for every |
Example. Request: "Add a Copy link button to the article page." Plan: (p1) add the button that copies the URL, (p2) a generic action registry, (p3) a new clipboard dependency, (p4) a global feature flag, (p5) reuse the existing clipboard helper. Expected: p1 and p5 keep; p2, p3, p4 simplify or review, each with a reason quoting the rationale.
How it decides
Primitive:
assessNecessityinsrc/jev/primitives/necessity.ts. One Jev request per item with six yes/no judgments:necessary,requested,convention(only when context was given),safety,speculative, anddroppable(only with other items: "would the rest of the plan alone satisfy the request?"). That last one is what makessimplifyconcrete: the smaller path is the plan minus this item.Policy:
decideinsrc/mcp/planning-checker.ts.keepat necessity ≥ 0.65, or when the request or a stated convention supports the item at ≥ 0.7.simplifyonly below 0.35 and droppable and no requested / convention / safety signal at ≥ 0.5 and more than one item. Everything else, including the mid-band and conflicting signals, isreviewwithuncertain: true. Reasons quote submitted text chosen by word overlap; nothing about the repository is invented.Fail open: a missing key, timeout, network, auth, or rate-limit error returns
status: "unavailable"with an emptyresultsarray and a short error category, and the agent proceeds with its plan. Invalid input returnsstatus: "invalid"with messages, never invented results.Limits: 25 items, 4,000 characters of request, 8,000 of context. Truncation is reported in
limitations.
Privacy
The key is read from TYPESAFE_API_KEY (or .env) and never logged or echoed. Only the request, the plan items, and the context are sent to the Jev API; the skill tells the agent not to include secrets or unrelated source. Nothing is written to disk unless BYTESJEV_CACHE=1. BYTESJEV_DEBUG=1 prints one stderr line per call with the status, item count, and milliseconds.
Layout
src/cli.ts `bytesjev` binary: mcp (default) | viz | install-skill
src/mcp/index.ts stdio entry point
src/mcp/server.ts the PlanningChecker tool: schema, description, logging
src/mcp/planning-checker.ts validation, limits, the decision policy, reasons
src/jev/primitives/necessity.ts the six Jev questions per item
src/jev/client.ts Jev client with timeouts, usage accounting, optional cache
src/viz/ the live dashboard: server, page, the messages the MCP server relays
src/example.ts the copy-link example used by the smoke test and the page
.claude/skills/bytesjev/ the Claude Code skill
test/ offline tests with a scripted Jev
scripts/ live smoke testsLicense
MIT. See LICENSE.
Available Tools
1 toolPlanningCheckerPlanning CheckerARead-onlyIdempotent
Advisory check for overengineering. Call it once after you have a plan and before you edit any file.
Send the user's ORIGINAL request verbatim (never reconstruct it), every meaningful proposed addition as its own plan item with your honest rationale, and the repository conventions or constraints that bear on those items. Do not send secrets, .env contents, or unrelated source: the payload is transmitted to the Jev API.
Each item comes back with a necessity_score (Jev's support for the specific question "is this needed by the request, a stated constraint, or correctness/security"; not a verified probability), a recommendation (keep / simplify / review), a reason that cites your submitted text, and an uncertain flag. "simplify" is only given when the other plan items alone would satisfy the request. "review" means the evidence is mixed or missing: decide yourself and name the concrete need if you keep the item. On status "unavailable" proceed with your plan as normal.
Limits: 25 items, 4000 chars of request, 8000 chars of context; truncation is reported in limitations.
| Name | Required | Description | Default |
|---|---|---|---|
| plan_items | Yes | Each meaningful proposed addition with a stable id and your rationale. | |
| repo_context | No | Existing conventions, utilities, or constraints relevant to the plan items. Facts you observed in the repository, not guesses. | |
| user_request | Yes | The user's original request, verbatim. |
Output Schema
| Name | Required | Description |
|---|---|---|
| error | No | |
| errors | No | |
| status | Yes | |
| results | Yes | |
| limitations | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint and openWorldHint, so the safety profile is covered. The description goes further with genuinely additive, non-obvious behavior: the payload is sent to the Jev API (external transmission), truncation is reported in limitations, and the meanings of the 'unavailable' status and 'simplify' vs 'review' outputs are disclosed. That is unusually rich for one paragraph, but it stops short of describing latency or rate-limit behavior, keeping it below the top of the range.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the purpose and the trigger, then the payload rules, then output semantics, then limits. Every sentence carries information, but the output-semantics sentence is long and reads as a single dense block rather than a scannable list, so it is slightly heavier than ideal.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Despite an output schema being present, the definition still explains what each returned field means and how to act on 'review' and 'unavailable', and it documents the input limits (25 items, 4000/8000 chars) with truncation reporting. Nothing an agent needs to invoke this tool correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the description adds real semantics the schema cannot: send the request verbatim, never reconstruct it; decompose each meaningful addition into its own item with an honest rationale; and a hard exclusion list (secrets, .env, unrelated source). These are usage constraints on the three parameters, not restatements.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('advisory check for overengineering') and pins the exact moment to call it: once, after planning, before editing files. An agent knows precisely what this tool is and is not (a verifier, not a planner or editor).
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicit timing ('once after you have a plan and before you edit any file') and an explicit failure-mode instruction ('On status "unavailable" proceed with your plan as normal'). No siblings exist, so no alternative routing is needed, and the when-not condition is still covered.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
1 tool update
v0.1.0- First observed
PlanningChecker
TDQS
Scored across 1 tool
With only one tool, there is no possibility of confusion or misselection between tools. The tool's purpose is clearly distinct by default.
A single tool named 'PlanningChecker' follows a consistent naming pattern by itself. There are no other tools to create inconsistency, so the naming is perfectly consistent within the set.
The server's purpose is a single advisory check, so one tool is appropriate and well-scoped. It is slightly below the typical 3-15 range, but the narrow focus justifies the minimal surface.
The tool covers the full stated purpose of checking a plan for overengineering, with no obvious missing operations for an advisory service. Minor gaps like history or batch processing are not required for its core function.
Maintenance
Related MCP Connectors
Advisory policy preflight for AI-agent spend requests; never executes payments or accesses wallets.
Check AI work against requirements and return structured verdicts, findings, and repair steps.
Security reviews for coding agents: diffs checked against your org policy and live infrastructure.
Score supplied buying-signal evidence and prepare review-only why-now briefs. Read-only, no auth.
Related MCP Servers
- FlicenseNot gradedqualityAmaintenanceProvides a read-only approval gate for AI agent commerce actions, reviewing up to five non-sensitive actions and returning decisions and required evidence without executing, paying, or signing.-
- FlicenseNot gradedqualityBmaintenanceProvides offline, evidence-based preflight risk assessment for agent actions, returning revision-pinned incident patterns and unresolved verification gates to help prevent past failures from recurring.-
- AlicenseNot gradedqualityAmaintenanceEnables coding agents to scout, rank, and preflight software work before implementation, returning evidence-backed ACT, VERIFY, or SKIP decisions for issues and pull requests.60 npm2MIT

Worth Sendingofficial
AlicenseAqualityCmaintenanceEnables agents to evaluate whether a product-adoption message is worth sending by applying a fixed, evidence-based rubric to a draft, recipient context, evidence, and the agent's own assessment, returning a send, revise, or hold recommendation with a score and reasons.2MIT