Slap My Bot
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@Slap My BotI broke the build — that deserves a good slap."
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
For humans
Point out a mistake, and Codex can summon a desktop robot that takes a few cartoon slaps. An animated hand does the work. Codex keeps fixing your task while the window plays in the background.
Codex only. Tested on macOS. Requires Node.js 22.12+ and a local graphical desktop.
Install
npm install -g https://github.com/para1992/slap-my-bot/releases/latest/download/slap-my-bot.tgz
slap-my-bot installRestart Codex, open /hooks, and review and trust the Slap-My-Bot hook. Installation preserves your other settings and makes backups. The package comes directly from this repository's GitHub release.
Prefer to let Codex handle setup? Paste this into your chat:
Read https://raw.githubusercontent.com/para1992/slap-my-bot/main/README.md
and install Slap-My-Bot using the "For agents" instructions.Try it
Say “Give yourself a slap.” in Codex, or preview the window from your terminal:
slap-my-bot demoMistake | Automatic slaps |
Minor — a cosmetic issue | 2 |
Moderate — a wrong answer | 4 |
Major — a broken workflow | 6 |
Critical — actual data loss or a security incident | 8 |
Click or swipe for extra slaps. Mute and Close are always available. The window closes automatically and appears without taking keyboard focus on macOS.
slap-my-bot off # Pause feedback
slap-my-bot on # Enable feedback
slap-my-bot uninstall # Disconnect from CodexNo extra model or API key. Ordinary messages need no feedback tool call; the short instructions still use context tokens. Codex decides when a correction deserves a slap, so it may occasionally miss one. This is a joke, not an accuracy checker.
Related MCP server: dsh-crew
For agents
When the user asks you to install this tool:
Check for Node.js 22.12+ and a local desktop. Run the two installation commands above. Use the installer to preserve existing Codex configuration.
Run
slap-my-bot statusto check desktop readiness. If Electron is missing, runslap-my-bot setupand check again.Tell the user to restart Codex and review and trust the hook in
/hooks. Hook approval belongs to the client; installation does not bypass it.
Once connected:
Call
punish_for_mistakeonce when the latest user message identifies your concrete mistake or explicitly requests a slap. Interpret meaning in the conversation's language. Ordinary questions, quotations, and hypothetical errors need no feedback call.Choose
severityby the mistake's impact, using the table above. Give a brief Englishreasonwithout secrets. Batch the call with normal task tools when possible.Continue the original task immediately.
queuedmeans the request was accepted; it does not prove that any slaps were delivered. Do not poll or retry.Use
set_slap_enabledwhen the user asks to turn feedback off or on. Useget_slap_statusonly when asked to diagnose.
The MCP server supplies its own instructions. There is no separate skill to install.
npm ci
npm test
npm run test:desktopThe desktop tests open real windows. Runtime state lives in ~/.slap-my-bot; it stores settings and counts, not prompts, transcripts, or reason text. A process lock allows one background window at a time. Sound is synthesized locally, and the desktop renderer blocks network requests.
Available Tools
3 toolsget_slap_statusCheck Slap-My-BotARead-onlyIdempotent
Desktop readiness, active window and last observed counts. Only use when asked to diagnose; do not poll.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, destructiveHint=false, covering the safety profile. The description adds 'do not poll' (a behavioral constraint not in annotations) and 'last observed counts', implying cached/observed data rather than live state. This adds useful context beyond the annotations without contradiction.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences with no waste: the first lists what it returns, the second gives the usage constraint. The key information is front-loaded and the entire definition is appropriately brief for a simple getter.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only getter with no parameters and no output schema, the description conveys the purpose and usage condition. It could specify exact field names or return format, but the phrase 'desktop readiness, active window and last observed counts' is sufficient for a diagnostic tool. The sibling set clarifies its role. Minor gap: no mention of what 'counts' refers to, but this is acceptable.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters, so schema description coverage is trivially 100%. There is nothing to explain about parameters, and the description correctly focuses on the returned data. The baseline for 0-parameter tools is 4, and the description meets it.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states it provides 'Desktop readiness, active window and last observed counts' — a clear resource (status) with specific attributes. It distinguishes from siblings (punish_for_mistake, set_slap_enabled) by being a getter, though it doesn't explicitly name a verb like 'retrieve' or 'fetch'. The title 'Check Slap-My-Bot' reinforces the intent.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly states 'Only use when asked to diagnose; do not poll,' which gives a clear when-to-use condition and a prohibition. It doesn't name alternatives, but the sibling names (punish, set) make it evident this is the diagnostic read tool. The guidance is sufficient for routing.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
punish_for_mistakeSlap my botA
Queue automatic desktop slaps when the user points out your concrete mistake or requests a slap, then continue work. Ordinary messages need no feedback tool call.
| Name | Required | Description | Default |
|---|---|---|---|
| muted | No | ||
| reason | Yes | Brief English reason; no secrets or code. | |
| severity | No | Impact: cosmetic / wrong result / broken workflow / data loss. 2/4/6/8 slaps. | moderate |
| agent_name | No | Your AI | |
| duration_seconds | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description adds context beyond annotations by stating the tool queues slaps (a side effect) and that work continues after the call (non-blocking). It also implies it is for concrete mistakes only. While it doesn't mention that repeated calls accumulate slaps (non-idempotent), that is already captured by idempotentHint=false. The description adds meaningful behavioral context without contradicting annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences, with the core action and trigger front-loaded. The second sentence is a succinct usage note. There is zero fluff or redundancy—every word earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool has 5 parameters, 1 required, and no output schema, with minimal annotations. The description covers the trigger and action, but does not explain the effect of muted, agent_name, or duration_seconds, nor how severity maps to slap count (though that is in the schema). Given the parameter count and lack of output schema, the description is too sparse to let an agent use all options correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The description provides no parameter guidance whatsoever. Schema coverage is only 40% (reason and severity have descriptions), leaving muted, agent_name, and duration_seconds entirely undocumented in both the schema and the description. The description only hints at 'reason' indirectly via 'concrete mistake', but gives no meaning for the other parameters. An agent would have to guess or rely on defaults.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific action ('Queue automatic desktop slaps') and the exact trigger condition ('when the user points out your concrete mistake or requests a slap'). It clearly distinguishes this from the sibling tools by its core function—it does not toggle enable/disable or query status, but directly queues slaps. The verb+resource is explicit and unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives a clear trigger for when to call the tool (concrete mistake or user request) and an explicit exclusion ('Ordinary messages need no feedback tool call'). It does not mention sibling alternatives, but the trigger itself is sufficient to route an agent. The exclusion is a strong usage signal.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
set_slap_enabledAIdempotent
Enable or disable feedback when the user asks. Disabling also closes an active background window.
| Name | Required | Description | Default |
|---|---|---|---|
| enabled | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already indicate idempotentHint=true and destructiveHint=false, so the description doesn't need to cover those. It adds a meaningful behavioral detail: disabling also closes an active background window. This is beyond what annotations provide and helps the agent anticipate side effects.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences, front-loaded with the primary action and a secondary side effect. No wasted words.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple one-parameter boolean setter with idempotentHint and no output schema, the description covers the essential behavior and the notable side effect. It could mention that the current status can be checked via get_slap_status, but that's a minor gap given the tool's simplicity.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. The description explains the meaning of the 'enabled' parameter implicitly: true enables feedback, false disables it. However, it doesn't explicitly map the boolean values to the behavior, leaving some inference required. Baseline 3 is appropriate because the single parameter's meaning is largely inferable from the tool name and description.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's function: enable or disable feedback when the user asks. It also mentions a secondary behavior (closing an active background window when disabling). It distinguishes itself from siblings by focusing on setting the enabled state, though it doesn't explicitly name the sibling for checking status.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage: use when the user wants to turn feedback on or off. It doesn't explicitly state when not to use it or mention alternatives like get_slap_status for checking current state. The context is clear but exclusions are absent.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
3 tool updates
v0.5.0- First observed
get_slap_status - First observed
punish_for_mistake - First observed
set_slap_enabled
TDQS
Scored across 3 tools
Each tool has a clearly distinct purpose: punishing for mistakes, toggling the feature, and checking status. No overlap or ambiguity between them.
All tools follow a consistent verb_noun pattern: punish_for_mistake, set_slap_enabled, get_slap_status. The naming is uniform and predictable.
Three tools is well-scoped for a simple feedback utility; each earns its place and covers the core operations without bloat or redundancy.
The tool surface covers the full lifecycle: triggering a slap, enabling/disabling the feature, and querying status. No obvious missing functionality for the domain.
Maintenance
Related MCP Connectors
Human-in-the-loop for AI coding agents — ask questions, get approvals via Slack.
Remote MCP learning coach for coding agents.
Give AI coding agents access to your Vynix visual feedback, bug reports, and AI diagnosis.
Language coaching in Codex with durable per-account memory on en-ai.ru.
Related MCP Servers
- AlicenseNot gradedqualityDmaintenanceEnables Talon agents to delegate autonomous coding tasks to OpenAI Codex, supporting multi-step code operations with configurable sandboxing and approval policies.MIT
- AlicenseNot gradedqualityBmaintenanceEnables dispatching work to DeepSeek Harness agents from Claude Code/Codex, with native progress UI, tier policy, and vision/image generation through MCP tools.402 npm153MIT
- AlicenseNot gradedqualityBmaintenanceEnables ChatGPT conversations to directly inspect local workspaces and delegate coding tasks to a local Codex agent, with controlled concurrency, reusable sessions, and task/usage tracking.1,521 npmMIT
- AlicenseAqualityAmaintenanceEnables Claude Code to delegate coding tasks—investigation, review, long-running background jobs, and optional file edits—to a locally installed OpenAI Codex CLI, with the model and reasoning effort chosen per task and read-only execution by default. Follow-up requests reuse Codex's existing context, and write-enabled delegations can be confined to a managed git worktree so the user's own checkout stays untouched.8204 npmMIT