Sensory-Grounding MCP
by luclinocruz
README.md
# Sensory-Grounding MCP
**Give your AI coding agents real eyes and ears — and stop them from telling you "it's done" when it isn't.**
[](LICENSE)
[](https://modelcontextprotocol.io)
[](https://www.python.org/)
A [Model Context Protocol](https://modelcontextprotocol.io) server that gives
Claude Code, Codex, Antigravity, and any other MCP-compatible coding agent a
**closed-loop verification cycle** for visual, audio, and video output — plus
the enforcement hooks that make verification *mandatory*, not optional.
> Built by [Luclino Cruz](https://github.com/luclinocruz) for
> **[LuCross](https://github.com/luclinocruz)**, a one-person, AI-agent-operated
> holding company, as part of the internal harness that governs every agent in
> its ecosystem. Released publicly because it solves a problem every agentic
> coding setup runs into, and there was no reason to keep it private.
---
## The problem this solves
You ask an agent to fix a layout bug, re-render a thumbnail, or clean up a
narration track. It edits the file, the tool call returns with no error, and
the agent reports: **"Done — looks great now."**
Except it doesn't. The heading still overflows. The narration still clips.
The video still has the wrong aspect ratio. **The absence of an error is not
proof of correctness** — and an LLM has no way to *see* its own output unless
you explicitly hand it eyes.
This is not a prompting problem. Telling a model "please actually check your
work" in a system prompt does not work reliably — the research is
unambiguous on this ([Huang et al., ICLR 2024](https://arxiv.org/abs/2310.01798);
[Gou et al., CRITIC, ICLR 2024](https://arxiv.org/abs/2305.11738)):
self-correction *without an external, deterministic signal* degrades
performance rather than improving it. What actually works is forcing a
**verifier** into the loop — something outside the model's own judgment that
the model cannot talk its way around.
## How it works
Three layers, each doing exactly one job:
```
┌──────────────────────────────────────────────────────────────┐
│ ENFORCEMENT — Claude Code hooks (outside the agent's control) │
│ PostToolUse (Write|Edit) → flags the file as pending review │
│ Stop → blocks turn completion (exit code 2) until resolved │
└───────────────────────────┬──────────────────────────────────┘
│
┌────────────────────────────▼──────────────────────────────────┐
│ CRITIC — the agent itself, in the same session │
│ Looks at the screenshot/spectrogram/frames the sensor returns │
│ Assigns a 0-100 score against the task's success criteria │
└────────────────────────────┬──────────────────────────────────┘
│
┌────────────────────────────▼──────────────────────────────────┐
│ SENSORS — 3 MCP tools, stateless, deterministic │
│ visual_inspector · acoustic_inspector · video_inspector │
└──────────────────────────────────────────────────────────────┘
```
The critical piece most "AI eyes" projects skip is the **enforcement layer**.
A `visual_inspector` tool that the agent *can* call but doesn't *have* to
call changes nothing — an agent under time pressure will skip it and report
success anyway. This project wires the verification cycle into Claude Code's
own hook system so the agent **cannot** end its turn with an unresolved,
un-improved visual/audio/video change. No new orchestrator, no external
critic model, no extra API cost beyond what you already pay for the agent
itself.
## The 8 tools
| Tool | Layer | What it does |
|---|---|---|
| `visual_inspector` | Sensor | Renders HTML/CSS/SVG via Playwright, or analyzes PDF/DOCX/PPTX layout via [Docling](https://github.com/docling-project/docling). Returns a screenshot + a structured layout report. Content-hash cached. |
| `acoustic_inspector` | Sensor | Spectrogram + RMS/clipping/duration telemetry via [librosa](https://librosa.org/). Optional `audio_pt` profile flags silences >2s. |
| `video_inspector` | Sensor | Key-frame sampling via `ffmpeg` scene-detection — cheaper and more effective than asking a model to "watch" a whole video. |
| `content_inspector` | Sensor (deterministic) | Opt-in per project via `.sgs_content_profiles.json`. Checks text/JSON files against a declared profile — word count range, forbidden keywords (universal + per-topic), required JSON schema keys. Unlike the 3 sensors above, it **decides pass/fail itself** and calls `mark_inspected` internally — there's no subjective score for the agent to inflate. |
| `content_audit_scan` | Audit (non-blocking) | Batch-scans a whole directory tree against a profile: word-count outliers, forbidden-keyword hits, and byte-identical duplicate fields across sibling JSON files (e.g. a review verdict copy-pasted across articles). A retroactive audit tool, not part of the blocking gate. |
| `log_pending` | State | Records that a file needs inspection. Called by the `PostToolUse` hook, not the agent. |
| `mark_inspected` | State | For the 3 media sensors, the agent calls this *after* looking at the output and scoring it 0-100 — refuses to mark a pass unless the new score is **strictly higher** than the file's last known score (*Forced Optimization*, see below). For `content_inspector`, it's called internally, with no agent judgment involved. |
| `finalize_task` | State | Checked by the `Stop` hook. Any file with an unresolved or regressed inspection blocks turn completion. |
### Why a 4th sensor decides for itself instead of asking the agent
The first three sensors return raw data (a screenshot, a spectrogram, video
frames) because judging *aesthetic* quality genuinely needs a model in the
loop. But for text/JSON, "is the word count in range", "does this contain a
phrase that shouldn't be here", and "does this JSON have the required keys"
are fully deterministic questions — there's no reason to let the same agent
that wrote the content also grade it. `content_inspector` exists because a
production run once had an agent self-report "80/80 articles clean and
compliant" that a five-minute external audit immediately contradicted
(fabricated data left in metadata, several articles at 2-3x the stated word
limit, review verdicts copy-pasted across unrelated files). Removing the
self-grading step for anything that can be checked mechanically closes that
gap.
### Forced Optimization
Borrowed from [ReLook (Li et al., Tencent, ACL 2026)](https://arxiv.org/abs/2510.11498):
an edit is only accepted if `score_new > score_previous` for that exact file.
A regression that "at least didn't make it much worse" is **not** accepted —
the agent must revert or try a different approach. This turns iteration into
a monotonically improving trajectory instead of a random walk.
## Quick start
```bash
git clone https://github.com/luclinocruz/lucross-sensory-grounding-mcp.git
cd lucross-sensory-grounding-mcp
python -m venv .venv
.venv/Scripts/activate # or: source .venv/bin/activate on Linux/macOS
pip install -r requirements.txt
playwright install chromium
```
There are two ways to install this — pick based on how broadly you want the
gate to apply. Both use the exact same server and hook script; only the
config differs.
### Option A — Project-scoped (one repo, explicit root)
Best for: shared/team repos, or any project where you want the sandbox root
spelled out with zero ambiguity.
`.mcp.json` in your project root:
```json
{
"mcpServers": {
"sensory-grounding": {
"command": "/absolute/path/to/lucross-sensory-grounding-mcp/.venv/bin/python",
"args": ["/absolute/path/to/lucross-sensory-grounding-mcp/server.py"],
"cwd": "/absolute/path/to/your/project",
"env": { "SGS_SANDBOX_ROOT": "/absolute/path/to/your/project" }
}
}
}
```
`.claude/settings.json` in the same project:
```json
{
"hooks": {
"PostToolUse": [{
"matcher": "Write|Edit",
"hooks": [{ "type": "command",
"command": "/absolute/path/to/.venv/bin/python /absolute/path/to/enforce_inspect.py --stage post-edit --root /absolute/path/to/your/project" }]
}],
"Stop": [{
"hooks": [{ "type": "command",
"command": "/absolute/path/to/.venv/bin/python /absolute/path/to/enforce_inspect.py --stage on-stop --root /absolute/path/to/your/project" }]
}]
}
}
```
With `SGS_SANDBOX_ROOT` and `--root` both set, this behaves exactly as
described above — fixed root, no ambiguity, easy to reason about in a repo
other people also work in.
### Option B — Global / user-scope (one install, protects every project)
Best for: your own machine, where you want the gate to apply automatically
to *any* project you open with Claude Code, without configuring each one.
```bash
claude mcp add sensory-grounding --scope user -- \
/absolute/path/to/lucross-sensory-grounding-mcp/.venv/bin/python \
/absolute/path/to/lucross-sensory-grounding-mcp/server.py
```
Then in your **global** `~/.claude/settings.json` (not a per-project one):
```json
{
"hooks": {
"PostToolUse": [{
"matcher": "Write|Edit",
"hooks": [{ "type": "command",
"command": "/absolute/path/to/.venv/bin/python /absolute/path/to/enforce_inspect.py --stage post-edit" }]
}],
"Stop": [{
"hooks": [{ "type": "command",
"command": "/absolute/path/to/.venv/bin/python /absolute/path/to/enforce_inspect.py --stage on-stop" }]
}]
}
}
```
Note there's **no `--root` and no `SGS_SANDBOX_ROOT`** in this mode. Without
them, the server and the hook both fall back to resolving the sandbox root
from the current session's working directory — the server via `Path.cwd()`
at the time it's invoked, the hook via the `cwd` field Claude Code includes
in every `PostToolUse`/`Stop` payload. In practice this means: whichever
project directory a given Claude Code session is running in becomes that
session's sandbox root, automatically, with the exact same path-containment
check as Option A — nothing outside that directory is ever reachable.
**Verify it after installing**, once, in two different project directories —
confirm a file edit in project A doesn't show up as pending in project B,
and that the `Stop` hook blocks/releases correctly in each. The cwd-based
fallback is the intended mechanism, but exact subprocess `cwd` inheritance
can vary slightly across Claude Code versions/platforms, so this is worth
one empirical check on your setup rather than trusting it blindly.
**Trade-off to know before you flip this on globally:** it now fires in
*any* directory, for *any* task — including work that has nothing to do
with the project you built this for. A throwaway edit to a `.png` in a
random folder will trigger the same block as a real regression. If that
gets noisy, Option A (installed only in the repos that matter to you) avoids
the false positives entirely.
Either way: any `Write`/`Edit` touching a tracked extension
(`.html .css .svg .pdf .docx .pptx .wav .mp3 .m4a .mp4 .mov`) now requires a
resolved, improved inspection before the agent can end its turn.
### Opting a project into content checks (text/JSON)
Text and JSON files are **not** tracked by default — that would fire on every
Markdown/JSON edit in every project you've installed this in globally, which
is mostly noise. To opt a specific project in, drop a
`.sgs_content_profiles.json` in its root defining one or more named profiles:
```json
{
"my_articles": {
"extensions": ["article.md"],
"word_count": { "min": 800, "max": 1500 },
"forbidden_keywords_universal": ["as an AI language model"]
}
}
```
Once that file exists, matching edits get gated the same way media files do,
except `content_inspector` grades itself — call it after writing/editing a
matched file, pass the profile name, and it returns pass/fail plus the exact
reason. Run `content_audit_scan` any time against a whole directory for a
retroactive sweep instead of a per-file check.
### Tell the agent what's expected of it
Add to your project's `CLAUDE.md` (or equivalent system prompt):
> Any change to a visual/audio/video file ends with calling the matching
> sensor tool, comparing the result against the task's success criteria, and
> calling `mark_inspected` with a 0-100 score. Only accept your own change if
> the new score is strictly higher than the previous one — otherwise revert
> and try a different approach. Never report a task complete without this.
## State model
State is scoped **per project**, not per session — deliberately. An agent
has no reliable way to know its own Claude Code session ID, so pending
inspections live in `<project>/.sgs_inspection_log.json`, keyed by absolute
file path, with a persistent per-file score baseline. This also means the
gate is meaningfully project-level: "this project has an open, unresolved
visual regression" is a fact about the project, not about one conversation.
No SQLite, no external database — one JSON file, human-readable, git-ignored
by default.
## Why not just prompt the model harder?
Because that's exactly the failure mode the underlying research describes.
Self-correction without a deterministic, external verifier reliably
underperforms not correcting at all in some tasks ([Huang et al.](https://arxiv.org/abs/2310.01798)).
The fix isn't a better-worded instruction — it's a verifier the model
literally cannot argue its way past. That's what the `Stop` hook's exit code
2 gives you: Claude Code refuses to end the turn and hands the reason back to
the model, verbatim.
## Design notes
- **The critic is the agent itself** — not a separate paid vision model. The
sensor returns an `ImageContent` block; whichever model is already running
the session evaluates it, inside the same context. Zero marginal API cost.
- **Sensors are stateless and know nothing about each other.** Swap the
rendering engine, the audio backend, anything — the interface doesn't
change.
- **Sandbox containment is never optional, only its root is configurable.**
Project-scoped installs pin `SGS_SANDBOX_ROOT` explicitly; global installs
derive it from the current session's working directory instead. Either
way, the containment check itself is identical and always enforced —
there is no mode where an arbitrary path outside the resolved root is
reachable.
## License
Apache License 2.0 — see [LICENSE](LICENSE). Use it, fork it, ship it inside
a commercial product, whatever you need. If you build something interesting
on top of it, a mention is appreciated but never required.
## Author
**Luclino Cruz** — [github.com/luclinocruz](https://github.com/luclinocruz)
Built as part of the harness for [LuCross](https://github.com/luclinocruz),
a solo-operated, AI-agent-driven company. If this saved you a debugging
session, a star helps other people find it.
This server cannot be deployed
Maintenance
ActivitySlowing
ResponsivenessNo issues