web perception
This server provides AI-powered audio perception tools for web applications, enabling detailed analysis and comparison of live audio output.
Capture audio: Record short audio clips (500ms–30s) from a web application's audio output, returning a unique capture ID.
Analyze audio: Perform signal analysis on a captured clip, obtaining metrics such as RMS, peak dB, clipping, spectral centroid, frequency band energy, estimated BPM, and timing jitter.
Describe audio: Generate a plain-English AI-powered description of a captured clip, useful for understanding complex audio characteristics or issues.
Compare audio clips (
diff_audio): Compare two captured clips to highlight changes in loudness, tone, timing, and clipping — ideal for identifying audio regressions after code changes.
Provides Express middleware to add audio capture endpoints to web applications, enabling the MCP server to tap into the Web Audio graph.
webear
Give your AI real senses — hear, see, and feel any web app.
An MCP server + browser SDK that gives AI coding assistants direct sensory access to a live web application. Audio, visuals, performance, network, security, and console — captured from the browser, analyzed in real time, delivered via MCP.
"The beat sounds muddy" → your AI captures 3 seconds, measures the spectral centroid at 580 Hz with 45% energy below 250 Hz, and tells you exactly why.

What It Does
Tool | Description |
| Record a short clip (500ms–30s) of what your web app is outputting right now |
| Signal analysis: RMS, peak dB, clipping, spectral centroid, frequency bands, BPM, timing jitter |
| Plain-English AI description — "the kick is boomy with heavy sub buildup around 80 Hz" |
| Compare two captures and flag what changed — loudness, tone, timing, clipping |
Related MCP server: broca-machina
How It Works
Browser (Web Audio API)
↓ MediaRecorder taps the AudioContext output node
↓ Uploads WebM blob via HTTP POST
Express Middleware (your dev server)
↓ Stores captures in memory, dispatches commands via SSE
MCP Server (stdio — runs inside your IDE)
↓ Retrieves captures, sends to CodedSwitch analysis API
AI Coding Assistant
→ "Your bass band is 42% of the mix (high), spectral centroid
is 580 Hz (muddy), and timing jitter is 23ms — the scheduler
is drifting under load."The key difference from every other audio MCP: this taps the Web Audio graph directly, bypassing room acoustics, microphone hardware, and the need to export files.
Quick Start
1. Install
npm install webear2. Add the Express middleware to your dev server
import express from 'express'
import { webearMiddleware } from 'webear/middleware'
const app = express()
app.use(express.json())
// Mount the audio debug bridge (automatically disabled in production)
app.use('/api/webear', webearMiddleware())
app.listen(5000)3. Add the client snippet to your web app
Option A — auto-detect everything (Tone.js or raw Web Audio)
import WebEar from 'webear/client'
WebEar.init()Option B — explicit AudioContext
const ctx = new AudioContext()
const masterGain = ctx.createGain()
masterGain.connect(ctx.destination)
WebEar.init({ audioContext: ctx, outputNode: masterGain })Option C — Tone.js project
import * as Tone from 'tone'
WebEar.init({ toneJs: true })Option D — Three.js WebGL Game
import * as THREE from 'three'
const listener = new THREE.AudioListener()
camera.add(listener)
WebEar.init({ tapNode: listener.getInput() })Option E — plain script tag
<script src="node_modules/webear/client-snippet.js"></script>
<script>WebEar.init()</script>4. Configure your IDE
Claude Code (.mcp.json in project root):
{
"mcpServers": {
"webear": {
"command": "npx",
"args": ["webear"],
"env": {
"WEBEAR_BASE_URL": "http://localhost:5000",
"CODEDSWITCH_API_KEY": "your-key-here"
}
}
}
}Cursor (.cursor/mcp.json):
{
"mcpServers": {
"webear": {
"command": "npx",
"args": ["webear"],
"env": {
"WEBEAR_BASE_URL": "http://localhost:5000",
"CODEDSWITCH_API_KEY": "your-key-here"
}
}
}
}Windsurf (mcp_config.json):
{
"webear": {
"command": "npx",
"args": ["webear"],
"disabled": false,
"env": {
"WEBEAR_BASE_URL": "http://localhost:5000",
"CODEDSWITCH_API_KEY": "your-key-here"
}
}
}5. Get an API key — optional, and not to start
analyze_audio works with no key and no account. If ffmpeg is on your PATH,
it decodes and analyzes the capture on your machine and returns a basic report:
duration, loudness, peak level and whether the audio is clipping. Nothing is
uploaded. Try the tool before you sign up for anything.
A key unlocks the parts that need more than arithmetic:
No key | With key | |
| ✓ | ✓ |
| Basic — duration, loudness, peak, clipping (local) | Full — spectral centroid, band energy, crest factor, BPM, timing jitter |
| — | ✓ |
| — | ✓ |
| — | ✓ |
To get one:
Create a free account at codedswitch.com.
Go to codedswitch.com/developer (also in the account menu as Developer API).
Click Generate API Key — that value is your
CODEDSWITCH_API_KEY. Keys start withwbr_.
Free tier: 50 analyses/day. No credit card required.
6. Start your dev server, open your app, play audio, then ask your AI:
"Capture 3 seconds and tell me why the bass sounds muddy."
"Compare the audio before and after my last commit."
"Is there any clipping in the high-frequency range?"
Example Output
analyze_audio
── Audio Analysis Report ──────────────────────────────
Duration: 3.02s
── Loudness ─────────────────────────────────────────
RMS: -12.4 dBFS
Peak: -1.2 dBFS
Dynamic range: 11.2 dB
Crest factor: 3.63
Clipping: none
── Tone ──────────────────────────────────────────────
Spectral centroid: 2847 Hz
DC offset: 0.00012 (ok)
── Frequency Bands ───────────────────────────────────
Sub (20-80 Hz): 8.2%
Bass (80-250 Hz): 22.1%
Mid (250-2k Hz): 38.4%
Hi-mid (2-6k Hz): 21.8%
High (6k+ Hz): 9.5%
── Rhythm ────────────────────────────────────────────
Estimated BPM: 92
Onset count: 12
Timing jitter: 4.2 ms std dev
── Summary ───────────────────────────────────────────
Loudness: -12.4 dBFS RMS, peak -1.2 dBFS. Tone: balanced (centroid 2847 Hz).
Band mix — sub: 8% | bass: 22% | mid: 38% | hi-mid: 22% | high: 10%.
Rhythm: estimated 92 BPM, 12 onsets detected. Timing: very tight (< 5 ms jitter).diff_audio
── Audio Diff: a1b2c3d4… → e5f6g7h8… ──
── Loudness ──────────────────────────────────────────
RMS: -14.2 dBFS → -12.4 dBFS (+1.8 dBFS)
⚠ Peak: -3.1 dBFS → -0.2 dBFS (+2.9 dBFS)
⚠ CLIPPING INTRODUCED — gain staging regression
── Tone ──────────────────────────────────────────────
⚠ Spectral centroid: 2847.0 Hz → 1920.0 Hz (-927.0 Hz)
── Interpretation ────────────────────────────────────
A gain bug was introduced that causes clipping.
Tonal character changed noticeably — EQ or filter behaviour may have shifted.Configuration
Environment Variables
Variable | Default | Description |
|
| URL of your dev server (where middleware is mounted) |
| — | API key from codedswitch.com — required for |
|
| Override the analysis API base (advanced / self-hosted) |
Middleware Options
webearMiddleware({
maxCaptures: 50, // Max captures in memory (default: 50)
maxAgeMins: 10, // Auto-evict after N minutes (default: 10)
maxUploadBytes: 50e6, // Max upload size (default: 50MB)
devOnly: true, // Disable in production (default: true)
})Client Options
WebEar.init({
audioContext: myCtx, // Your AudioContext instance
outputNode: myGainNode, // The node to tap (defaults to destination)
toneJs: true, // Auto-detect Tone.js context
bridgeBase: '/api/webear', // Override API path
devOnly: true, // Only init outside of production (default: true)
})Requirements
Node.js >= 18
A browser that supports
MediaRecorder(Chrome, Firefox, Edge, Safari 14+)A
CODEDSWITCH_API_KEYfor analysis (free at codedswitch.com)
Who Is This For?
Web Audio / Tone.js developers — debug beats, synths, effects, and mixing without leaving your IDE
Game audio developers — verify sound effects, spatial audio, and mixing in real-time
Music app builders — catch regressions between code changes with
diff_audioPodcast / streaming apps — validate audio quality, levels, and encoding
Anyone whose app makes sound — if it has a Web Audio graph, your AI can now hear it
Why Not Just Use the Microphone?
Microphone MCPs capture room sound — your fan noise, chair creaks, and room reverb are all in the recording. webear taps the Web Audio API before it hits the DAC, giving you a clean digital signal with no room artifacts.
Web Perception — Full Sensor Suite
WebEar started as audio-only. Web Perception expands it to 6 senses:
Sensor | What it perceives |
WebEar | Audio — mix quality, rhythm, instruments, clipping |
WebEye | Visual — canvas, UI layout, animations, screenshots |
WebSense | Performance — frame rate, memory, audio latency |
WebNerve | Network — API latencies, connection quality, storage |
WebShield | Security — cookies, storage exposure, CSP, framing |
WebLog | Console — logs, warnings, errors, uncaught exceptions |
Install the full browser SDK
import { WebPerception } from 'webear/perception'
WebPerception.init({
apiKey: 'wbr_YOUR_API_KEY',
relayUrl: 'https://www.codedswitch.com',
sensors: ['ear', 'eye', 'sense', 'nerve', 'shield', 'log'],
})Or use a single sensor:
import { WebEar } from 'webear/perception'
WebEar.init({
apiKey: 'wbr_YOUR_API_KEY',
ear: { audioContext: myCtx, audioNode: masterGain },
})Connect via MCP (hosted relay — no local server required)
{
"mcpServers": {
"webear": {
"url": "https://www.codedswitch.com/api/webear/mcp/sse",
"headers": {
"Authorization": "Bearer wbr_YOUR_API_KEY"
}
}
}
}Available MCP Tools
Sensor | Tool | Credits | Description |
Ear |
| Free | Record live tab audio |
Ear |
| 1 | BPM, loudness, frequency bands, clipping, dynamic range |
Ear |
| 2 | AI plain-English description — instruments, genre, mood, mix notes |
Ear |
| 1 | Compare two captures — loudness, tone, timing deltas |
Ear |
| 2 | Grid alignment, swing factor, consistency (0–100%) |
Ear |
| 1 | Capture + analysis in one call |
Ear |
| 3 | Structured mixing feedback |
Eye |
| Free | Record canvas/video from the tab |
Eye |
| 2 | AI visual description — layout, colors, bugs |
Eye |
| 2 | Compare two visual captures |
Sense |
| Free | FPS, memory, layout shifts, audio latency |
Sense |
| 1 | Frame drops, memory pressure, audio underruns |
Nerve |
| Free | API timings, connection quality, storage size |
Nerve |
| 1 | Slow APIs, connection quality, storage bloat |
Shield |
| Free | Cookies, CSP, storage exposure, framing |
Shield |
| 1 | CORS issues, non-HttpOnly cookies, missing CSP |
Log |
| Free | Console output + uncaught exceptions |
Log |
| 1 | Error patterns, stack traces, repeated warnings |
Get an API Key
Create a free account at codedswitch.com.
Open codedswitch.com/developer — also linked as Developer API in the account menu.
Click Generate API Key and copy it. Keys start with
wbr_.
Free tier: 50 analyses/day, no credit card required.
Changelog
2.0.1
Fixed the getting-started path for API keys. The previous instruction ("Settings → WebEar") was wrong — there is no WebEar section under Settings. Keys live at codedswitch.com/developer (linked as Developer API in the account menu). Both the Quick Start and the Web Perception sections now point to the correct place.
The SDK's "missing API key" console error now links straight to the key page.
Contributing
See CONTRIBUTING.md.
License
MIT — see LICENSE
Author
Built by @asume21 — CodedSwitch
Available Tools
4 toolsanalyze_audioA
Run signal analysis on a captured audio clip. Returns RMS, peak dB, clipping, spectral centroid, frequency band energy, estimated BPM, and timing jitter.
| Name | Required | Description | Default |
|---|---|---|---|
| capture_id | Yes | The capture ID returned by capture_audio |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description fully carries the burden of behavioral disclosure. It lists the analysis outputs and implies a non-destructive, read-only operation, which is appropriate for a signal analysis tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is concise at one sentence, front-loaded with the action ('Run signal analysis'), and efficiently enumerates outputs without extraneous content.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given only one parameter and no output schema, the description adequately lists the return values, providing enough context for an agent to understand what the tool produces.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% for the single parameter (capture_id), which already has a description. The tool description adds that the parameter should come from capture_audio, but this is marginal value beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly specifies the verb ('Run signal analysis'), resource ('captured audio clip'), and lists specific metrics returned (RMS, peak dB, clipping, etc.), distinguishing it from sibling tools like capture_audio or describe_audio.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies the tool requires a prior captured audio clip (via capture_id), but does not explicitly state when to prefer this tool over alternatives like describe_audio or diff_audio, leaving room for ambiguity.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
capture_audioA
Record a short clip of what the running web app is currently outputting. Returns a capture ID you can pass to analyze_audio or describe_audio.
| Name | Required | Description | Default |
|---|---|---|---|
| duration_ms | No | How many milliseconds to record (default 3000, max 30000) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It states 'short clip' and mentions the duration parameter, but does not disclose what audio source is captured (e.g., system output, microphone), potential failure modes, or any destructive effects.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, front-loaded with the verb and resource, no redundant information. Every sentence adds value.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple tool with one optional parameter and no output schema, the description fully covers the purpose, return value, and relationship to sibling tools. Nothing critical is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% (duration_ms fully documented). The main description provides no additional parameter semantics beyond what the schema already states (default 3000, max 30000). Baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Clearly states the tool records a short clip of the running web app's audio output and returns a capture ID. Distinguishes itself from sibling tools (analyze_audio, describe_audio, diff_audio) by specifying that the ID can be passed to them.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly says to use the capture ID with analyze_audio or describe_audio, implying a workflow. However, it does not provide explicit when-not-to-use scenarios or alternatives beyond the mentioned siblings.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
describe_audioA
Send a captured audio clip to Gemini or GPT-4o to get a plain-English description of what it sounds like — useful when something sounds wrong but you cannot describe it.
| Name | Required | Description | Default |
|---|---|---|---|
| capture_id | Yes | The capture ID returned by capture_audio to describe |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description must fully disclose behavior. It mentions the AI models but omits side effects like cost, latency, or external dependencies. This is insufficient for a tool that sends audio to external APIs.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, efficient sentence that conveys purpose, usage context, and outcome without any wasted words.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's simplicity (one parameter, no output schema), the description covers the essential information. Minor gaps exist regarding output format or latency, but overall it is complete enough for the complexity level.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema already describes the single parameter with 100% coverage. The description adds minimal value beyond restating the purpose, but no further semantic information is needed given the schema's completeness.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the verb ('send'), resource ('captured audio clip'), and outcome ('plain-English description'). It differentiates from sibling tools like analyze_audio, capture_audio, and diff_audio by focusing on description generation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides a usage context ('useful when something sounds wrong but you cannot describe it') but does not explicitly mention when not to use it or compare to alternatives like analyze_audio.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
diff_audioA
Compare two audio captures and flag what changed — loudness, tone, timing, clipping. Use this before and after a code change to verify the audio impact.
| Name | Required | Description | Default |
|---|---|---|---|
| capture_id_a | Yes | First capture ID (the "before") | |
| capture_id_b | Yes | Second capture ID (the "after") |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It states the tool 'flags what changed,' implying a read-only operation, but does not explicitly disclose whether it is safe, destructive, or requires permissions. Given the simple nature of a diff tool, the lack of explicit behavioral disclosure is a minor gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences: the first clearly states the purpose and what is flagged, the second gives a concise use case. No unnecessary words. Well-structured and efficient.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description covers the purpose and usage but lacks information about return values or potential errors. Given the tool's simplicity and the absence of an output schema, a hint at the output format would improve completeness. Sibling tools are explained in their own descriptions, so differentiation is adequate.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, with each parameter described as the 'before' and 'after' capture ID. The tool description adds context about comparing captures and the aspects examined, but does not significantly enhance the semantic meaning beyond the schema. With high schema coverage, baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool compares two audio captures and lists specific aspects (loudness, tone, timing, clipping). It distinguishes itself from siblings (capture_audio, analyze_audio, describe_audio) by focusing on comparative analysis and providing a concrete use case ('before and after a code change').
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly says when to use this tool: 'before and after a code change to verify the audio impact.' It provides clear context but does not explicitly state when not to use it or mention alternative tools, though the use case implies a comparison scenario.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
TDQS
Each tool has a clearly distinct purpose: capture records audio, analyze does signal analysis, describe provides a language description, and diff compares two captures. No overlap.
All tool names follow the consistent verb_noun pattern (capture_audio, analyze_audio, describe_audio, diff_audio), making them predictable and easy to understand.
Four tools cover the essential operations for audio perception without being too few or too many, perfectly scoped for the server's purpose.
The tool set covers capture, analysis, description, and comparison, providing a complete workflow for assessing audio output. No obvious missing operations.
Maintenance
Resources
Unclaimed servers have limited discoverability.
Looking for Admin?
If you are the server author, to access and configure the admin panel.
Related MCP Connectors
Give an AI agent a body in a Zero 3D voxel world: perceive, move, build, chat, remember.
Give your AI a face, a voice, and a personality. 3D avatars with custom personas.
OCR, transcription, file extraction, and image generation for AI agents via MCP.
Give your AI agents the tools to build, manage, and run automation workflows.
Related MCP Servers
- AlicenseAqualityAmaintenanceGive your AI assistant eyes and ears — analyze any video, audio, or image, entirely on your machine.2802Apache 2.0
- AlicenseNot gradedqualityAmaintenanceGives any text-based AI a voice and ears inside a Discord voice channel by transcribing speech, relaying to an LLM/agent, and speaking replies back.MIT
- AlicenseNot gradedqualityAmaintenanceGive your AI agents the ability to listen. Microphone capture and speech-to-text tools for MCP-compatible agents.1078Apache 2.0
- AlicenseNot gradedqualityDmaintenanceEnables AI systems to control the Reachy Mini robot—speak, listen, see, and express emotions through physical movement. Compatible with Claude, GPT, Grok, and other MCP-compatible AIs.MIT
Latest Blog Posts
- Who's Calling? MCP Hosts Are an Identity Blind Spot (And the Spec Knows It)By Om-Shree-0709 on .mcpAgent IdentityOAuth 2.1
- Your AI Chatbot Just Exposed Your CEO's Salary to an InternBy Om-Shree-0709 on .Agent IdentityMCP SecurityOAuth Delegation
- Why MCP Servers Need Execution Sandboxing (And Why Your Current Stack Isn't Enough)By Om-Shree-0709 on .Agentic AiPrompt InjectionWebAssembly
MCP directory API
We provide all the information about MCP servers via our MCP API.
curl -X GET 'https://glama.ai/api/mcp/v1/servers/asume21/webear'
If you have feedback or need assistance with the MCP directory API, please join our Discord server