Skip to main content
Glama

webear

npm version npm downloads License: MIT MCP Compatible

Give your AI real senses — hear, see, and feel any web app.

An MCP server + browser SDK that gives AI coding assistants direct sensory access to a live web application. Audio, visuals, performance, network, security, and console — captured from the browser, analyzed in real time, delivered via MCP.

"The beat sounds muddy" → your AI captures 3 seconds, measures the spectral centroid at 580 Hz with 45% energy below 250 Hz, and tells you exactly why.


AI Web Perception Demo


What It Does

Tool

Description

capture_audio

Record a short clip (500ms–30s) of what your web app is outputting right now

analyze_audio

Signal analysis: RMS, peak dB, clipping, spectral centroid, frequency bands, BPM, timing jitter

describe_audio

Plain-English AI description — "the kick is boomy with heavy sub buildup around 80 Hz"

diff_audio

Compare two captures and flag what changed — loudness, tone, timing, clipping

Related MCP server: broca-machina

How It Works

Browser (Web Audio API)
    ↓ MediaRecorder taps the AudioContext output node
    ↓ Uploads WebM blob via HTTP POST
Express Middleware (your dev server)
    ↓ Stores captures in memory, dispatches commands via SSE
MCP Server (stdio — runs inside your IDE)
    ↓ Retrieves captures, sends to CodedSwitch analysis API
AI Coding Assistant
    → "Your bass band is 42% of the mix (high), spectral centroid
       is 580 Hz (muddy), and timing jitter is 23ms — the scheduler
       is drifting under load."

The key difference from every other audio MCP: this taps the Web Audio graph directly, bypassing room acoustics, microphone hardware, and the need to export files.


Quick Start

1. Install

npm install webear

2. Add the Express middleware to your dev server

import express from 'express'
import { webearMiddleware } from 'webear/middleware'

const app = express()
app.use(express.json())

// Mount the audio debug bridge (automatically disabled in production)
app.use('/api/webear', webearMiddleware())

app.listen(5000)

3. Add the client snippet to your web app

Option A — auto-detect everything (Tone.js or raw Web Audio)

import WebEar from 'webear/client'
WebEar.init()

Option B — explicit AudioContext

const ctx = new AudioContext()
const masterGain = ctx.createGain()
masterGain.connect(ctx.destination)

WebEar.init({ audioContext: ctx, outputNode: masterGain })

Option C — Tone.js project

import * as Tone from 'tone'
WebEar.init({ toneJs: true })

Option D — Three.js WebGL Game

import * as THREE from 'three'
const listener = new THREE.AudioListener()
camera.add(listener)
WebEar.init({ tapNode: listener.getInput() })

Option E — plain script tag

<script src="node_modules/webear/client-snippet.js"></script>
<script>WebEar.init()</script>

4. Configure your IDE

Claude Code (.mcp.json in project root):

{
  "mcpServers": {
    "webear": {
      "command": "npx",
      "args": ["webear"],
      "env": {
        "WEBEAR_BASE_URL": "http://localhost:5000",
        "CODEDSWITCH_API_KEY": "your-key-here"
      }
    }
  }
}

Cursor (.cursor/mcp.json):

{
  "mcpServers": {
    "webear": {
      "command": "npx",
      "args": ["webear"],
      "env": {
        "WEBEAR_BASE_URL": "http://localhost:5000",
        "CODEDSWITCH_API_KEY": "your-key-here"
      }
    }
  }
}

Windsurf (mcp_config.json):

{
  "webear": {
    "command": "npx",
    "args": ["webear"],
    "disabled": false,
    "env": {
      "WEBEAR_BASE_URL": "http://localhost:5000",
      "CODEDSWITCH_API_KEY": "your-key-here"
    }
  }
}

5. Get an API key — optional, and not to start

analyze_audio works with no key and no account. If ffmpeg is on your PATH, it decodes and analyzes the capture on your machine and returns a basic report: duration, loudness, peak level and whether the audio is clipping. Nothing is uploaded. Try the tool before you sign up for anything.

A key unlocks the parts that need more than arithmetic:

No key

With key

capture_audio

analyze_audio

Basic — duration, loudness, peak, clipping (local)

Full — spectral centroid, band energy, crest factor, BPM, timing jitter

describe_audio — what it SOUNDS like

mix_coach — measured + heard

diff_audio — before/after

To get one:

  1. Create a free account at codedswitch.com.

  2. Go to codedswitch.com/developer (also in the account menu as Developer API).

  3. Click Generate API Key — that value is your CODEDSWITCH_API_KEY. Keys start with wbr_.

Free tier: 50 analyses/day. No credit card required.

6. Start your dev server, open your app, play audio, then ask your AI:

"Capture 3 seconds and tell me why the bass sounds muddy."

"Compare the audio before and after my last commit."

"Is there any clipping in the high-frequency range?"


Example Output

analyze_audio

── Audio Analysis Report ──────────────────────────────
Duration:          3.02s

── Loudness ─────────────────────────────────────────
RMS:               -12.4 dBFS
Peak:              -1.2 dBFS
Dynamic range:     11.2 dB
Crest factor:      3.63
Clipping:          none

── Tone ──────────────────────────────────────────────
Spectral centroid: 2847 Hz
DC offset:         0.00012 (ok)

── Frequency Bands ───────────────────────────────────
Sub  (20-80 Hz):   8.2%
Bass (80-250 Hz):  22.1%
Mid  (250-2k Hz):  38.4%
Hi-mid (2-6k Hz):  21.8%
High (6k+ Hz):     9.5%

── Rhythm ────────────────────────────────────────────
Estimated BPM:     92
Onset count:       12
Timing jitter:     4.2 ms std dev

── Summary ───────────────────────────────────────────
Loudness: -12.4 dBFS RMS, peak -1.2 dBFS. Tone: balanced (centroid 2847 Hz).
Band mix — sub: 8% | bass: 22% | mid: 38% | hi-mid: 22% | high: 10%.
Rhythm: estimated 92 BPM, 12 onsets detected. Timing: very tight (< 5 ms jitter).

diff_audio

── Audio Diff: a1b2c3d4… → e5f6g7h8… ──

── Loudness ──────────────────────────────────────────
  RMS: -14.2 dBFS → -12.4 dBFS  (+1.8 dBFS)
⚠ Peak: -3.1 dBFS → -0.2 dBFS  (+2.9 dBFS)
⚠ CLIPPING INTRODUCED — gain staging regression

── Tone ──────────────────────────────────────────────
⚠ Spectral centroid: 2847.0 Hz → 1920.0 Hz  (-927.0 Hz)

── Interpretation ────────────────────────────────────
A gain bug was introduced that causes clipping.
Tonal character changed noticeably — EQ or filter behaviour may have shifted.

Configuration

Environment Variables

Variable

Default

Description

WEBEAR_BASE_URL

http://localhost:4000

URL of your dev server (where middleware is mounted)

CODEDSWITCH_API_KEY

API key from codedswitch.com — required for analyze_audio and describe_audio

MCP_API_URL

https://www.codedswitch.com

Override the analysis API base (advanced / self-hosted)

Middleware Options

webearMiddleware({
  maxCaptures: 50,       // Max captures in memory (default: 50)
  maxAgeMins: 10,        // Auto-evict after N minutes (default: 10)
  maxUploadBytes: 50e6,  // Max upload size (default: 50MB)
  devOnly: true,         // Disable in production (default: true)
})

Client Options

WebEar.init({
  audioContext: myCtx,             // Your AudioContext instance
  outputNode: myGainNode,          // The node to tap (defaults to destination)
  toneJs: true,                    // Auto-detect Tone.js context
  bridgeBase: '/api/webear',  // Override API path
  devOnly: true,                   // Only init outside of production (default: true)
})

Requirements

  • Node.js >= 18

  • A browser that supports MediaRecorder (Chrome, Firefox, Edge, Safari 14+)

  • A CODEDSWITCH_API_KEY for analysis (free at codedswitch.com)


Who Is This For?

  • Web Audio / Tone.js developers — debug beats, synths, effects, and mixing without leaving your IDE

  • Game audio developers — verify sound effects, spatial audio, and mixing in real-time

  • Music app builders — catch regressions between code changes with diff_audio

  • Podcast / streaming apps — validate audio quality, levels, and encoding

  • Anyone whose app makes sound — if it has a Web Audio graph, your AI can now hear it


Why Not Just Use the Microphone?

Microphone MCPs capture room sound — your fan noise, chair creaks, and room reverb are all in the recording. webear taps the Web Audio API before it hits the DAC, giving you a clean digital signal with no room artifacts.


Web Perception — Full Sensor Suite

WebEar started as audio-only. Web Perception expands it to 6 senses:

Sensor

What it perceives

WebEar

Audio — mix quality, rhythm, instruments, clipping

WebEye

Visual — canvas, UI layout, animations, screenshots

WebSense

Performance — frame rate, memory, audio latency

WebNerve

Network — API latencies, connection quality, storage

WebShield

Security — cookies, storage exposure, CSP, framing

WebLog

Console — logs, warnings, errors, uncaught exceptions

Install the full browser SDK

import { WebPerception } from 'webear/perception'

WebPerception.init({
  apiKey: 'wbr_YOUR_API_KEY',
  relayUrl: 'https://www.codedswitch.com',
  sensors: ['ear', 'eye', 'sense', 'nerve', 'shield', 'log'],
})

Or use a single sensor:

import { WebEar } from 'webear/perception'

WebEar.init({
  apiKey: 'wbr_YOUR_API_KEY',
  ear: { audioContext: myCtx, audioNode: masterGain },
})

Connect via MCP (hosted relay — no local server required)

{
  "mcpServers": {
    "webear": {
      "url": "https://www.codedswitch.com/api/webear/mcp/sse",
      "headers": {
        "Authorization": "Bearer wbr_YOUR_API_KEY"
      }
    }
  }
}

Available MCP Tools

Sensor

Tool

Credits

Description

Ear

capture_audio

Free

Record live tab audio

Ear

analyze_audio

1

BPM, loudness, frequency bands, clipping, dynamic range

Ear

describe_audio

2

AI plain-English description — instruments, genre, mood, mix notes

Ear

diff_audio

1

Compare two captures — loudness, tone, timing deltas

Ear

groove_score

2

Grid alignment, swing factor, consistency (0–100%)

Ear

capture_and_analyze

1

Capture + analysis in one call

Ear

mix_coach

3

Structured mixing feedback

Eye

capture_video

Free

Record canvas/video from the tab

Eye

describe_video

2

AI visual description — layout, colors, bugs

Eye

diff_visuals

2

Compare two visual captures

Sense

capture_telemetry

Free

FPS, memory, layout shifts, audio latency

Sense

analyze_telemetry

1

Frame drops, memory pressure, audio underruns

Nerve

capture_nerve

Free

API timings, connection quality, storage size

Nerve

analyze_nerve

1

Slow APIs, connection quality, storage bloat

Shield

capture_shield

Free

Cookies, CSP, storage exposure, framing

Shield

analyze_shield

1

CORS issues, non-HttpOnly cookies, missing CSP

Log

capture_logs

Free

Console output + uncaught exceptions

Log

analyze_logs

1

Error patterns, stack traces, repeated warnings

Get an API Key

  1. Create a free account at codedswitch.com.

  2. Open codedswitch.com/developer — also linked as Developer API in the account menu.

  3. Click Generate API Key and copy it. Keys start with wbr_.

Free tier: 50 analyses/day, no credit card required.


Changelog

2.0.1

  • Fixed the getting-started path for API keys. The previous instruction ("Settings → WebEar") was wrong — there is no WebEar section under Settings. Keys live at codedswitch.com/developer (linked as Developer API in the account menu). Both the Quick Start and the Web Perception sections now point to the correct place.

  • The SDK's "missing API key" console error now links straight to the key page.

Contributing

See CONTRIBUTING.md.

License

MIT — see LICENSE

Author

Built by @asume21CodedSwitch

Available Tools

4 tools
analyze_audioA

Run signal analysis on a captured audio clip. Returns RMS, peak dB, clipping, spectral centroid, frequency band energy, estimated BPM, and timing jitter.

ParametersJSON Schema
NameRequiredDescriptionDefault
capture_idYesThe capture ID returned by capture_audio

TDQS

A3.9/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description fully carries the burden of behavioral disclosure. It lists the analysis outputs and implies a non-destructive, read-only operation, which is appropriate for a signal analysis tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise at one sentence, front-loaded with the action ('Run signal analysis'), and efficiently enumerates outputs without extraneous content.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given only one parameter and no output schema, the description adequately lists the return values, providing enough context for an agent to understand what the tool produces.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% for the single parameter (capture_id), which already has a description. The tool description adds that the parameter should come from capture_audio, but this is marginal value beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly specifies the verb ('Run signal analysis'), resource ('captured audio clip'), and lists specific metrics returned (RMS, peak dB, clipping, etc.), distinguishing it from sibling tools like capture_audio or describe_audio.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies the tool requires a prior captured audio clip (via capture_id), but does not explicitly state when to prefer this tool over alternatives like describe_audio or diff_audio, leaving room for ambiguity.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

capture_audioA

Record a short clip of what the running web app is currently outputting. Returns a capture ID you can pass to analyze_audio or describe_audio.

ParametersJSON Schema
NameRequiredDescriptionDefault
duration_msNoHow many milliseconds to record (default 3000, max 30000)

TDQS

A4.1/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden. It states 'short clip' and mentions the duration parameter, but does not disclose what audio source is captured (e.g., system output, microphone), potential failure modes, or any destructive effects.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, front-loaded with the verb and resource, no redundant information. Every sentence adds value.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple tool with one optional parameter and no output schema, the description fully covers the purpose, return value, and relationship to sibling tools. Nothing critical is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% (duration_ms fully documented). The main description provides no additional parameter semantics beyond what the schema already states (default 3000, max 30000). Baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Clearly states the tool records a short clip of the running web app's audio output and returns a capture ID. Distinguishes itself from sibling tools (analyze_audio, describe_audio, diff_audio) by specifying that the ID can be passed to them.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly says to use the capture ID with analyze_audio or describe_audio, implying a workflow. However, it does not provide explicit when-not-to-use scenarios or alternatives beyond the mentioned siblings.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

describe_audioA

Send a captured audio clip to Gemini or GPT-4o to get a plain-English description of what it sounds like — useful when something sounds wrong but you cannot describe it.

ParametersJSON Schema
NameRequiredDescriptionDefault
capture_idYesThe capture ID returned by capture_audio to describe

TDQS

A3.6/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description must fully disclose behavior. It mentions the AI models but omits side effects like cost, latency, or external dependencies. This is insufficient for a tool that sends audio to external APIs.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, efficient sentence that conveys purpose, usage context, and outcome without any wasted words.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's simplicity (one parameter, no output schema), the description covers the essential information. Minor gaps exist regarding output format or latency, but overall it is complete enough for the complexity level.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema already describes the single parameter with 100% coverage. The description adds minimal value beyond restating the purpose, but no further semantic information is needed given the schema's completeness.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the verb ('send'), resource ('captured audio clip'), and outcome ('plain-English description'). It differentiates from sibling tools like analyze_audio, capture_audio, and diff_audio by focusing on description generation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides a usage context ('useful when something sounds wrong but you cannot describe it') but does not explicitly mention when not to use it or compare to alternatives like analyze_audio.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

diff_audioA

Compare two audio captures and flag what changed — loudness, tone, timing, clipping. Use this before and after a code change to verify the audio impact.

ParametersJSON Schema
NameRequiredDescriptionDefault
capture_id_aYesFirst capture ID (the "before")
capture_id_bYesSecond capture ID (the "after")

TDQS

A3.9/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It states the tool 'flags what changed,' implying a read-only operation, but does not explicitly disclose whether it is safe, destructive, or requires permissions. Given the simple nature of a diff tool, the lack of explicit behavioral disclosure is a minor gap.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences: the first clearly states the purpose and what is flagged, the second gives a concise use case. No unnecessary words. Well-structured and efficient.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description covers the purpose and usage but lacks information about return values or potential errors. Given the tool's simplicity and the absence of an output schema, a hint at the output format would improve completeness. Sibling tools are explained in their own descriptions, so differentiation is adequate.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, with each parameter described as the 'before' and 'after' capture ID. The tool description adds context about comparing captures and the aspects examined, but does not significantly enhance the semantic meaning beyond the schema. With high schema coverage, baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool compares two audio captures and lists specific aspects (loudness, tone, timing, clipping). It distinguishes itself from siblings (capture_audio, analyze_audio, describe_audio) by focusing on comparative analysis and providing a concrete use case ('before and after a code change').

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly says when to use this tool: 'before and after a code change to verify the audio impact.' It provides clear context but does not explicitly state when not to use it or mention alternative tools, though the use case implies a comparison scenario.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

TDQS

A4.2/5.0
Disambiguation5/5

Each tool has a clearly distinct purpose: capture records audio, analyze does signal analysis, describe provides a language description, and diff compares two captures. No overlap.

Naming Consistency5/5

All tool names follow the consistent verb_noun pattern (capture_audio, analyze_audio, describe_audio, diff_audio), making them predictable and easy to understand.

Tool Count5/5

Four tools cover the essential operations for audio perception without being too few or too many, perfectly scoped for the server's purpose.

Completeness5/5

The tool set covers capture, analysis, description, and comparison, providing a complete workflow for assessing audio output. No obvious missing operations.

Maintenance

ActivityMaintained
ResponsivenessSyncing

Resources

Unclaimed servers have limited discoverability.

Looking for Admin?

If you are the server author, to access and configure the admin panel.

Related MCP Connectors

Related MCP Servers

Latest Blog Posts

MCP directory API

We provide all the information about MCP servers via our MCP API.

curl -X GET 'https://glama.ai/api/mcp/v1/servers/asume21/webear'

If you have feedback or need assistance with the MCP directory API, please join our Discord server