Skip to main content
Glama
README.md
# meeting-transcriber-mcp

A [Model Context Protocol](https://modelcontextprotocol.io) stdio server for
[Meeting Transcriber](https://github.com/pasrom/meeting-transcriber), the local-first meeting
transcriber for macOS. It lets an agent submit an audio file, read the diarized transcript, name
the speakers, and turn meeting detection on or off — without any of the audio leaving the Mac.

This is a thin client over the app's own
[Local Automation API](https://github.com/pasrom/meeting-transcriber/blob/main/docs/automation-api.md),
a localhost-only HTTP surface the app ships for exactly this purpose. No audio, models or
transcripts pass through any cloud service: the transcription runs on the Apple Neural Engine
and this server only moves file paths and text over the loopback interface.

## Requirements

- macOS 14.2+ with Meeting Transcriber installed from **Homebrew** or built from source. The
  automation API is compiled out of the App Store variant, because that sandbox forbids the
  `network.server` entitlement the listener needs.
- Node.js 20 or newer.
- The automation API turned on. It is off by default:
  - **Settings → Advanced → "Local Automation API"** for a persistent toggle, or
  - launch the app with `MEETINGTRANSCRIBER_DEBUG_RPC=1` for one session.

The app writes a bearer token to
`~/Library/Application Support/MeetingTranscriber/.rpc-token` (mode `0600`) on first launch of
the API. This server reads it from there, so there is nothing to configure by hand. The token
rotates whenever the API is toggled off and on; the server re-reads the file and retries once
when the app rejects a stale copy.

## Setup

### Cursor

Add to `~/.cursor/mcp.json`:

```json
{
  "mcpServers": {
    "meeting-transcriber": {
      "command": "npx",
      "args": ["-y", "meeting-transcriber-mcp"]
    }
  }
}
```

### Claude Code

```bash
claude mcp add meeting-transcriber -- npx -y meeting-transcriber-mcp
```

### From source

```bash
git clone https://github.com/msvargas/meeting-transcriber-mcp.git
cd meeting-transcriber-mcp
npm install && npm run build
```

Then point the `command` at `node` with `args` of `["/absolute/path/to/dist/index.js"]`.

## Tools

| Tool | What it does |
| --- | --- |
| `transcribe_file` | Submit one file and wait for the diarized transcript. Runs headless, so a multi-speaker recording finishes on its own with auto-assigned names |
| `enqueue_files` | Queue one or more files and get job ids back immediately |
| `get_job` | Read a job's state, result paths, transcript and optionally the generated Markdown protocol |
| `get_naming` | Read the speaker labels a job is waiting to have named, with the app's suggestions |
| `confirm_naming` | Assign real names to diarization labels so a parked job can finish |
| `skip_naming` | Let a job finish with the names the app assigned itself |
| `get_watch_status` | Read whether the app is watching for meetings, and whether its permissions are healthy |
| `set_watch` | Start or stop automatic meeting detection |

### Two deliberate omissions

**No microphone recording control.** The app's API can start and stop a microphone-only
recording for an in-person meeting, and this server does not expose it. Nobody in the room can
see an agent decide to record them, and the app's own docs note that a start can raise a
permission dialog "unannounced". Drive that from the menu bar or a Stream Deck key instead.

**No `toggle` on `set_watch`.** A toggle applies a delta to a state the caller cannot see
reliably, so if the meeting ended or somebody used the menu bar in between, it does the opposite
of what was intended and stays inverted. `start` and `stop` express the desired end state and
converge no matter what happened before.

## Configuration

Every variable is optional.

| Variable | Default | Purpose |
| --- | --- | --- |
| `MEETING_TRANSCRIBER_BASE_URL` | `http://127.0.0.1:9876` | Where the app's API listens |
| `MEETING_TRANSCRIBER_TOKEN_PATH` | The app's token file under `Application Support` | Read the bearer token from somewhere else |
| `MEETING_TRANSCRIBER_TOKEN` | unset | Pin the token directly instead of reading a file. Disables the re-read-on-401 recovery |
| `MEETING_TRANSCRIBER_TIMEOUT_MS` | `30000` | Budget for the short endpoints. `transcribe_file` derives its own from `maxWaitSeconds` |

## Reading the results

Two fields are easy to misread, so the server spells them out in prose.

**An absent echo verdict is not a clean one.** On a loudspeaker recording the far end lands on
the microphone track too, and the app measures that before transcribing. When the verdict is
missing, nothing was measured — the job was single-source, a track was silent, or the tracks
overlapped for less than one analysis window. Only an explicit "not detected" means analysed and
clean, and this server never collapses the two.

**A job in `error` is not always final.** A user can retry it from the menu bar, which moves the
same id back to `waiting`. A poller that sees `error` and keeps polling may well watch the job
run again and end in `done`.

## Inherited limitations

These come from the app's API, not from this server:

- **Polling only.** There is no webhook or push channel, so a client polls `get_job`. On
  loopback a 3–5 second interval costs effectively nothing.
- **No upload.** A path must already be readable on the Mac running the app. Submitting a file
  that only exists on another machine fails.
- **The speaker database is read-only on this path.** Already-enrolled voices are recognized,
  but no endpoint here enrolls new ones.
- **Starting to watch can raise a macOS permission prompt.** Grant microphone and screen
  recording once interactively before relying on `set_watch`.

## Development

```bash
npm run typecheck
npm test          # 26 tests, no running app required: fetch is injected
npm run check     # typecheck + tests + repository hygiene
npm run build
bash scripts/smoke.sh   # drives the built server over stdio against a real app
```

The tests connect a real MCP client to the server over an in-memory transport and stub the HTTP
layer, so they cover the tool schemas, the status-code mapping and the rendering without
touching Meeting Transcriber.

## License

MIT

TDQS

A4.1/5.0

Scored across 8 tools

Disambiguation4/5

Most tools have clear distinct purposes: the naming cluster (get_naming/confirm_naming/skip_naming) and the watch cluster (set_watch/get_watch_status) are separable, and get_job is unique. transcribe_file vs enqueue_files share the same transcription goal and differ mainly in blocking behavior, which is the one spot an agent could misselect, though descriptions clarify it well.

Naming Consistency5/5

All tools follow a consistent snake_case verb_noun pattern (transcribe_file, enqueue_files, get_job, set_watch, get_naming, confirm_naming, skip_naming, get_watch_status). The get_/set_/confirm_/skip_ prefixes are used predictably.

Tool Count5/5

Eight tools is well-scoped for a transcription-and-watch domain, with no redundant or filler endpoints. Each tool maps to a distinct capability (sync/async transcription, job polling, naming resolution, watch control and status).

Completeness4/5

The surface covers the core lifecycle: submit, queue, poll, resolve naming, and watch automation, and get_job returns transcript results. Minor gaps exist (no job cancellation, no list/enumerate jobs, no speaker enrollment), but the descriptions explicitly flag some of these as intentional and workflows remain achievable.

Maintenance

ActivityMaintained
ResponsivenessNo issues