Skip to main content
Glama

meeting-transcriber-mcp

A Model Context Protocol stdio server for Meeting Transcriber, the local-first meeting transcriber for macOS. It lets an agent submit an audio file, read the diarized transcript, name the speakers, and turn meeting detection on or off — without any of the audio leaving the Mac.

This is a thin client over the app's own Local Automation API, a localhost-only HTTP surface the app ships for exactly this purpose. No audio, models or transcripts pass through any cloud service: the transcription runs on the Apple Neural Engine and this server only moves file paths and text over the loopback interface.

Requirements

  • macOS 14.2+ with Meeting Transcriber installed from Homebrew or built from source. The automation API is compiled out of the App Store variant, because that sandbox forbids the network.server entitlement the listener needs.

  • Node.js 20 or newer.

  • The automation API turned on. It is off by default:

    • Settings → Advanced → "Local Automation API" for a persistent toggle, or

    • launch the app with MEETINGTRANSCRIBER_DEBUG_RPC=1 for one session.

The app writes a bearer token to ~/Library/Application Support/MeetingTranscriber/.rpc-token (mode 0600) on first launch of the API. This server reads it from there, so there is nothing to configure by hand. The token rotates whenever the API is toggled off and on; the server re-reads the file and retries once when the app rejects a stale copy.

Related MCP server: ParrotScribe MCP Server

Setup

Cursor

Add to ~/.cursor/mcp.json:

{
  "mcpServers": {
    "meeting-transcriber": {
      "command": "npx",
      "args": ["-y", "meeting-transcriber-mcp"]
    }
  }
}

Claude Code

claude mcp add meeting-transcriber -- npx -y meeting-transcriber-mcp

From source

git clone https://github.com/msvargas/meeting-transcriber-mcp.git
cd meeting-transcriber-mcp
npm install && npm run build

Then point the command at node with args of ["/absolute/path/to/dist/index.js"].

Tools

Tool

What it does

transcribe_file

Submit one file and wait for the diarized transcript. Runs headless, so a multi-speaker recording finishes on its own with auto-assigned names

enqueue_files

Queue one or more files and get job ids back immediately

get_job

Read a job's state, result paths, transcript and optionally the generated Markdown protocol

get_naming

Read the speaker labels a job is waiting to have named, with the app's suggestions

confirm_naming

Assign real names to diarization labels so a parked job can finish

skip_naming

Let a job finish with the names the app assigned itself

get_watch_status

Read whether the app is watching for meetings, and whether its permissions are healthy

set_watch

Start or stop automatic meeting detection

Two deliberate omissions

No microphone recording control. The app's API can start and stop a microphone-only recording for an in-person meeting, and this server does not expose it. Nobody in the room can see an agent decide to record them, and the app's own docs note that a start can raise a permission dialog "unannounced". Drive that from the menu bar or a Stream Deck key instead.

No toggle on set_watch. A toggle applies a delta to a state the caller cannot see reliably, so if the meeting ended or somebody used the menu bar in between, it does the opposite of what was intended and stays inverted. start and stop express the desired end state and converge no matter what happened before.

Configuration

Every variable is optional.

Variable

Default

Purpose

MEETING_TRANSCRIBER_BASE_URL

http://127.0.0.1:9876

Where the app's API listens

MEETING_TRANSCRIBER_TOKEN_PATH

The app's token file under Application Support

Read the bearer token from somewhere else

MEETING_TRANSCRIBER_TOKEN

unset

Pin the token directly instead of reading a file. Disables the re-read-on-401 recovery

MEETING_TRANSCRIBER_TIMEOUT_MS

30000

Budget for the short endpoints. transcribe_file derives its own from maxWaitSeconds

Reading the results

Two fields are easy to misread, so the server spells them out in prose.

An absent echo verdict is not a clean one. On a loudspeaker recording the far end lands on the microphone track too, and the app measures that before transcribing. When the verdict is missing, nothing was measured — the job was single-source, a track was silent, or the tracks overlapped for less than one analysis window. Only an explicit "not detected" means analysed and clean, and this server never collapses the two.

A job in error is not always final. A user can retry it from the menu bar, which moves the same id back to waiting. A poller that sees error and keeps polling may well watch the job run again and end in done.

Inherited limitations

These come from the app's API, not from this server:

  • Polling only. There is no webhook or push channel, so a client polls get_job. On loopback a 3–5 second interval costs effectively nothing.

  • No upload. A path must already be readable on the Mac running the app. Submitting a file that only exists on another machine fails.

  • The speaker database is read-only on this path. Already-enrolled voices are recognized, but no endpoint here enrolls new ones.

  • Starting to watch can raise a macOS permission prompt. Grant microphone and screen recording once interactively before relying on set_watch.

Development

npm run typecheck
npm test          # 26 tests, no running app required: fetch is injected
npm run check     # typecheck + tests + repository hygiene
npm run build
bash scripts/smoke.sh   # drives the built server over stdio against a real app

The tests connect a real MCP client to the server over an in-memory transport and stub the HTTP layer, so they cover the tool schemas, the status-code mapping and the rendering without touching Meeting Transcriber.

License

MIT

Available Tools

8 tools
confirm_namingA

Confirm the real names behind a job's diarization labels so it can finish. A 409 means the job exists but is no longer awaiting naming. This does not enroll new voices in the app's speaker database; no endpoint here does.

ParametersJSON Schema
NameRequiredDescriptionDefault
jobIDYesThe job awaiting speaker naming.
mappingYesDiarization label to real name, for example {"Speaker 1": "Alice"}. Read the labels from get_naming first.

TDQS

A3.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations cover the safety profile (readOnlyHint=false, idempotentHint=false, openWorldHint=false) but say nothing about errors or side effects. The description adds real behavioral context beyond that: a 409 means the job exists but is no longer awaiting naming, and it explicitly scopes out voice-database enrollment.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short sentences, front-loaded with the action and its effect, then error semantics, then scope exclusion. No sentence is redundant and nothing is buried.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a two-parameter mutation with no output schema and annotations that cover the safety hints, the description supplies the missing pieces an agent needs: the 409 condition and the explicit non-enrollment boundary. It does not say what a successful response looks like, but that gap is minor.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and already explains both jobID and the label-to-name mapping, including 'Read the labels from get_naming first.' The description reinforces the diarization-label framing but adds no syntax or format detail beyond the schema, so the baseline of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('confirm the real names behind a job's diarization labels') and the consequence ('so it can finish'), which tells the agent exactly what effect the call has. It is clear but never names the sibling it pairs with (get_naming), so discrimination from skip_naming rests on the schema rather than the description.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is implied by 'a job's diarization labels' and the 409 note, which tells the agent the call only applies while a job is awaiting naming. There is no explicit 'call this after get_naming' or 'use skip_naming instead if' guidance in the description itself.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

enqueue_filesA
Idempotent

Queue one or more audio files and return their job ids immediately. Poll each one with get_job. Unlike transcribe_file these jobs can park on the speaker-naming step, which you resolve with confirm_naming or skip_naming.

ParametersJSON Schema
NameRequiredDescriptionDefault
pathsYesAbsolute paths to audio or video files readable on the Mac running Meeting Transcriber.
idempotencyKeyNoReuse the same key when retrying so the repeat returns the original job ids instead of enqueuing duplicates.

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare readOnlyHint=false and idempotentHint=true, and the description adds genuinely non-obvious behavior: the call is asynchronous and jobs can 'park' on a speaker-naming step requiring a separate resolution. It stops short of covering failure modes or what happens with duplicate/invalid paths, so it is strong but not exhaustive.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short sentences, front-loaded with the core action and outcome, then the follow-up workflow and the sibling contrast. No filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description still states what comes back (job ids) and outlines the full downstream lifecycle (poll, resolve naming). Nothing essential for invoking and using the result is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so both parameters (paths, idempotencyKey) are already documented in the schema. The description only reinforces plurality ('one or more') and adds no syntax or format detail beyond that, matching the baseline for schema-documented params.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Names a specific verb+resource ('Queue one or more audio files') and states the immediate outcome (returns job ids). It explicitly differentiates itself from the sibling transcribe_file, so an agent can route without opening either schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives explicit next steps (poll with get_job) and the alternative resolution paths (confirm_naming or skip_naming), plus a direct contrast with transcribe_file. This is when-to-use and what-to-do-next guidance, not just implied context.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_jobA
Read-only

Read a job's state, result paths and transcript. Answers for live jobs and for finished ones the app has already reaped. A 404 means the id was never enqueued, aged out of the app's terminal store, or was cancelled before finishing. A job in error is not always final: a user can retry it from the menu bar, which moves the same id back to waiting.

ParametersJSON Schema
NameRequiredDescriptionDefault
jobIDYesThe job id to read.
includeProtocolNoAlso read the generated Markdown protocol from the path the app reported. Defaults to false.
includeTranscriptNoFold the transcript text into the answer. Defaults to true. A job that is still running has no transcript yet.

TDQS

A4.5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, and the description adds substantial non-obvious behavior: 404 causes (never enqueued, aged out, cancelled before finishing) and the fact that an errored job is not always final because a manual retry moves the same id back to waiting. That lifecycle detail is exactly the kind of context annotations cannot carry.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences, front-loaded with the core purpose, then increasingly specific lifecycle facts. No filler, no restatement of the name, and each sentence carries distinct operational value.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description compensates by naming what is returned (state, result paths, transcript) and by explaining error/retry and 404 semantics. For a read-only three-parameter tool, nothing an agent needs to call it correctly appears to be missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so every parameter including includeProtocol and includeTranscript is already documented, and the schema itself notes that running jobs have no transcript. The description repeats that point but adds no new syntax, format, or default information beyond the schema, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Read a job's state, result paths and transcript') and enumerates what the read returns. This clearly separates it from siblings like enqueue_files (creates) and transcribe_file (starts work), so an agent can pick it without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explains the contexts in which reads succeed (live jobs and jobs the app has already reaped) and what a 404 implies, which tells the agent when a result is meaningful. It does not name an alternative tool or an explicit when-not-to-call condition, so it stops short of full routing guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_namingA
Read-only

Read the speaker labels a job is waiting to have named, with the app's own suggestion and how long each one spoke. Only meaningful while the job is in speakerNamingPending; any other state answers 404. Voice embeddings and audio are deliberately not exposed.

ParametersJSON Schema
NameRequiredDescriptionDefault
jobIDYesThe job awaiting speaker naming.

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already establish readOnlyHint=true and openWorldHint=false, and the description adds real value beyond them: the state-dependent 404 behavior and the deliberate omission of voice embeddings and audio. It stops short of describing response shape or whether the pending set is paginated, but the safety and failure profile is well covered.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three compact sentences: what is returned, when it is valid, and what is intentionally excluded. The scoping constraint is front-loaded and no sentence is filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description carries the return-value burden and does so by naming the suggestion and per-speaker duration, and it covers the failure mode and exclusions. For a single-parameter read tool, an agent has everything needed to call it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so jobID's type and requirement are already documented, which sets the baseline at 3. The description adds meaning beyond the schema by stating that the job must be in the speakerNamingPending state for the parameter to be usable, tying the argument to a runtime precondition.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (read) and resource (speaker labels awaiting naming) plus the payload contents (app suggestion, per-speaker duration). It is clearly distinguished from the sibling writers confirm_naming/skip_naming without needing to name them, since the pending-read framing implies the write counterparts.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly scopes usage: only meaningful while the job is in speakerNamingPending, and other states return 404. That is a clear use condition, though it never names an alternative tool (e.g., confirm_naming to finalize) or what to do instead when the state is wrong.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_watch_statusA
Read-only

Read whether Meeting Transcriber is watching for Teams, Zoom and Webex meetings, what the menu bar badge shows, and whether its permissions are healthy. Cheap enough to poll, and it answers even while the app is still starting up, which also makes it the liveness check for the automation API.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already establish readOnlyHint=true and openWorldHint=false, so safety is covered. The description adds real context beyond that: polling is cheap, the call succeeds even while the app is starting up, and it doubles as an API liveness check — genuinely useful operational traits.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, both front-loaded with the return contents first and the operational guidance second. Slightly dense in the second clause, but every phrase (poll cost, startup availability, liveness role) carries information an agent would otherwise lack.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no input parameters and no output schema, the description carries the burden of describing what comes back — and it does, listing watch state, badge, and permission health. Nothing needed to call this zero-arg read tool correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so there is nothing for the description to disambiguate; baseline 4 applies. Schema coverage is 100% with an empty object schema, leaving no gaps.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific read verb and enumerates the exact resource facets returned: watch state for Teams/Zoom/Webex, menu bar badge, and permission health. This is immediately separable from set_watch and the transcription siblings, which operate on jobs and files rather than app state.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives clear usage context ('cheap enough to poll', usable as a liveness check for the automation API) and notes it works during startup. It does not explicitly name set_watch as the mutating counterpart, so the alternative routing is implied rather than stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

set_watchA
Idempotent

Start or stop automatic meeting detection, the same thing the menu bar's Start Watching item does. A 409 means a manual recording owns the watch loop. The first start on a fresh install can raise a macOS microphone or screen-recording prompt that somebody has to answer, so grant those once interactively before relying on this.

ParametersJSON Schema
NameRequiredDescriptionDefault
actionYesThe state you want the app to end up in. Toggle is deliberately not offered: it applies a delta to a state the caller cannot see reliably, so it inverts silently when the meeting ended or someone used the menu bar in between.

TDQS

A4.5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations cover the safety profile (readOnlyHint=false, idempotentHint=true, openWorldHint=false), and the description adds genuinely new behavioral context: a 409 signals that a manual recording owns the watch loop, and a first start on a fresh install can trigger a macOS microphone/screen-recording prompt requiring interactive grant. These are exactly the failure modes an agent cannot infer from the structured fields.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three tight sentences, front-loaded with the core action and followed by error and permission context in priority order. No filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema exists, but for a one-parameter toggle the description covers the operation, the key error case, and the external permission prerequisite. Nothing an agent needs to invoke it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and the single 'action' enum is richly documented in the schema itself (including why toggle is omitted). The description adds nothing beyond that, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (start/stop) and resource (automatic meeting detection / watch loop), and anchors it to a concrete UI equivalent ('the same thing the menu bar's Start Watching item does'). An agent can separate this from get_watch_status (read-only status) without opening either schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives clear context for invocation (start vs stop, the 409 conflict case, and the interactive-permission prerequisite before relying on it). It does not explicitly route to a sibling alternative such as get_watch_status for checking state first, so it stops short of full when/when-not guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

skip_namingA

Let a job finish with the speaker names the app assigned itself. A 409 means the job exists but is no longer awaiting naming.

ParametersJSON Schema
NameRequiredDescriptionDefault
jobIDYesThe job awaiting speaker naming.

TDQS

A3.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The annotations already indicate this is a non-read-only, non-idempotent operation, so the description adds useful behavioral detail by explaining that a 409 response means the job exists but is no longer awaiting naming. It does not disclose auth requirements, but the error-state semantics go beyond the structured annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two concise sentences with no wasted words. The main effect is front-loaded, and the 409 detail follows as a useful secondary point.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a one-parameter action tool with annotations and full schema coverage, the description covers the core action and one notable error state. It remains thin on workflow context, particularly how this tool relates to sibling tools such as confirm_naming and get_naming.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, and the single parameter jobID is already defined as 'The job awaiting speaker naming.' The description does not add any syntax, format, or additional meaning beyond the schema, so the baseline of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific effect — allowing a job to finish using the app-assigned speaker names — which is a clear skip-naming action. It is understandable apart from the tool name, though it does not explicitly contrast itself with similar naming tools like confirm_naming.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies the tool is used when the caller wants to accept auto-assigned names, but it gives no explicit when-to-use condition or alternative. It also does not tell the agent when to use confirm_naming or get_naming instead.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

transcribe_fileA
Idempotent

Submit one audio or video file to Meeting Transcriber and wait for the diarized transcript. Runs headless, so a multi-speaker recording finishes on its own with auto-assigned speaker names instead of parking on the naming step. Use enqueue_files plus get_job when you do not want to block, or when you want to assign speaker names yourself.

ParametersJSON Schema
NameRequiredDescriptionDefault
pathYesAbsolute path to an audio or video file readable on the Mac running Meeting Transcriber. There is no upload: a path that exists only on another machine is rejected.
idempotencyKeyNoReuse the same key when retrying so the repeat returns the original job instead of enqueuing a duplicate. The app dedupes sequential retries within one session of its automation API.
maxWaitSecondsNoHow long the app blocks before answering. Defaults to 600. Use 0 to enqueue and get the current status back immediately.
includeTranscriptNoFold the transcript text into the answer instead of returning only a file path. Defaults to true, which is what you want unless the transcript is large and you only need the paths.

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Adds substantial behavior context beyond annotations: 'runs headless,' auto-assigned speaker names, no parking on the naming step, and the idempotency/dedup semantics are reinforced. Annotations declare idempotentHint=true, which the description aligns with rather than contradicts.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences, no filler, front-loaded with the primary action, then behavioral note, then alternative routing. Every sentence earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Completes the picture for a mutation-side submitter: how it behaves (headless, speaker naming), how it differs from siblings, and timeout semantics via schema. No output schema, so it does not need to describe return shape, and the schema handles return-adjacent details like includeTranscript.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so all four parameters are already documented in the schema, including the path-absolute/local-machine restriction and idempotency reuse guidance. The description adds little param-specific syntax or format detail beyond what the schema provides; baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a precise verb+resource+scope: submit one audio/video file to Meeting Transcriber and wait for the diarized transcript. Clearly contrasts with enqueue_files and get_job as the non-blocking alternative.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly names enqueue_files plus get_job as the alternative and gives the condition (not wanting to block, or wanting to assign speaker names yourself). It does not cover when blocking itself is inappropriate, but the routing guidance is clear.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 8 tool updatesv0.1.0
    • First observedconfirm_naming
    • First observedenqueue_files
    • First observedget_job
    • First observedget_naming
    • First observedget_watch_status
    • First observedset_watch
    • First observedskip_naming
    • First observedtranscribe_file

TDQS

A4.1/5.0

Scored across 8 tools

Disambiguation4/5

Most tools have clear distinct purposes: the naming cluster (get_naming/confirm_naming/skip_naming) and the watch cluster (set_watch/get_watch_status) are separable, and get_job is unique. transcribe_file vs enqueue_files share the same transcription goal and differ mainly in blocking behavior, which is the one spot an agent could misselect, though descriptions clarify it well.

Naming Consistency5/5

All tools follow a consistent snake_case verb_noun pattern (transcribe_file, enqueue_files, get_job, set_watch, get_naming, confirm_naming, skip_naming, get_watch_status). The get_/set_/confirm_/skip_ prefixes are used predictably.

Tool Count5/5

Eight tools is well-scoped for a transcription-and-watch domain, with no redundant or filler endpoints. Each tool maps to a distinct capability (sync/async transcription, job polling, naming resolution, watch control and status).

Completeness4/5

The surface covers the core lifecycle: submit, queue, poll, resolve naming, and watch automation, and get_job returns transcript results. Minor gaps exist (no job cancellation, no list/enumerate jobs, no speaker enrollment), but the descriptions explicitly flag some of these as intentional and workflows remain achievable.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • A
    license
    A
    quality
    F
    maintenance
    Enables AI agents to interact with the ParrotScribe transcription service on macOS, providing tools to start/stop transcription, retrieve real-time and historical transcripts, and search across sessions.
    8
    9 npm
    MIT
  • A
    license
    A
    quality
    C
    maintenance
    Connects to the TypeWhisper macOS app to let coding agents transcribe local files, inspect model status, search history, and manage dictionary terms and corrections.
    10
    14 npm
    2
    GPL 3.0
  • A
    license
    A
    quality
    B
    maintenance
    Enables automated audio restoration, transcription, and speaker diarization via MCP tools for queuing files, monitoring progress, and retrieving speaker-labeled transcripts.
    9
    MIT