Spoken
This server provides podcast transcript retrieval and search capabilities via three tools:
search_podcasts: Find published podcast episodes using a free-text query (e.g. "huberman sleep") or a direct episode URL (Spotify, YouTube, etc.). Returns episode metadata including
id,title,podcast, anddate. Free to use — no credits consumed.get_transcript: Fetch a full podcast episode transcript as clean Markdown, complete with real speaker names (not "Speaker 1") and timestamps. Requires an episode
idfromsearch_podcasts. Costs 1 credit on the first fetch; repeat fetches are free and errors are never charged.get_balance: Check your current credit balance, account email, and recent usage history for your configured API key. Free to use — no credits consumed.
Spoken — podcast transcripts as clean Markdown, built for AI agents
Spoken is a transcript API that turns any published podcast into clean Markdown with real speaker names — not "Speaker 1." One API call returns named, timestamped text, ready for LLMs, RAG pipelines, summarizers, and search.
It's a transcript retrieval API, not a speech-to-text service: it works on already-published podcasts, so you skip uploading audio, running diarization, and mapping anonymous speaker labels by hand. For published shows that's typically 5–10× cheaper than running the audio through a transcription service.
🎙️ Real speaker names, resolved automatically
📄 Clean Markdown with timestamps, tuned for LLM context windows and RAG chunking
🔎 Search by text query or paste a Spotify/YouTube URL
💳 Pay-per-use credits — no subscription, failed calls never charged, repeat fetches free
🤖 Agent-native — ships with an Agent Skill,
agents.md,llms.txt, and an OpenAPI spec
Get a key at spoken.md — or try it free with the demo key pt_demo (search works fully; transcripts limited to the demo episode).
Quickstart
# 1. Find an episode (by text, or paste a Spotify/YouTube URL)
curl -s 'https://spoken.md/search?q=huberman+sleep' \
-H 'x-api-key: pt_demo'
# 2. Fetch the transcript as Markdown
curl -s 'https://spoken.md/transcripts/1000651996090' \
-H 'x-api-key: pt_demo'The transcript comes back as Markdown with named speakers and timestamps:
**John Smith** (0:00)
Welcome to the show. Today we're talking about...
**Jane Doe** (0:15)
Thanks for having me.Related MCP server: Lenny's Podcast MCP
Endpoints
Method & path | What it does | Credits |
| Find episodes; returns | 0 |
| List a show's full back catalog; returns every episode's | 0 |
| Return the Markdown transcript | 1 on first fetch, 0 on repeat |
| Current credit balance + usage history | 0 |
| New-key checkout (Stripe) | — |
| Returning-customer top-up (Stripe) | — |
Auth is the x-api-key header. Responses include X-Credits-Remaining and X-Credits-Charged. See agents.md for the full error table and response shapes.
Examples
examples/podcast_summarizer.py— fetch a transcript and summarize itexamples/rag_pipeline.py— chunk a transcript for a vector store / RAGexamples/quickstart.sh— search → transcript in two curl callsexamples/archive-show.sh— archive a show's entire back catalogue, one file per episode
Use as an MCP server
This repo includes spoken-mcp, a Model Context Protocol server that exposes Spoken to MCP-compatible agents (Claude Desktop, Cursor, Cline, …). It provides four tools:
Tool | Description |
| Find episodes by text or a pasted Spotify/YouTube URL |
| List a show's entire back-catalog from a |
| Fetch an episode's transcript as Markdown with real speaker names |
| Check remaining credits |
Add it to your MCP client config (e.g. Claude Desktop's claude_desktop_config.json):
{
"mcpServers": {
"spoken": {
"command": "npx",
"args": ["-y", "spoken-mcp"],
"env": { "SPOKEN_API_KEY": "pt_your_key" }
}
}
}SPOKEN_API_KEY defaults to pt_demo (search works fully; transcripts limited to the demo episode). Get a real key at spoken.md.
Run from source instead:
npm install && npm run build
SPOKEN_API_KEY=pt_your_key node dist/index.jsUse with AI agents
Spoken is designed to be called by agents. Point your agent at the Agent Skill (also served at https://spoken.md/.well-known/skills/spoken-md/SKILL.md), or hand it agents.md. The OpenAPI spec makes it easy to wrap as a tool for any function-calling or MCP-compatible client (Claude, GPT, Cursor).
Pricing
Pay-per-use credits, no subscription. New keys: 100 for $15, 500 for $50, 2,000 for $160. Machine-readable at spoken.md/pricing.md.
Links
Website & docs: https://spoken.md
Agent instructions: https://spoken.md/agents.md
OpenAPI spec: https://spoken.md/.well-known/openapi.json
LLM-friendly overview: https://spoken.md/llms.txt
Spoken is built and maintained at spoken.md.
Available Tools
3 toolsget_balanceGet credit balanceA
Check the current Spoken credit balance, account email, and recent usage for the configured API key. Does not consume credits.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description provides key behavioral trait: 'Does not consume credits' and lists what is checked (balance, email, usage). It appropriately discloses read-only nature without side effects.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single, concise sentence that front-loads the action and includes important side-effect information. Every word earns its place with no redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a parameterless read-only tool, the description fully covers what the tool does and returns (balance, email, usage). It explains safety (no credits consumed) and distinguishes from siblings.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
No parameters exist, so schema coverage is effectively 100%. The baseline for 0 parameters is 4, and the description adds no parameter info since none needed. Score 5 as it fully satisfies the dimension.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool checks the Spoken credit balance, account email, and recent usage. The verb 'Check' and specific resources make the purpose unambiguous and distinct from sibling tools (get_transcript, search_podcasts).
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies when to use (to check balance) and explicitly states it does not consume credits, indicating safety. While it does not mention alternatives, sibling tools are sufficiently different that no confusion arises.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_transcriptGet transcriptA
Fetch a podcast episode's transcript as clean Markdown with real speaker names and timestamps. Pass an episode id from search_podcasts. Costs 1 credit on the first fetch of an episode; repeat fetches are free and errors are never charged.
| Name | Required | Description | Default |
|---|---|---|---|
| episode_id | Yes | Episode id returned by search_podcasts. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Despite no annotations, the description discloses that first fetch costs 1 credit, repeats are free, and errors are not charged, providing helpful behavioral context for a read operation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two succinct sentences that front-load the action and output, with no extraneous information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a single-parameter tool with no output schema, the description adequately covers usage, return format, and cost policy; lacks mention of limits or size.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% with episode_id already described; the description only reinforces this without adding extra semantic detail.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool fetches a transcript as cleaned Markdown with speaker names and timestamps, distinguishing it from sibling tools like get_balance and search_podcasts.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly instructs to pass an episode id from search_podcasts and mentions credit cost policy, but does not specify when not to use or alternative tools.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
search_podcastsSearch podcastsA
Search published podcast episodes by text query, or paste an episode URL (Spotify, YouTube, etc.). Returns matching episodes with their id, title, podcast, and date. Use the id with get_transcript. Does not consume credits.
| Name | Required | Description | Default |
|---|---|---|---|
| query | Yes | Free text (e.g. 'huberman sleep') or a pasted episode URL. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries full burden. It states it returns specific fields and does not consume credits, which implies read-only behavior. However, it omits details on pagination, result limits, or error conditions.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is three sentences, front-loaded with the main action, and every sentence serves a purpose with no wasted words.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given one parameter, no output schema, and no annotations, the description covers input format, return fields, and linkage to get_transcript. It lacks details on result limits or pagination, but is mostly complete for a simple search tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% and the schema description already explains the query parameter. The tool description repeats the same information without adding new meaning beyond the schema, so baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states it searches published podcast episodes by text or URL. It distinguishes from siblings: get_balance is unrelated, get_transcript is a follow-up using the returned id.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description specifies when to use the tool: to search episodes. It mentions the follow-up tool get_transcript, but lacks explicit guidance on when not to use it or alternatives, though siblings are few and clear.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
TDQS
Each tool targets a distinct concern: account balance, transcript retrieval, and episode search. No overlap in functionality.
All tools follow a consistent verb_noun pattern with snake_case (get_balance, get_transcript, search_podcasts).
Three tools is perfectly scoped for the server's purpose of podcast transcript retrieval, covering the core workflow without unnecessary complexity.
The tool surface covers the primary use case (search → transcript) and adds account balance. A minor gap is the lack of a tool to list all available podcasts or get episode details without searching, but this is acceptable.
Maintenance
Resources
Unclaimed servers have limited discoverability.
Looking for Admin?
If you are the server author, to access and configure the admin panel.
Related MCP Connectors
Podcast intelligence for agents: transcripts, clips, speaker diarization, mention tracking.
Search speech in podcasts, government meetings, and your own audio: speakers, entities, timestamps.
YouTube transcripts, search, channels, playlists and bulk transcript jobs for AI agents. 14 tools.
Podcast search, metadata, chapters, and transcripts for AI agents — from $15/mo
Related MCP Servers
- AlicenseAqualityDmaintenanceEnables access to Jupiter Broadcasting podcast episodes through RSS feed parsing. Supports searching episodes by date, hosts, or content, retrieving detailed episode information, and fetching transcripts when available.45MIT
- FlicenseAqualityDmaintenanceEnables searching and retrieving transcripts from over 280 episodes of Lenny's Podcast to access expert product and growth insights. It allows users to query by topic, list available episodes, and fetch full interview transcripts directly through Claude.335
- FlicenseNot gradedqualityCmaintenanceConnects AI assistants to Pocket Casts accounts for browsing subscriptions, reading episode details, and retrieving transcripts with automatic transcription via AssemblyAI when no native transcript exists.
- AlicenseNot gradedqualityCmaintenanceWraps the Podcast Index API (podcastindex.org) to enable AI agents to search and retrieve podcast episodes and metadata.14MIT
Latest Blog Posts
- Who's Calling? MCP Hosts Are an Identity Blind Spot (And the Spec Knows It)By Om-Shree-0709 on .mcpAgent IdentityOAuth 2.1
- Your AI Chatbot Just Exposed Your CEO's Salary to an InternBy Om-Shree-0709 on .Agent IdentityMCP SecurityOAuth Delegation
- Why MCP Servers Need Execution Sandboxing (And Why Your Current Stack Isn't Enough)By Om-Shree-0709 on .Agentic AiPrompt InjectionWebAssembly
MCP directory API
We provide all the information about MCP servers via our MCP API.
curl -X GET 'https://glama.ai/api/mcp/v1/servers/spokenmd/spoken'
If you have feedback or need assistance with the MCP directory API, please join our Discord server