freeaudiototext-mcp
by double2dev
README.md
# 🎙️ FreeAudioToText MCP Server
An official Model Context Protocol (MCP) server for [FreeAudioToText.com](https://freeaudiototext.com), enabling AI agents like **DeepSeek**, Claude, and Cursor to instantly transcribe any audio/video file or YouTube/TikTok URL into text with speaker diarization.
Our core transcription service is **100% free with unlimited usage**. It runs on high-performance local Apple Silicon hardware via the Cloudflare Edge, providing ultra-fast inference with state-of-the-art accuracy across 90+ languages.
## 🔗 Important Links
- **Website / Mac App**: [https://freeaudiototext.com](https://freeaudiototext.com)
- **Developer API**: [https://freeaudiototext.com/speech-to-text-api](https://freeaudiototext.com/speech-to-text-api)
- **Get API Key**: No API key is required for basic MCP tool usage! Our core transcription is fully free.
## 🛠️ Available Tools
This MCP server exposes the following tools to your AI assistant:
- `transcribe_audio`: Uploads a local audio or video file (e.g., MP3, M4A, WAV, MP4) for transcription. Returns a unique `job_id`.
- `transcribe_from_url`: Submits a YouTube or TikTok URL for extraction and transcription. Returns a `job_id`.
- `get_job_status`: Checks if the submitted `job_id` is "pending" or "completed".
- `get_transcript`: Retrieves the full generated transcript (with speaker labels like `[SPK_0]`) for a completed job.
- `generate_ai_report`: Generates a deep LLM analysis report (Summary, Action Items, Outline, or Speaker Analysis) for your transcript.
## 🚀 Installation & Setup
### Using npx (Recommended for Claude Desktop & Cursor)
Add the following to your MCP client configuration file (e.g., `claude_desktop_config.json` on macOS):
```json
{
"mcpServers": {
"freeaudiototext": {
"command": "npx",
"args": [
"-y",
"freeaudiototext-mcp"
]
}
}
}
```
### Using Smithery CLI
```bash
npx -y @smithery/cli install freeaudiototext-mcp --client cursor
```
## 🤖 Integration Examples
### Example 1: Using with DeepSeek Models
Since this server follows the standard MCP specification, you can use any DeepSeek-compatible MCP client (like Cursor, Cline, or Roo Code) to combine our transcription with DeepSeek's powerful reasoning.
Just ask your DeepSeek-powered agent:
> *"Use the FreeAudioToText tool to transcribe `/Users/myname/Downloads/board_meeting.m4a`. Once it finishes, act as an executive assistant and use your DeepSeek-R1 reasoning to summarize the key decisions and output an action item list."*
### Example 2: Using with Claude
> *"I have an interview recording at `https://youtube.com/watch?v=xxxx`. Can you transcribe it from the URL and give me a detailed speaker-by-speaker outline?"*
The agent will automatically:
1. Call `transcribe_audio` or `transcribe_from_url`.
2. Poll `get_job_status` until completion.
3. Retrieve the text via `get_transcript` and perform the advanced analysis.
## 📜 License
MIT License. See [LICENSE](LICENSE) for more information.
TDQS
A4/5.0
Scored across 5 tools
Disambiguation5/5
Each tool serves a distinct purpose: two input methods for transcription, status polling, transcript retrieval, and report generation. No ambiguity between them.
Naming Consistency5/5
All tool names follow a consistent verb_noun pattern in lowercase snake_case (transcribe_audio, get_job_status, etc.), making them predictable.
Tool Count5/5
Five tools cover the essential transcription workflow without being excessive or insufficient. The scope is well-suited for the server's purpose.
Completeness4/5
The tool set covers the core workflow (submit, poll, retrieve, analyze). Minor gaps like job cancellation or listing are absent but not critical for primary use.
Maintenance
ActivityMaintained
ResponsivenessSyncing