Skip to main content
Glama
README.md
# 🎙️ FreeAudioToText MCP Server

An official Model Context Protocol (MCP) server for [FreeAudioToText.com](https://freeaudiototext.com), enabling AI agents like **DeepSeek**, Claude, and Cursor to instantly transcribe any audio/video file or YouTube/TikTok URL into text with speaker diarization.

Our core transcription service is **100% free with unlimited usage**. It runs on high-performance local Apple Silicon hardware via the Cloudflare Edge, providing ultra-fast inference with state-of-the-art accuracy across 90+ languages.

## 🔗 Important Links

- **Website / Mac App**: [https://freeaudiototext.com](https://freeaudiototext.com)
- **Developer API**: [https://freeaudiototext.com/speech-to-text-api](https://freeaudiototext.com/speech-to-text-api)
- **Get API Key**: No API key is required for basic MCP tool usage! Our core transcription is fully free.

## 🛠️ Available Tools

This MCP server exposes the following tools to your AI assistant:

- `transcribe_audio`: Uploads a local audio or video file (e.g., MP3, M4A, WAV, MP4) for transcription. Returns a unique `job_id`.
- `transcribe_from_url`: Submits a YouTube or TikTok URL for extraction and transcription. Returns a `job_id`.
- `get_job_status`: Checks if the submitted `job_id` is "pending" or "completed".
- `get_transcript`: Retrieves the full generated transcript (with speaker labels like `[SPK_0]`) for a completed job.
- `generate_ai_report`: Generates a deep LLM analysis report (Summary, Action Items, Outline, or Speaker Analysis) for your transcript.

## 🚀 Installation & Setup

### Using npx (Recommended for Claude Desktop & Cursor)

Add the following to your MCP client configuration file (e.g., `claude_desktop_config.json` on macOS):

```json
{
  "mcpServers": {
    "freeaudiototext": {
      "command": "npx",
      "args": [
        "-y",
        "freeaudiototext-mcp"
      ]
    }
  }
}
```

### Using Smithery CLI
```bash
npx -y @smithery/cli install freeaudiototext-mcp --client cursor
```

## 🤖 Integration Examples

### Example 1: Using with DeepSeek Models
Since this server follows the standard MCP specification, you can use any DeepSeek-compatible MCP client (like Cursor, Cline, or Roo Code) to combine our transcription with DeepSeek's powerful reasoning.

Just ask your DeepSeek-powered agent:
> *"Use the FreeAudioToText tool to transcribe `/Users/myname/Downloads/board_meeting.m4a`. Once it finishes, act as an executive assistant and use your DeepSeek-R1 reasoning to summarize the key decisions and output an action item list."*

### Example 2: Using with Claude
> *"I have an interview recording at `https://youtube.com/watch?v=xxxx`. Can you transcribe it from the URL and give me a detailed speaker-by-speaker outline?"*

The agent will automatically:
1. Call `transcribe_audio` or `transcribe_from_url`.
2. Poll `get_job_status` until completion.
3. Retrieve the text via `get_transcript` and perform the advanced analysis.

## 📜 License

MIT License. See [LICENSE](LICENSE) for more information.

TDQS

A4/5.0

Scored across 5 tools

Disambiguation5/5

Each tool serves a distinct purpose: two input methods for transcription, status polling, transcript retrieval, and report generation. No ambiguity between them.

Naming Consistency5/5

All tool names follow a consistent verb_noun pattern in lowercase snake_case (transcribe_audio, get_job_status, etc.), making them predictable.

Tool Count5/5

Five tools cover the essential transcription workflow without being excessive or insufficient. The scope is well-suited for the server's purpose.

Completeness4/5

The tool set covers the core workflow (submit, poll, retrieve, analyze). Minor gaps like job cancellation or listing are absent but not critical for primary use.

Maintenance

ActivityMaintained
ResponsivenessSyncing