Skip to main content
Glama

MCP-Audio Plugin

mcp-audio is an AIO-2030 compliant MCP plugin that performs voice-to-text transcription using the Audio speech recognition API.

It exposes the identify_voice method via both multipart/form-data and base64 formats, supports the AIO tools.call protocol, and returns JSON-RPC structured outputs.


Features

  • Fully AIO-compliant MCP plugin (/tools.call, /help)

  • Converts .wav/.mp3 audio files to transcripts using SiliconFlow

  • API key managed securely via .env file

  • Docker-compatible and minimal dependencies

  • Registration-ready for AIO endpoint registry


Related MCP server: Voice to Text MCP Server

Setup (Local)

1. Clone and Install

git clone git@github.com:AIO-2030/mcp-audio.git
cd mcp-audio
python -m venv venv && source venv/bin/activate
pip install -r requirements.txt

2. Add .env file

cp .env.example .env

Set your audio URL and API key:

AUDIO_URL=https--xxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxx
API_KEY=sk-xxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxx

3. Run the MCP server

python src/mcp_server.py

4. Docker

4.1 Build and Run

docker build -t mcp-audio .
docker run --env-file .env -p 8080:8080 mcp-audio

API Overview

POST /api/v1/mcp/voice_model

Upload audio file directly. Response:

{
  "transcript": "hello world",
  "confidence": 0.91,
  "audio_hash": "a1b2c3..."
}

POST /api/v1/mcp/tools.call (AIO Protocol)

JSON-RPC format with base64-encoded audio. Response:

{
  "method": "tools.call",
  "params": {
    "method": "identify_voice",
    "inputs": [
      {
        "type": "audio",
        "value": "<base64-audio>"
      }
    ]
  }
}

GET /api/v1/mcp/help

Auto-serves contents of mcp_audio_registration.json. Used by Queen AI for MCP discovery and service indexing.

Testing Tools

Base64 Voice Test

python test/test_audio_base64.py

Health Check

python health_check.py

MCP Registration (to AIO Endpoint Canister)

./register_mcp.sh

Requires jq, dfx, and a running endpoint_registry canister.

A
license - permissive license
-
quality - not tested
D
maintenance

Maintenance

Maintainers
Response time
Release cycle
Releases (12mo)
Commit activity

Resources

Unclaimed servers have limited discoverability.

Looking for Admin?

If you are the server author, to access and configure the admin panel.

Related MCP Servers

  • A
    license
    -
    quality
    D
    maintenance
    Provides voice recognition and text extraction capabilities with support for both stdio and MCP modes, processing audio files or base64 encoded data and returning structured results with language, emotion, and speaker information.
    MIT
  • F
    license
    -
    quality
    D
    maintenance
    A powerful speech-to-text MCP server that supports multiple audio formats and recognition engines including remote APIs (Bailian, OpenAI Whisper, iFLYTEK), Google Speech Recognition, and CMU Sphinx.
    1
  • A
    license
    A
    quality
    D
    maintenance
    An MCP server that enables transcribing local audio files and Telegram voice messages using OpenAI's Whisper via local inference or cloud API. It supports multiple audio formats, automatic language detection, and optional word-level timestamps for AI-powered audio analysis.
    5
    1
    MIT

View all related MCP servers

Related MCP Connectors

  • AI-manageable audio CDN: upload, transcode, normalize, stream & deliver audio, plus grounded docs.

  • Transcripts from YouTube, TikTok, Instagram and podcasts (Spotify, Apple, RSS), as clean JSON.

  • Hosted pay-per-use TTS: 54 neural voices, 9 languages incl. Brazilian Portuguese. $10 free credits.

View all MCP Connectors

Latest Blog Posts

MCP directory API

We provide all the information about MCP servers via our MCP API.

curl -X GET 'https://glama.ai/api/mcp/v1/servers/AIO-2030/mcp-audio'

If you have feedback or need assistance with the MCP directory API, please join our Discord server