Audio Playback MCP Server
# Audio Playback MCP Server
This repository provides an MCP server named `audio-playback-server` that exposes a single `audio_playback` tool. The tool lets an MCP client start, stop, and inspect playback of local audio files routed through a configurable virtual audio output device.
## Features
- Play audio files from a configurable root directory using `ffplay`.
- Get audio file duration (in seconds) when playing files using `ffprobe`.
- Stop current playback.
- Pause playback (saves position); resume continues from that position.
- Query current status with a position estimate.
- List audio files hosted locally under `AUDIO_ROOT_DIR`.
- Path safety enforcement to prevent leaving the configured root directory.
- Run as either a local stdio MCP server or a long-running HTTP MCP server.
## Configuration
The server reads configuration from environment variables and an optional JSON file. Environment variables take precedence over JSON values.
Required:
- `AUDIO_ROOT_DIR`: Root directory containing allowed audio files.
- `AUDIO_OUTPUT_DEVICE`: Identifier of the virtual output device.
Optional playback settings:
- `DEFAULT_FORMAT`: File extension to append when none is provided (default `wav`).
- `FFPLAY_PATH`: Path to the `ffplay` binary (default `ffplay`).
- `AUDIO_PLAYBACK_CONFIG`: Path to a JSON config file containing any of the above keys.
Optional MCP transport settings:
- `MCP_TRANSPORT`: `stdio` (default) or `http`.
- `MCP_HTTP_HOST`: HTTP bind host for `http` transport (default `0.0.0.0`).
- `MCP_HTTP_PORT`: HTTP bind port for `http` transport (default `8765`).
- `MCP_HTTP_PATH`: MCP endpoint path for `http` transport (default `/mcp`).
- `MCP_DNS_REBINDING_PROTECTION`: `true`/`false` (default `true`).
- `MCP_ALLOWED_HOSTS`: Comma-separated allow-list of host headers for HTTP mode.
An example config file is provided at `config/audio_playback_config.example.json`.
## Installation
```bash
python -m venv .venv
source .venv/bin/activate
pip install -e .
```
## Running the server (stdio)
```bash
AUDIO_ROOT_DIR=/path/to/audio \
AUDIO_OUTPUT_DEVICE="Virtual Cable" \
python -m audio_playback_server
```
## Running the server (HTTP, for remote clients)
```bash
AUDIO_ROOT_DIR=/path/to/audio \
AUDIO_OUTPUT_DEVICE="Virtual Cable" \
MCP_TRANSPORT=http \
MCP_HTTP_HOST=0.0.0.0 \
MCP_HTTP_PORT=8765 \
MCP_HTTP_PATH=/mcp \
python -m audio_playback_server
```
In HTTP mode, the server stays running and accepts MCP requests at `http://<host>:<port><path>` (for example `http://192.168.1.10:8765/mcp`).
### Remote access notes (Windows host + Raspberry Pi client)
1. Bind to `0.0.0.0` so the server listens on your LAN interface.
2. Allow inbound TCP on `MCP_HTTP_PORT` in Windows Defender Firewall.
3. If MCP host-header checks block requests, set `MCP_ALLOWED_HOSTS` to include the hostname/IP your client uses, or disable checks with `MCP_DNS_REBINDING_PROTECTION=false` for trusted networks only.
4. Configure your Raspberry Pi MCP client/Openclaw to call your Windows machine at `http://<windows-lan-ip>:8765/mcp`.
## Tool schema
The `audio_playback` tool accepts the following JSON input:
- `action`: `play`, `stop`, `pause`, `resume`, `status`, or `list_files` (required)
- `filename`: Relative path under `AUDIO_ROOT_DIR` (required for `play`)
- `start_offset_ms`: Start offset in milliseconds (default `0`)
- `list_limit`: Maximum number of files returned when `action=list_files` (default `200`)
Responses always include `success`, `message`, and a `state` object with `status`, `current_file`, `started_at_ms`, and `position_estimate_ms`. When playing a file, the response also includes `duration_seconds` (float) in the `state` object, indicating the total duration of the audio file in seconds. The duration is also included in the `message` text for convenience. For `list_files`, responses also include a `files` payload containing `root_dir`, `count`, `limit`, and `files` entries (`filename`, `size_bytes`).
TDQS
Scored across 1 tool
With only one tool, there is no possibility of confusion or overlap between tools, as there are no other tools to compare it to. The single tool has a clear and distinct purpose focused on audio playback for testing.
Since there is only one tool, naming consistency is inherently perfect—there are no other tool names to be inconsistent with. The tool name 'audio_playback' follows a clear verb_noun pattern.
A single tool is too few for a server with the purpose of controlling audio playback for automated testing, as it suggests a very limited scope that may not cover essential operations like pausing, stopping, or managing audio files. This is borderline for the apparent domain.
The tool surface is severely incomplete for audio playback control; it only offers playback without supporting basic operations such as pause, stop, volume control, or file management. This will likely cause agent failures in handling more complex testing scenarios.