Skip to main content
Glama
YuvrajSinghBhadoria2

OpenCode Voice MCP Server

README.md
<div align="center">

# 🎀 OpenCode Voice MCP Server

### Voice Input for AI Coding Assistants

[![License: MIT](https://img.shields.io/badge/License-MIT-blue.svg)](LICENSE)
[![npm version](https://img.shields.io/npm/v/@yuvarjbhado/voice-mcp.svg)](https://www.npmjs.com/package/@yuvarjbhado/voice-mcp)
[![GitHub stars](https://img.shields.io/github/stars/YuvrajSinghBhadoria2/opencode-voice-mcp)](https://github.com/YuvrajSinghBhadoria2/opencode-voice-mcp)
[![GitHub issues](https://img.shields.io/github/issues/YuvrajSinghBhadoria2/opencode-voice-mcp)](https://github.com/YuvrajSinghBhadoria2/opencode-voice-mcp)
[![MCP Compatible](https://img.shields.io/badge/MCP-Compatible-green.svg)](https://modelcontextprotocol.io)

**Speak your prompts. No typing required.**

[Installation](#installation) β€’ [Quick Start](#quick-start) β€’ [Tools](#tools) β€’ [Configuration](#configuration) β€’ [Architecture](#architecture) β€’ [Contributing](#contributing)

</div>

---

## ✨ Features

| Feature | Description |
|---------|-------------|
| 🎀 **Voice Recording** | Record audio from microphone with configurable duration |
| πŸ—£οΈ **Speech-to-Text** | Transcribe using local Whisper model (100% offline) |
| ⌨️ **Auto-Typing** | Type transcribed text at cursor position |
| πŸ”’ **Privacy First** | No cloud API β€” audio never leaves your machine |
| 🌍 **Multi-Language** | Support for 99+ languages via Whisper |
| πŸ”Œ **MCP Standard** | Works with OpenCode, Claude Code, Cursor, and more |

## πŸ“¦ Installation

```bash
# Install globally
npm install -g @yuvarjbhado/voice-mcp

# Or use with npx (no install required)
npx @yuvarjbhado/voice-mcp
```

### Prerequisites

<details>
<summary><strong>macOS</strong></summary>

```bash
# Install recording tool
brew install sox

# Install transcription engine
pip install faster-whisper
```
</details>

<details>
<summary><strong>Linux (Ubuntu/Debian)</strong></summary>

```bash
# Install recording tool
sudo apt install sox

# Install transcription engine
pip install faster-whisper
```
</details>

<details>
<summary><strong>Windows</strong></summary>

```bash
# Install FFmpeg (via scoop)
scoop install ffmpeg

# Install transcription engine
pip install faster-whisper
```
</details>

## πŸš€ Quick Start

### Step 1: Configure MCP Server

Add to your MCP config file:

| Tool | Config Location |
|------|-----------------|
| **OpenCode** | `~/.config/opencode/config.json` |
| **Claude Code** | `~/.claude/claude_desktop_config.json` |
| **Cursor** | `~/.cursor/mcp.json` |

**OpenCode** (use `mcp` key):
```json
{
  "mcp": {
    "voice": {
      "command": "npx",
      "args": ["-y", "@yuvarjbhado/voice-mcp"]
    }
  }
}
```

**Claude Desktop / Cursor** (use `mcpServers` key):
```json
{
  "mcpServers": {
    "voice": {
      "command": "npx",
      "args": ["-y", "@yuvarjbhado/voice-mcp"]
    }
  }
}
```

### Step 2: Restart Your Tool

Restart OpenCode, Claude Code, or Cursor to load the MCP server.

### Step 3: Use Voice Input

```
@voice voice_transcribe
@voice voice_type
@voice voice_status
```

## πŸ› οΈ Tools

### `voice_transcribe`

Record audio from microphone and transcribe to text.

```json
{
  "name": "voice_transcribe",
  "arguments": {
    "duration": 10,
    "language": "en"
  }
}
```

| Parameter | Type | Default | Description |
|-----------|------|---------|-------------|
| `duration` | number | `10` | Recording duration in seconds |
| `language` | string | `auto` | Language code (e.g., `en`, `es`, `fr`) |

**Returns:** Transcribed text as string.

---

### `voice_type`

Record audio, transcribe to text, and type it at the cursor position.

```json
{
  "name": "voice_type",
  "arguments": {
    "duration": 10,
    "language": "en"
  }
}
```

| Parameter | Type | Default | Description |
|-----------|------|---------|-------------|
| `duration` | number | `10` | Recording duration in seconds |
| `language` | string | `auto` | Language code |

**Returns:** Confirmation message with typed text.

---

### `voice_status`

Check if voice recording and transcription are available.

```json
{
  "name": "voice_status",
  "arguments": {}
}
```

**Returns:**
```json
{
  "recording": "rec",
  "transcription": "faster-whisper (local)",
  "platform": "darwin",
  "ready": true
}
```

## βš™οΈ Configuration

### Environment Variables

| Variable | Description | Default |
|----------|-------------|---------|
| `WHISPER_MODEL` | Whisper model size | `base` |
| `WHISPER_DEVICE` | Device to use (`cpu`, `cuda`, `auto`) | `auto` |
| `WHISPER_COMPUTE` | Compute type (`int8`, `float16`, `float32`) | `int8` |

### Model Sizes

| Model | Size | Speed | Accuracy | VRAM |
|-------|------|-------|----------|------|
| `tiny` | ~75MB | ⚑⚑⚑⚑ | ⭐⭐ | ~1GB |
| `base` | ~150MB | ⚑⚑⚑ | ⭐⭐⭐ | ~1GB |
| `small` | ~500MB | ⚑⚑ | ⭐⭐⭐⭐ | ~2GB |
| `medium` | ~1.5GB | ⚑ | ⭐⭐⭐⭐⭐ | ~5GB |
| `large-v3` | ~3GB | 🐌 | ⭐⭐⭐⭐⭐ | ~10GB |

**Recommendation:** Use `base` for best balance of speed and accuracy.

## πŸ—οΈ Architecture

```
β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚                        MCP Client                               β”‚
β”‚                  (OpenCode / Claude Code / Cursor)              β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                           β”‚ JSON-RPC
                           β–Ό
β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚                    Voice MCP Server                             β”‚
β”‚                     (Node.js)                                   β”‚
β”œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€
β”‚  β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”  β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”  β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”         β”‚
β”‚  β”‚ voice_record β”‚  β”‚voice_transcribeβ”‚ β”‚  voice_type  β”‚         β”‚
β”‚  β”‚    Tool      β”‚  β”‚     Tool     β”‚  β”‚    Tool      β”‚         β”‚
β”‚  β””β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”˜  β””β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”˜  β””β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”˜         β”‚
β”‚         β”‚                 β”‚                 β”‚                   β”‚
β”‚         β–Ό                 β–Ό                 β–Ό                   β”‚
β”‚  β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”       β”‚
β”‚  β”‚              Audio Recording Layer                   β”‚       β”‚
β”‚  β”‚         (sox / ffmpeg / macOS rec)                  β”‚       β”‚
β”‚  β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜       β”‚
β”‚                            β”‚                                    β”‚
β”‚                            β–Ό                                    β”‚
β”‚  β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”       β”‚
β”‚  β”‚           Transcription Engine                      β”‚       β”‚
β”‚  β”‚      (faster-whisper / OpenAI API)                  β”‚       β”‚
β”‚  β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜       β”‚
β”‚                            β”‚                                    β”‚
β”‚                            β–Ό                                    β”‚
β”‚  β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”       β”‚
β”‚  β”‚              Output Layer                           β”‚       β”‚
β”‚  β”‚    β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”  β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”        β”‚       β”‚
β”‚  β”‚    β”‚ Return Text β”‚  β”‚ Type at Cursor      β”‚        β”‚       β”‚
β”‚  β”‚    β”‚   (MCP)     β”‚  β”‚ (osascript/xdotool) β”‚        β”‚       β”‚
β”‚  β”‚    β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜  β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜        β”‚       β”‚
β”‚  β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜       β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
```

### Data Flow

```
β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”    β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”    β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”    β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚  User   │───▢│  Micro- │───▢│  Whisper    │───▢│  Text   β”‚
β”‚  Speaks β”‚    β”‚  phone  β”‚    β”‚  Transcribe β”‚    β”‚  Output β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜    β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜    β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜    β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
     β”‚              β”‚               β”‚                β”‚
     β”‚              β”‚               β”‚                β”‚
     β–Ό              β–Ό               β–Ό                β–Ό
  "Hello      Records 16kHz    Processes with    Returns text
   world"     mono audio      base model        or types it
```

## πŸ”§ Development

### Setup

```bash
# Clone repository
git clone https://github.com/YuvrajSinghBhadoria2/opencode-voice-mcp.git
cd opencode-voice-mcp

# Install dependencies
npm install

# Build
npm run build

# Run in development
npm run dev
```

### Project Structure

```
opencode-voice-mcp/
β”œβ”€β”€ src/
β”‚   └── index.ts          # MCP server implementation
β”œβ”€β”€ dist/
β”‚   └── index.js          # Compiled output
β”œβ”€β”€ package.json          # Package configuration
β”œβ”€β”€ tsconfig.json         # TypeScript config
β”œβ”€β”€ build.sh              # Build script
└── publish.sh            # npm publish script
```

### Available Scripts

| Command | Description |
|---------|-------------|
| `npm run build` | Compile TypeScript to JavaScript |
| `npm run dev` | Run in development mode with tsx |
| `npm run start` | Run compiled server |
| `./publish.sh` | Build and publish to npm |

## 🀝 Contributing

Contributions are welcome! Please follow these steps:

1. **Fork** the repository
2. **Create** a feature branch (`git checkout -b feat/amazing-feature`)
3. **Commit** your changes (`git commit -m 'feat: add amazing feature'`)
4. **Push** to the branch (`git push origin feat/amazing-feature`)
5. **Open** a Pull Request

### Development Guidelines

- Follow TypeScript best practices
- Add tests for new features
- Update documentation as needed
- Use conventional commit messages

## πŸ“„ License

This project is licensed under the MIT License - see the [LICENSE](LICENSE) file for details.

## πŸ™ Acknowledgments

- [Model Context Protocol](https://modelcontextprotocol.io) - Standard for AI tool integration
- [faster-whisper](https://github.com/SYSTRAN/faster-whisper) - Fast Whisper implementation
- [OpenCode](https://github.com/anomalyco/opencode) - AI coding assistant
- [OpenAI Whisper](https://github.com/openai/whisper) - Speech recognition model

---

<div align="center">

**Built with ❀️ for the developer community**

[Report Bug](https://github.com/YuvrajSinghBhadoria2/opencode-voice-mcp/issues) β€’ [Request Feature](https://github.com/YuvrajSinghBhadoria2/opencode-voice-mcp/issues) β€’ [Discussions](https://github.com/YuvrajSinghBhadoria2/opencode-voice-mcp/discussions)

</div>

TDQS

A3.7/5.0

Scored across 3 tools

Disambiguation3/5

voice_transcribe and voice_type both perform recording and transcription, differing only in whether the result is typed into the active window. This creates some overlap, though voice_status is clearly distinct. Descriptions clarify the intended use.

Naming Consistency4/5

The voice_ prefix is consistent across all tools, but the suffix alternates between a verb (transcribe, type) and a noun (status), slightly breaking the uniform verb pattern. Overall, the naming is predictable and readable.

Tool Count5/5

With only three tools, the server is tightly scoped to the core voice capture and transcription workflows. Each tool serves a clear purpose, and the count feels appropriate for the narrow domain.

Completeness4/5

The server covers the primary actions of transcribing and typing, plus a status check. Missing optional features like language selection or audio device configuration, but these are not essential to the main workflow.

Maintenance

ActivityMaintained
ResponsivenessNo issues