FunASR MCP Server
OfficialProvides an OpenAI-compatible API server for integrating FunASR speech recognition with LangChain, enabling AI agents to transcribe audio.
Quick Start
No local setup? Open the Colab quickstart to transcribe a public sample or upload your own audio in a browser.
Found FunASR useful? Star the project so more builders can find it.
# CPU-only installs can use the default PyPI wheels.
pip install torch torchaudio
pip install funasrFor GPU quickstarts, install the PyTorch and torchaudio wheels that match your NVIDIA driver from pytorch.org before installing FunASR. After installation, confirm the GPU is visible:
python - <<'PY'
import torch
print(torch.cuda.is_available())
PYOnly use device="cuda" when this prints True; otherwise use device="cpu"
or reinstall PyTorch with the correct CUDA wheel.
Flagship model — Fun-ASR-Nano (LLM-ASR for Chinese, English, and Japanese, plus Chinese dialect groups and regional accents; needs a GPU):
from funasr import AutoModel
model = AutoModel(model="FunAudioLLM/Fun-ASR-Nano-2512", device="cuda")
result = model.generate(input="https://isv-data.oss-cn-hangzhou.aliyuncs.com/ics/MaaS/ASR/test_audio/asr_example_zh.wav")
print(result[0]["text"])
# 欢迎大家来体验达摩院推出的语音识别模型。For the separate 31-language checkpoint, use Fun-ASR-MLT-Nano-2512. Language coverage is checkpoint-specific, so Nano and MLT-Nano should be treated as distinct model choices.
On CPU (or for five-language ASR plus emotion and audio-event tags), use SenseVoiceSmall. The pipeline below composes SenseVoiceSmall with FSMN-VAD and CAM++; diarization is provided by the separate CAM++ model, not by the SenseVoiceSmall checkpoint: See the SenseVoice paper, Hugging Face checkpoint, and GGUF edge checkpoint.
from funasr import AutoModel
from funasr.utils.postprocess_utils import rich_transcription_postprocess
model = AutoModel(model="iic/SenseVoiceSmall", vad_model="fsmn-vad", spk_model="cam++", device="cuda") # use device="cpu" if you don't have a GPU
result = model.generate(
input="https://isv-data.oss-cn-hangzhou.aliyuncs.com/ics/MaaS/ASR/test_audio/asr_example_zh.wav",
batch_size_s=300,
)
# The AutoModel pipeline returns VAD segments with speaker ids and timestamps:
for seg in result[0]["sentence_info"]:
print(f"[{seg['start']/1000:.1f}s] Speaker {seg['spk']}: {rich_transcription_postprocess(seg['sentence'])}")Output — structured text with speaker labels, timestamps, and punctuation:
[0.6s] Speaker 0: 欢迎大家来体验达摩院推出的语音识别模型One AutoModel pipeline call coordinates the configured ASR, VAD, and speaker
models and returns the combined result.
Scale & deploy the flagship
At scale, accelerate Fun-ASR-Nano with vLLM (batch processing):
from funasr.auto.auto_model_vllm import AutoModelVLLM
model = AutoModelVLLM(model="FunAudioLLM/Fun-ASR-Nano-2512", tensor_parallel_size=1)
results = model.generate(["audio1.wav", "audio2.wav"], language="auto")Deploy as API server:
funasr-server --device cuda→ OpenAI-compatible endpoint at localhost:8000Use with AI agents: MCP Server for Claude/Cursor · OpenAI API for LangChain/Dify/AutoGen
Use with voice agents: OpenClaw realtime plugin for self-hosted Talk and Voice Call transcription
Why FunASR?
Whisper is a single model; FunASR is a toolkit — you pick the right model per job: Fun-ASR-Nano (Chinese, English, Japanese, and Chinese dialects; GPU), Fun-ASR-MLT-Nano (31 languages), SenseVoiceSmall (five-language ASR plus emotion and audio events), and Paraformer (low-latency streaming). The table shows toolkit-level capabilities and names the model or pipeline that provides each one:
FunASR (toolkit) | Whisper | Cloud APIs | |
Top speed | 340x realtime (Fun-ASR-Nano + vLLM) | 13x realtime | ~1x realtime |
Speaker ID | ✅ via VAD + CAM++ pipeline | ❌ Needs pyannote | ✅ Extra cost |
Emotion | ✅ via SenseVoice | ❌ | ❌ |
Languages | Checkpoint-specific (for example Qwen3-ASR 52, MLT-Nano 31, Nano zh/en/ja) | 57 | Varies |
Streaming | ✅ WebSocket (Paraformer) | ❌ | ✅ |
CPU viable | ✅ 17x realtime (SenseVoice) | ❌ Too slow | N/A |
Self-hosted | ✅ Yes (toolkit: MIT; model licenses vary) | ✅ MIT license | ❌ Cloud only |
Cost | Free | Free | $0.006/min+ |
Trying FunASR for the first time? Use the Colab quickstart before setting up a local environment. Choosing a first model? Start with the model selection guide. Planning a switch from Whisper or a cloud ASR provider? Use the migration guide and benchmark example to test representative audio, map features, and roll out safely.
Related MCP server: simple-asr-mcp
Installation
pip install funasrgit clone https://github.com/modelscope/FunASR.git && cd FunASR
pip install -e ./Requirements: Python ≥ 3.8. Install PyTorch + torchaudio first (pytorch.org), then pip install funasr.
Model Zoo
Model | Task | Languages | Params | Links |
Fun-ASR-Nano | ASR | zh/en/ja + Chinese dialects and accents | 800M | |
Fun-ASR-MLT-Nano | ASR | 31 languages | 800M | |
SenseVoiceSmall | ASR + emotion + events | zh/en/ja/ko/yue | 234M | |
Paraformer-zh | ASR + timestamps | zh/en | 220M | |
Paraformer-zh-streaming | Streaming ASR | zh/en | 220M | |
Qwen3-ASR | ASR, 52 languages | multilingual | 1.7B | |
GLM-ASR-Nano | ASR, 17 languages | multilingual | 1.5B | |
Whisper-large-v3 | ASR + translation | multilingual | 1550M | |
Whisper-large-v3-turbo | ASR + translation | multilingual | 809M | |
ct-punc | Punctuation | zh/en | 290M | |
fsmn-vad | VAD | zh/en | 0.4M | |
cam++ | Speaker diarization | — | 7.2M | |
emotion2vec+large | Emotion recognition | — | 300M |
Usage
Full examples with parameter docs: Tutorial →
from funasr import AutoModel
# Chinese production (VAD + ASR + punctuation + speaker)
model = AutoModel(model="paraformer-zh", vad_model="fsmn-vad", punc_model="ct-punc", spk_model="cam++", device="cuda")
result = model.generate(input="https://isv-data.oss-cn-hangzhou.aliyuncs.com/ics/MaaS/ASR/test_audio/asr_example_zh.wav", hotword="关键词 20")
# Optional Silero VAD (install first: python -m pip install "funasr[silero]")
model = AutoModel(
model="paraformer-zh", vad_model="silero-vad", device="cuda",
vad_kwargs={"silero_threshold": 0.5, "silero_min_silence_duration_ms": 100},
)
result = model.generate(input="audio.wav")
# Streaming real-time (feed audio chunk by chunk)
import soundfile as sf
model = AutoModel(model="paraformer-zh-streaming", device="cuda")
audio, sr = sf.read("speech.wav", dtype="float32") # 16 kHz mono
chunk_size = [0, 10, 5] # 600 ms chunks
chunk_stride = chunk_size[1] * 960
cache = {}
n_chunks = (len(audio) - 1) // chunk_stride + 1
for i in range(n_chunks):
chunk = audio[i * chunk_stride : (i + 1) * chunk_stride]
res = model.generate(input=chunk, cache=cache, is_final=(i == n_chunks - 1),
chunk_size=chunk_size, encoder_chunk_look_back=4, decoder_chunk_look_back=1)
if res[0]["text"]:
print(res[0]["text"], end="", flush=True)
# Emotion recognition
model = AutoModel(model="emotion2vec_plus_large", device="cuda")
result = model.generate(input="audio.wav", granularity="utterance")CLI (Agent-Friendly)
# Transcribe audio (simplest)
funasr audio.wav
# JSON output (for AI agents)
funasr audio.wav --output-format json
# SRT subtitles
funasr audio.wav --output-format srt --output-dir ./subs
# Speaker diarization + timestamps
funasr audio.wav --spk --timestamps -f json
# Choose model and language
funasr audio.wav --model paraformer --language zh
# Batch transcribe
funasr *.wav --output-format srt --output-dir ./outputAvailable models: sensevoice (default), paraformer, paraformer-en, fun-asr-nano
Deploy
# OpenAI-compatible API (recommended)
pip install torch torchaudio
pip install funasr vllm fastapi uvicorn python-multipart
funasr-server --device cuda
# → POST /v1/audio/transcriptions at localhost:8000
# Joint long-form ASR + anonymous speaker labels (offline HTTP):
funasr-server --model moss-transcribe-diarize --device cuda:0MOSS service, Docker, Kubernetes, vLLM, SGLang, LocalAI, and FunClip guide →
Verify it with a public sample:
curl -L https://isv-data.oss-cn-hangzhou.aliyuncs.com/ics/MaaS/ASR/test_audio/BAC009S0764W0121.wav -o sample.wav
curl http://localhost:8000/v1/audio/transcriptions \
-F file=@sample.wav \
-F model=sensevoice \
-F response_format=verbose_json# Docker streaming service
docker pull registry.cn-hangzhou.aliyuncs.com/funasr_repo/funasr:funasr-runtime-sdk-online-cpu-0.1.12CPU / Edge — llama.cpp / GGUF (no GPU, no Python)
Run SenseVoice / Paraformer / Fun-ASR-Nano as a single self-contained binary on CPU and edge devices — this is to FunASR what whisper.cpp is to Whisper, but with ~3× lower CER than whisper.cpp on Chinese. Built-in FSMN-VAD, no Python at runtime.
# Linux / macOS: run from the extracted release directory
bash download-funasr-model.sh sensevoice ./gguf # or: paraformer | nano
./llama-funasr-sensevoice -m ./gguf/sensevoice-small-q8.gguf --vad ./gguf/fsmn-vad.gguf -a audio.wav
# → 欢迎大家来体验达摩院推出的语音识别模型# Windows PowerShell: run from the extracted archive root (with the `hf` CLI installed)
hf download FunAudioLLM/SenseVoiceSmall-GGUF sensevoice-small-q8.gguf --local-dir .\gguf
hf download FunAudioLLM/fsmn-vad-GGUF fsmn-vad.gguf --local-dir .\gguf
.\llama-funasr-sensevoice.exe -m .\gguf\sensevoice-small-q8.gguf --vad .\gguf\fsmn-vad.gguf -a audio.wav
# Use the windows-x64-vulkan package with a current AMD, Intel, or NVIDIA Vulkan driver:
.\llama-funasr-sensevoice.exe -m .\gguf\sensevoice-small-q8.gguf --vad .\gguf\fsmn-vad.gguf -a audio.wav --backend vulkan
# Use the windows-x64-cuda package on RTX 30-class GPUs:
.\llama-funasr-sensevoice.exe -m .\gguf\sensevoice-small-q8.gguf --vad .\gguf\fsmn-vad.gguf -a audio.wav --backend cudaUse funasr-llamacpp-linux-x64-vulkan.tar.gz on Linux GPU systems with a
working Vulkan driver/ICD:
./llama-funasr-sensevoice -m ./gguf/sensevoice-small-q8.gguf --vad ./gguf/fsmn-vad.gguf -a audio.wav --backend vulkanThe Windows Vulkan ZIP uses the system Vulkan loader supplied by the GPU driver; installing the Vulkan SDK is only necessary when building from source. Both Vulkan packages currently accelerate SenseVoiceSmall.
Tagged releases provide two Windows CUDA packages. The standard
windows-x64-cuda ZIP targets CUDA architecture 86, while
windows-x64-cuda-blackwell targets architecture 120 (sm_120) for RTX 50 /
Blackwell GPUs. Both ZIPs bundle the required cuBLAS DLLs and use the static MSVC
runtime, so users need a compatible NVIDIA driver but not a separate CUDA Toolkit
installation. CI verifies the architecture and package boundary; it does not prove
inference on physical Blackwell hardware.
Prebuilt binaries: Releases · v0.2.6 · Linux Vulkan tarball · Windows Vulkan zip · Windows CUDA zip · Windows Blackwell CUDA zip · Download & quickstart: funasr.com/deploy/llama-cpp · GGUF models: Hugging Face · Docs & benchmarks: runtime/llama.cpp/
OpenAI API example → · Gradio demo → · Client recipes → · JavaScript/TypeScript recipes → · Kubernetes template → · Workflow recipes → · Postman collection → · OpenAPI spec → · Security guide → · Deployment matrix → · Deployment docs → · Agent integration →
Benchmark
184 long-form audio files (192 min). Full report → · RTFx and reproducibility notes →
Model | Chinese CER ↓ | GPU Speed | CPU Speed | vs Whisper-large-v3 |
Fun-ASR-Nano (vLLM) | 8.20% | 340x realtime | — | 🚀 26x faster |
SenseVoice-Small | 7.81% | 170x realtime | 17x realtime | 🚀 13x faster |
Paraformer-Large | 10.18% | 120x realtime | 15x realtime | 🚀 9x faster |
Whisper-large-v3-turbo | 21.71% | 46x realtime | ❌ | 3.4x faster |
Whisper-large-v3 | 20.02% | 13x realtime | ❌ | baseline |
Key takeaway: FunASR models run on CPU faster than Whisper runs on GPU.
What's new
MOSS-Transcribe-Diarize brings long-form ASR, timestamps, and anonymous speaker labels to FunASR services, Docker, Kubernetes, vLLM/SGLang workflows, and FunClip. Deploy MOSS ->
FunASR 1.4.13 pins the core NumPy dependency below 2 to prevent binary ABI import failures. Install with
python -m pip install -U "funasr==1.4.13". Release ->Production deployment now includes faster, more resilient realtime serving and verified llama.cpp archives for ten Linux, macOS, and Windows targets. GPU services -> · CPU/edge packages ->
See GitHub Releases for the complete changelog and downloadable assets.
Community
🐛 Issues | |
Star History
License
FunASR toolkit source code in this repository: MIT License.
Pretrained model weights are licensed separately. Check the license shown on each model card; when a model card links to the FunASR Model Open Source License Agreement, those terms apply.
Citations
@inproceedings{gao2023funasr,
author={Zhifu Gao and others},
title={FunASR: A Fundamental End-to-End Speech Recognition Toolkit},
booktitle={INTERSPEECH},
year={2023}
}This server cannot be installed
Maintenance
Resources
Unclaimed servers have limited discoverability.
Looking for Admin?
If you are the server author, to access and configure the admin panel.
Related MCP Connectors
MCP server for Speech-to-Text
MCP server exposing the AceDataCloud Fish Audio API (text-to-speech with voice conditioning)
Remote MCP server for supportsheep: run AI interviews and manage support content for your blog.
Related MCP Servers
- AlicenseNot gradedqualityCmaintenanceMCP server for offline speech-to-text and speaker diarization, enabling AI agents to transcribe audio locally without cloud APIs.3MIT
- AlicenseNot gradedqualityDmaintenanceMinimal MCP server for local speech recognition using faster-whisper. Runs on CPU, no cloud required.MIT
- AlicenseNot gradedqualityBmaintenanceLocal-only transcription server for MCP agents, powered by CrispASR. Transcribes audio/video files without cloud uploads, supporting English and Chinese.1MIT
- AlicenseAqualityCmaintenanceLocal, private audio transcription MCP server enabling AI agents to transcribe audio files entirely on-device without uploading data.3MIT
Latest Blog Posts
- Who's Calling? MCP Hosts Are an Identity Blind Spot (And the Spec Knows It)By Om-Shree-0709 on .mcpAgent IdentityOAuth 2.1
- Your AI Chatbot Just Exposed Your CEO's Salary to an InternBy Om-Shree-0709 on .Agent IdentityMCP SecurityOAuth Delegation
- Why MCP Servers Need Execution Sandboxing (And Why Your Current Stack Isn't Enough)By Om-Shree-0709 on .Agentic AiPrompt InjectionWebAssembly
MCP directory API
We provide all the information about MCP servers via our MCP API.
curl -X GET 'https://glama.ai/api/mcp/v1/servers/modelscope/FunASR'
If you have feedback or need assistance with the MCP directory API, please join our Discord server