dtelecom-stt
OfficialClick on "Install Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@dtelecom-stttranscribe the audio file meeting.wav"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
@dtelecom/stt-mcp
MCP (Model Context Protocol) server for dTelecom real-time speech-to-text with x402 micropayments.
Lets AI assistants (Claude, Cursor, etc.) transcribe audio files using dTelecom STT — pay-per-use with USDC, no API keys needed.
Tools
Tool | Description |
| Transcribe a WAV file (PCM16, 16kHz, mono) to text |
| Get current pricing ($0.005/min) |
| Check service health |
Related MCP server: Deepgram MCP Server
Setup
1. Install
npm install -g @dtelecom/stt-mcp2. Get a wallet
You need a private key with USDC. Either:
EVM (Base): Ethereum private key (0x hex) with USDC on Base — MetaMask, etc.
Solana: Solana private key (base58) with USDC on Solana — Phantom, Solflare, etc.
3. Configure your AI assistant
Claude Code (~/.claude.json):
{
"mcpServers": {
"dtelecom-stt": {
"command": "dtelecom-stt-mcp",
"env": {
"DTELECOM_PRIVATE_KEY": "YOUR_PRIVATE_KEY"
}
}
}
}Claude Desktop (claude_desktop_config.json):
{
"mcpServers": {
"dtelecom-stt": {
"command": "npx",
"args": ["-y", "@dtelecom/stt-mcp"],
"env": {
"DTELECOM_PRIVATE_KEY": "YOUR_PRIVATE_KEY"
}
}
}
}Cursor (Settings > MCP Servers > Add):
{
"dtelecom-stt": {
"command": "npx",
"args": ["-y", "@dtelecom/stt-mcp"],
"env": {
"DTELECOM_PRIVATE_KEY": "YOUR_PRIVATE_KEY"
}
}
}4. Convert audio (if needed)
The tool accepts WAV files in PCM16 16kHz mono format. Convert with:
ffmpeg -i input.mp3 -ar 16000 -ac 1 -acodec pcm_s16le output.wavEnvironment Variables
Variable | Required | Default | Description |
| Yes | — | EVM key (0x hex) or Solana key (base58) |
| No |
| STT service URL |
Pricing
$0.005/min, billed per session
Minimum 5 minutes ($0.025)
Paid in USDC on Base or Solana via x402 protocol
No accounts, no API keys, no subscriptions
Links
License
MIT
Available Tools
3 toolsstt_healthA
Check if the dTelecom STT service is running and healthy.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. The word 'check' implies a read-only, non-destructive operation, which is helpful. However, the description does not disclose what the response looks like or what 'healthy' means operationally, leaving some behavioral ambiguity for a health-check endpoint.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single sentence that is direct and free of filler. It states the resource, the action, and the expected result with no redundant words.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple zero-parameter health check, the description is largely sufficient. It names the service and the intent, and with no input schema requirements, an agent can invoke the tool without additional setup. The only minor gap is the lack of explicit return value information, but the purpose is clear enough for a health check.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters, so the schema already fully covers parameter expectations. The description adds no parameter details, but none are needed; the baseline for a zero-parameter tool is 4.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('check') with a clear resource ('dTelecom STT service') and a specific outcome ('running and healthy'). This makes the tool's purpose immediately obvious and distinguishes it from its siblings (transcribe_file and stt_pricing), which handle different concerns.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description clearly implies when to use this tool: when an agent needs to verify the health or running status of the STT service. It does not explicitly mention alternatives or exclusions, but the sibling tools are so semantically distinct that no routing ambiguity is introduced.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
stt_pricingA
Get current dTelecom STT pricing information. No payment required.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are present, so the description carries the full burden. It conveys a read-only intent through 'Get' and adds the useful behavioral fact 'No payment required'. However, it does not disclose response format, data freshness, or any other operational constraints, though the simplicity of a zero-parameter pricing lookup reduces the risk.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is extremely concise: one short main clause stating the action and resource, plus a brief relevant note about payment. Every word earns its place, and the core purpose is front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter, no-output-schema tool this description is largely sufficient: it explains what the tool provides and that it is free. It could mention the exact shape of the returned pricing data, but the triviality of the tool lowers the burden on the description.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema has zero parameters, so there is nothing for the description to add about parameter meanings. The baseline for parameterless tools is 4, and the description appropriately provides context about the resource being queried.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses the specific verb 'Get' and clearly identifies the resource as 'current dTelecom STT pricing information'. This unambiguously distinguishes the tool from siblings transcribe_file and stt_health, which focus on transcription and health status respectively.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The context is clear: the tool is for retrieving pricing information, and the note 'No payment required' clarifies that invoking it has no cost. It does not explicitly contrast with sibling tools, but the focused purpose makes the intended use easy to infer.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
transcribe_fileA
Transcribe a WAV audio file to text using dTelecom STT. The file must be PCM16, 16kHz, mono. Convert with: ffmpeg -i input.mp3 -ar 16000 -ac 1 -acodec pcm_s16le output.wav
| Name | Required | Description | Default |
|---|---|---|---|
| path | Yes | Absolute path to a WAV file (PCM16, 16kHz, mono) | |
| minutes | No | Session duration in minutes (5-120). Billed at $0.005/min | |
| language | No | Language code, e.g. en, ru, de, fr, es | en |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the behavioral disclosure burden. It does so by stating that only a specific WAV encoding is accepted and by offering a concrete workaround. It also indicates the output is text and implies an external STT dependency. It does not discuss latency or failure behavior, but transcription is an expected, non-mutating operation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences with no wasted words: the first states the action and required format, the second provides a practical command to prepare files. The most important information is front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a three-parameter transcription tool with no annotations and no output schema, the description gives the input constraints, the conversion path, and the expected result ('to text'). Billing and language details are left to the schema, which already documents them well.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents path, minutes, and language. The description adds value beyond the schema with the ffmpeg conversion command and reinforces the exact WAV requirements for the path parameter, helping an agent prepare a valid input.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Transcribe'), a resource ('WAV audio file'), and the service ('dTelecom STT'). This clearly distinguishes it from the siblings stt_pricing and stt_health, so an agent can tell at a glance what the tool does.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It clearly specifies the required input format (PCM16, 16kHz, mono) and even provides the exact ffmpeg command to convert an unsupported file to a valid input. It does not explicitly contrast with the siblings, but those are pricing and health endpoints, so the transcription use case is unambiguous.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections. Dates show when Glama detected each change.
3 tool updates
v0.2.0- First observed
stt_health - First observed
stt_pricing - First observed
transcribe_file
TDQS
Each tool has a single, clearly distinct purpose: transcribing audio, checking pricing, and checking service health. There is no overlap or ambiguity between them.
transcribe_file follows an action_noun pattern, while stt_pricing and stt_health follow a prefix_noun pattern. The names are readable, but the conventions are mixed.
Three tools is well-scoped for a focused STT service: one core operation plus two supporting informational tools. Each tool earns its place and there is no bloat.
The core transcription action is covered, and pricing and health checks round out the expected surface for this narrow service. No obvious missing operations are apparent from the provided scope.
Maintenance
Resources
Unclaimed servers have limited discoverability.
Looking for Admin?
If you are the server author, to access and configure the admin panel.
Related MCP Connectors
Transcribe audio & video to text for AI agents: 100+ languages, speaker labels, webhooks.
AI transcription from URLs or files. 119 languages, diarization, SRT/VTT/text export.
Pay-per-call profanity/explicit-content detection for AI agents. $0.005 USDC per call, no signup.
247 LLMs + image/video/voice/music gen + crypto/DeFi/markets/web-search. Pay-per-call USDC, no key.
Related MCP Servers
- AlicenseNot gradedqualityCmaintenanceThis service provides fast and reliable transcriptions for audio/video files and voice memos. It allows LLMs to interact with the text content of audio/video file.8MIT
- AlicenseNot gradedqualityDmaintenanceEnables speech-to-text transcription, text-to-speech synthesis, and audio analysis using Deepgram's AI models. Supports features like speaker diarization, sentiment analysis, language detection, and various audio processing capabilities.2MIT
- AlicenseAqualityFmaintenanceEnables AI assistants to transcribe audio files from URLs or local paths using AssemblyAI's services, with support for speaker diarization, language detection, and asynchronous job management through a standardized MCP interface.4252MIT
- AlicenseAqualityCmaintenanceEnables AI agents to transcribe audio and video with speaker labels, timestamps, and captions via Pepys API.963MIT
Latest Blog Posts
- Who's Calling? MCP Hosts Are an Identity Blind Spot (And the Spec Knows It)By Om-Shree-0709 on .mcpAgent IdentityOAuth 2.1
- Your AI Chatbot Just Exposed Your CEO's Salary to an InternBy Om-Shree-0709 on .Agent IdentityMCP SecurityOAuth Delegation
- Why MCP Servers Need Execution Sandboxing (And Why Your Current Stack Isn't Enough)By Om-Shree-0709 on .Agentic AiPrompt InjectionWebAssembly
MCP directory API
We provide all the information about MCP servers via our MCP API.
curl -X GET 'https://glama.ai/api/mcp/v1/servers/dTelecom/stt-mcp'
If you have feedback or need assistance with the MCP directory API, please join our Discord server