Skip to main content
Glama

@dtelecom/stt-mcp

MCP (Model Context Protocol) server for dTelecom real-time speech-to-text with x402 micropayments.

Lets AI assistants (Claude, Cursor, etc.) transcribe audio files using dTelecom STT — pay-per-use with USDC, no API keys needed.

Tools

Tool

Description

transcribe_file

Transcribe a WAV file (PCM16, 16kHz, mono) to text

stt_pricing

Get current pricing ($0.005/min)

stt_health

Check service health

Related MCP server: Deepgram MCP Server

Setup

1. Install

npm install -g @dtelecom/stt-mcp

2. Get a wallet

You need a private key with USDC. Either:

  • EVM (Base): Ethereum private key (0x hex) with USDC on Base — MetaMask, etc.

  • Solana: Solana private key (base58) with USDC on Solana — Phantom, Solflare, etc.

3. Configure your AI assistant

Claude Code (~/.claude.json):

{
  "mcpServers": {
    "dtelecom-stt": {
      "command": "dtelecom-stt-mcp",
      "env": {
        "DTELECOM_PRIVATE_KEY": "YOUR_PRIVATE_KEY"
      }
    }
  }
}

Claude Desktop (claude_desktop_config.json):

{
  "mcpServers": {
    "dtelecom-stt": {
      "command": "npx",
      "args": ["-y", "@dtelecom/stt-mcp"],
      "env": {
        "DTELECOM_PRIVATE_KEY": "YOUR_PRIVATE_KEY"
      }
    }
  }
}

Cursor (Settings > MCP Servers > Add):

{
  "dtelecom-stt": {
    "command": "npx",
    "args": ["-y", "@dtelecom/stt-mcp"],
    "env": {
      "DTELECOM_PRIVATE_KEY": "YOUR_PRIVATE_KEY"
    }
  }
}

4. Convert audio (if needed)

The tool accepts WAV files in PCM16 16kHz mono format. Convert with:

ffmpeg -i input.mp3 -ar 16000 -ac 1 -acodec pcm_s16le output.wav

Environment Variables

Variable

Required

Default

Description

DTELECOM_PRIVATE_KEY

Yes

EVM key (0x hex) or Solana key (base58)

DTELECOM_STT_URL

No

https://x402stt.dtelecom.org

STT service URL

Pricing

  • $0.005/min, billed per session

  • Minimum 5 minutes ($0.025)

  • Paid in USDC on Base or Solana via x402 protocol

  • No accounts, no API keys, no subscriptions

License

MIT

Available Tools

3 tools
stt_healthA

Check if the dTelecom STT service is running and healthy.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. The word 'check' implies a read-only, non-destructive operation, which is helpful. However, the description does not disclose what the response looks like or what 'healthy' means operationally, leaving some behavioral ambiguity for a health-check endpoint.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single sentence that is direct and free of filler. It states the resource, the action, and the expected result with no redundant words.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple zero-parameter health check, the description is largely sufficient. It names the service and the intent, and with no input schema requirements, an agent can invoke the tool without additional setup. The only minor gap is the lack of explicit return value information, but the purpose is clear enough for a health check.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has zero parameters, so the schema already fully covers parameter expectations. The description adds no parameter details, but none are needed; the baseline for a zero-parameter tool is 4.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('check') with a clear resource ('dTelecom STT service') and a specific outcome ('running and healthy'). This makes the tool's purpose immediately obvious and distinguishes it from its siblings (transcribe_file and stt_pricing), which handle different concerns.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description clearly implies when to use this tool: when an agent needs to verify the health or running status of the STT service. It does not explicitly mention alternatives or exclusions, but the sibling tools are so semantically distinct that no routing ambiguity is introduced.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

stt_pricingA

Get current dTelecom STT pricing information. No payment required.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are present, so the description carries the full burden. It conveys a read-only intent through 'Get' and adds the useful behavioral fact 'No payment required'. However, it does not disclose response format, data freshness, or any other operational constraints, though the simplicity of a zero-parameter pricing lookup reduces the risk.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is extremely concise: one short main clause stating the action and resource, plus a brief relevant note about payment. Every word earns its place, and the core purpose is front-loaded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-parameter, no-output-schema tool this description is largely sufficient: it explains what the tool provides and that it is free. It could mention the exact shape of the returned pricing data, but the triviality of the tool lowers the burden on the description.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema has zero parameters, so there is nothing for the description to add about parameter meanings. The baseline for parameterless tools is 4, and the description appropriately provides context about the resource being queried.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses the specific verb 'Get' and clearly identifies the resource as 'current dTelecom STT pricing information'. This unambiguously distinguishes the tool from siblings transcribe_file and stt_health, which focus on transcription and health status respectively.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The context is clear: the tool is for retrieving pricing information, and the note 'No payment required' clarifies that invoking it has no cost. It does not explicitly contrast with sibling tools, but the focused purpose makes the intended use easy to infer.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

transcribe_fileA

Transcribe a WAV audio file to text using dTelecom STT. The file must be PCM16, 16kHz, mono. Convert with: ffmpeg -i input.mp3 -ar 16000 -ac 1 -acodec pcm_s16le output.wav

ParametersJSON Schema
NameRequiredDescriptionDefault
pathYesAbsolute path to a WAV file (PCM16, 16kHz, mono)
minutesNoSession duration in minutes (5-120). Billed at $0.005/min
languageNoLanguage code, e.g. en, ru, de, fr, esen

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the behavioral disclosure burden. It does so by stating that only a specific WAV encoding is accepted and by offering a concrete workaround. It also indicates the output is text and implies an external STT dependency. It does not discuss latency or failure behavior, but transcription is an expected, non-mutating operation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two sentences with no wasted words: the first states the action and required format, the second provides a practical command to prepare files. The most important information is front-loaded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a three-parameter transcription tool with no annotations and no output schema, the description gives the input constraints, the conversion path, and the expected result ('to text'). Billing and language details are left to the schema, which already documents them well.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents path, minutes, and language. The description adds value beyond the schema with the ffmpeg conversion command and reinforces the exact WAV requirements for the path parameter, helping an agent prepare a valid input.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('Transcribe'), a resource ('WAV audio file'), and the service ('dTelecom STT'). This clearly distinguishes it from the siblings stt_pricing and stt_health, so an agent can tell at a glance what the tool does.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It clearly specifies the required input format (PCM16, 16kHz, mono) and even provides the exact ffmpeg command to convert an unsupported file to a valid input. It does not explicitly contrast with the siblings, but those are pricing and health endpoints, so the transcription use case is unambiguous.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections. Dates show when Glama detected each change.

  1. 3 tool updatesv0.2.0
    • First observedstt_health
    • First observedstt_pricing
    • First observedtranscribe_file

TDQS

A4.4/5.0
Disambiguation5/5

Each tool has a single, clearly distinct purpose: transcribing audio, checking pricing, and checking service health. There is no overlap or ambiguity between them.

Naming Consistency3/5

transcribe_file follows an action_noun pattern, while stt_pricing and stt_health follow a prefix_noun pattern. The names are readable, but the conventions are mixed.

Tool Count5/5

Three tools is well-scoped for a focused STT service: one core operation plus two supporting informational tools. Each tool earns its place and there is no bloat.

Completeness5/5

The core transcription action is covered, and pricing and health checks round out the expected surface for this narrow service. No obvious missing operations are apparent from the provided scope.

Maintenance

ActivityInactive
ResponsivenessUnresponsive

Resources

Unclaimed servers have limited discoverability.

Looking for Admin?

If you are the server author, to access and configure the admin panel.

Related MCP Connectors

Related MCP Servers

Latest Blog Posts

MCP directory API

We provide all the information about MCP servers via our MCP API.

curl -X GET 'https://glama.ai/api/mcp/v1/servers/dTelecom/stt-mcp'

If you have feedback or need assistance with the MCP directory API, please join our Discord server