Skip to main content
Glama
97,560 servers. Updated

Matching MCP tools:

Matching MCP Connectors:

"A tool for scraping websites with images" matching MCP servers:

GET /v1/servers – MCP directory API reference
  • A
    license
    Not graded
    quality
    D
    maintenance
    Enables chat-driven audio analysis and enhancement using local Claude, including denoising, EQ, compression, and loudness normalization, with an A/B viewer for synchronized comparison.
    1
    MIT
  • A
    license
    Not graded
    quality
    B
    maintenance
    Enables voice assistants to find nearby service providers, connect users to live video representatives for natural conversation, and handle appointment booking. It also returns confirmed bookings so the assistant can add them to its calendar.
    MIT
  • A
    license
    A
    quality
    C
    maintenance
    MCP server that provides a transcribe_audio tool to convert voice messages from channels into text using OpenAI Whisper, enabling Claude Code to process audio attachments.
    1
    MIT
  • A
    license
    A
    quality
    A
    maintenance
    Your agent seeks what search can't find. A self-hosted perception MCP server that transcribes speech, reads behind logins, sees images and video frames, crosses languages, and remembers.
    18
    48 PyPI
    80
    Apache 2.0
  • A
    license
    A
    quality
    B
    maintenance
    Enables generating images, video, audio, and speech from MCP clients using your own Vidofy account, with access to hundreds of models for text-to-video, image-to-video, image editing, lipsync, text-to-speech, and voice cloning.
    9
    82 npm
    2
    MIT
  • A
    license
    A
    quality
    B
    maintenance
    Translate PDF, Word (DOCX), Excel (XLSX) and PowerPoint (PPTX) files, images, subtitles and text. Turn audio and video into transcripts or translated subtitles. Preserve layout where supported. Requires an Equalang API key and credits.
    8
    61 npm
    1
    Apache 2.0
  • A
    license
    A
    quality
    B
    maintenance
    Enables AI agents to generate spoken audio locally on a GPU via the Chatterbox models, with zero-shot voice cloning from reference clips, multilingual output, and sentence-aligned chunking for long text.
    5
    1
    MIT
  • A
    license
    A
    quality
    B
    maintenance
    Provides local vision and audio perception for MCP-compatible agents, enabling them to read images, transcribe text from visual media, analyze videos, and convert speech to text entirely on-device. It is privacy-focused with no cloud upload or API keys by default, using Ollama for inference.
    6
    1
    MIT
  • A
    license
    A
    quality
    B
    maintenance
    Real-time English ↔ Mandarin Chinese speech translation for Claude. Transcribes audio locally with Whisper, translates via Claude API, and synthesises speech locally with Piper TTS. Pass a WAV file path. Claude handles the rest.
    3
    66 npm
    3
    MIT
  • F
    license
    A
    quality
    F
    maintenance
    Provides various AI capabilities through DeepInfra's OpenAI-compatible API including image generation, text processing, embeddings, speech recognition, object detection, and classification tasks. Enables users to access multiple AI models for diverse tasks like generating images from prompts, transcribing audio, analyzing text sentiment, and performing computer vision operations.
    10
    2
    -
  • A
    license
    B
    quality
    D
    maintenance
    A Node.js server that enables AI assistants to interact with Bouyomi-chan's text-to-speech functionality through Model Context Protocol (MCP), allowing for voice reading of text with adjustable parameters.
    1
    2
    MIT