Skip to main content
Glama
97,619 servers. Updated
20 Best Browser Automation MCP Servers: compared and ranked, October 2026Ranked from 3,678 matching servers on stars, growth, downloads and maintenance. Updated .

Matching MCP tools:

Matching MCP Connectors:

"Interacting with a webpage using a browser plugin and live browser instance" matching MCP servers:

GET /v1/servers – MCP directory API reference
  • A
    license
    B
    quality
    F
    maintenance
    Enables text-to-speech functionality on macOS using the say command, offering extensive control over speech parameters like voice, rate, volume, and pitch for a customizable auditory experience.
    2
    14 npm
    20
    MIT
  • A
    license
    B
    quality
    D
    maintenance
    A Model Context Protocol server that provides text-to-speech functionality for AI agents using Microsoft Edge's text-to-speech technology, supporting multiple voices, languages, and voice customization.
    2
    8
    MIT
  • A
    license
    B
    quality
    Not graded
    maintenance
    A PowerShell-based MCP server that enables Claude Desktop to convert text to speech using Windows' built-in Speech API, offering features like playback control, speed and volume adjustment.
    10
    MIT
  • F
    license
    C
    quality
    C
    maintenance
    An MCP server that exposes speech-to-text and text-to-speech capabilities using a local speaches instance, allowing AI assistants to transcribe audio and generate speech.
    2
    -
  • A
    license
    Not graded
    quality
    C
    maintenance
    Enables assistants to open scripts in a centered, multilingual reader with direct paste, locale-aware word tracking, manual sentence navigation, pause/auto-scroll, and optional browser or OpenAI speech recognition.
    MIT
  • A
    license
    Not graded
    quality
    D
    maintenance
    Enables AI assistants to convert text to high-quality speech audio using MeloTTS. Automatically splits long texts into segments, generates WAV files, and merges them using ffmpeg with support for multiple languages and customizable speech parameters.
    MIT
  • A
    license
    Not graded
    quality
    D
    maintenance
    Enables agents to convert text to speech using OpenAI's TTS models with voice selection, delivery instructions, and queue-based audio playback. Supports both blocking and non-blocking modes for flexible audio generation and playback control.
    3
    BSD 3-Clause
  • A
    license
    Not graded
    quality
    B
    maintenance
    A text-to-speech MCP server with 48 voices across 9 languages, supporting emotion spans, SFX tags, and multi-speaker dialogue. Deployable via a single npx command with built-in guardrails and swappable backends.
    MIT
  • A
    license
    Not graded
    quality
    B
    maintenance
    Enables MCP clients to speak text aloud on macOS using the built-in say command, with no external speech API or keys, and to list available voices and control rate or voice selection. Requests from multiple client threads share a single FIFO playback worker, so speech plays one item at a time and can be inspected, stopped, or muted in hold or discard mode.
    3 npm
    MIT
  • A
    license
    Not graded
    quality
    B
    maintenance
    Enables AI agents to ring or text a user's actual iPhone, read a question or update aloud, and return the spoken answer back to the agent as text. Calls and messages land in a single titled thread per agent, with tools for polling results, waiting for replies, and labeling conversations.
    22
    MIT
  • A
    license
    Not graded
    quality
    B
    maintenance
    Enables Alexa+ and other voice agents to tell personalized, feelings-aware bedtime stories using a child's saved hero profile, with interactive choices, timed wind-down stories, and optional MP3 narration.
    MIT
  • A
    license
    Not graded
    quality
    B
    maintenance
    Enables coding agents to hold spoken conversations by exposing speech-to-text and text-to-speech tools over MCP, routing them to a local daemon that drives a Chrome window using on-device speech recognition and Piper synthesis. It supports listening, speaking, and reading files aloud within a Claude Code plugin or an MCP gateway.
    4 npm
    MIT
  • A
    license
    Not graded
    quality
    A
    maintenance
    Enables MCP-compatible agents to control a desktop mascot by switching its appearance and voice, showing speech-bubble dialogue, and managing state such as affinity.
    31 npm
    1
    MIT