Skip to main content
Glama
97,619 servers. Updated
20 Best Browser Automation MCP Servers: compared and ranked, October 2026Ranked from 3,678 matching servers on stars, growth, downloads and maintenance. Updated .

Matching MCP tools:

Matching MCP Connectors:

"A browser focused on stealth or privacy features" matching MCP servers:

GET /v1/servers – MCP directory API reference
  • A
    license
    B
    quality
    Not graded
    maintenance
    A PowerShell-based MCP server that enables Claude Desktop to convert text to speech using Windows' built-in Speech API, offering features like playback control, speed and volume adjustment.
    10
    MIT
  • A
    license
    Not graded
    quality
    C
    maintenance
    Enables assistants to open scripts in a centered, multilingual reader with direct paste, locale-aware word tracking, manual sentence navigation, pause/auto-scroll, and optional browser or OpenAI speech recognition.
    MIT
  • A
    license
    Not graded
    quality
    B
    maintenance
    Enables MCP clients to speak text aloud on macOS using the built-in say command, with no external speech API or keys, and to list available voices and control rate or voice selection. Requests from multiple client threads share a single FIFO playback worker, so speech plays one item at a time and can be inspected, stopped, or muted in hold or discard mode.
    3 npm
    MIT
  • A
    license
    Not graded
    quality
    C
    maintenance
    Enables AI agents to ring or text a user's actual iPhone, read a question or update aloud, and return the spoken answer back to the agent as text. Calls and messages land in a single titled thread per agent, with tools for polling results, waiting for replies, and labeling conversations.
    21
    MIT
  • A
    license
    Not graded
    quality
    B
    maintenance
    Exposes a single transcribe tool over streamable HTTP so containerised agents can send a media file name and receive text transcribed locally by MacWhisper on the host Mac's GPU, with token-gated access and no uploads or API keys. Callers place media in a configured directory, and the blocking call returns the finished transcript.
    MIT
  • A
    license
    Not graded
    quality
    B
    maintenance
    Enables coding agents to hold spoken conversations by exposing speech-to-text and text-to-speech tools over MCP, routing them to a local daemon that drives a Chrome window using on-device speech recognition and Piper synthesis. It supports listening, speaking, and reading files aloud within a Claude Code plugin or an MCP gateway.
    7 npm
    MIT
  • A
    license
    Not graded
    quality
    C
    maintenance
    An MCP server for AI voice synthesis with an inline audio player, allowing users to give their AI assistant a custom cloned voice using DashScope or ElevenLabs.
    MIT
  • A
    license
    Not graded
    quality
    C
    maintenance
    Enables text-to-speech synthesis using VOICEVOX Web API with customizable speaker selection. Features a specialized tool for generating speech as Asuka Langley from Evangelion and provides access to available speaker lists.
    Apache 2.0
  • F
    license
    Not graded
    quality
    D
    maintenance
    Enables Claude to act as an autonomous radio DJ, generating live-coded music with Strudel, making text-to-speech announcements, and responding to audience requests through a browser UI.
    -
  • F
    license
    Not graded
    quality
    D
    maintenance
    A desktop avatar companion for Claude Code that provides a visible on-screen presence with text-to-speech, speech-to-text, and an interactive avatar overlay.
    -
  • A
    license
    A
    quality
    A
    maintenance
    Enables spoken conversations with Claude on a Mac: Claude speaks through speakers, listens to the user's natural replies, and transcribes them locally. No audio leaves the computer, and it includes tools for voice setup and a hands-free voice mode.
    2
    632 npm
    MIT
  • A
    license
    A
    quality
    A
    maintenance
    Voice interface for Claude Code: you talk, the agent listens, codes, and talks back while it works. Live speech-to-text with turn-taking, Grok/xAI voices with per-subagent personas, and a real-time HUD dashboard.
    4
    12
    MIT
  • A
    license
    A
    quality
    D
    maintenance
    Enables local text-to-speech synthesis for Claude and Cursor using Supertonic 3, with support for multiple voices, expressions, and languages. No API key or cloud required.
    3
    MIT