Browser Agent MCP
Агент браузера MCP
Функции
Расширенная автоматизация браузера
Переходите к любому URL-адресу с помощью настраиваемых стратегий загрузки
Делайте снимки экрана всей страницы или отдельных элементов
Выполнять точные взаимодействия с DOM (щелчок, заполнение, выбор, наведение)
Выполнение произвольного JavaScript в контексте браузера с записью журналов консоли
Мощный API-клиент
Выполнение HTTP-запросов (GET, POST, PUT, PATCH, DELETE)
Настройте заголовки запроса и содержимое тела
Обработка данных ответа с использованием форматирования JSON
Обработка ошибок с подробной обратной связью
Управление ресурсами MCP
Доступ к журналам консоли браузера как к ресурсам
Извлечение снимков экрана через интерфейс ресурсов MCP
Постоянный сеанс с загруженным экземпляром браузера
Возможности ИИ-агента
Объединяйте несколько операций браузера для решения сложных задач
Следуйте многошаговым инструкциям с интеллектуальным устранением ошибок
Автоматизация технических задач с помощью инструкций на естественном языке
Related MCP server: Browserbeam MCP Server
Демо
Нажмите на любую временную метку, чтобы перейти к соответствующему разделу видео.
00:00 - Поиск MCP в Google
Переход на домашнюю страницу Google и поиск по запросу «Model Context Protocol». Демонстрация Claude Desktop с использованием интеграции MCP для выполнения базового веб-поиска и обработки результатов.
00:33 - Снимок экрана
Создание снимка экрана результатов поиска с пользовательским именем файла и демонстрация его в Finder. Показывает, как Claude может захватывать и сохранять визуальный контент с веб-страниц во время автоматизации браузера.
01:00 - Поиск в Википедии
Переход на Wikipedia.org и поиск по запросу «Model Context Protocol». Демонстрирует способность Клода взаимодействовать с различными веб-сайтами и их поисковой функциональностью посредством интеграции MCP.
01:38 - Взаимодействие с выпадающим меню I
Переход на тестовый веб-сайт (the-internet.herokuapp.com/dropdown) и выбор «Варианта 1» из выпадающего меню. Демонстрирует способность Клода взаимодействовать с элементами формы и делать выбор.
01:56 - Взаимодействие с выпадающим меню II
Изменение выбора на «Вариант 2» из того же выпадающего меню. Демонстрирует способность Клода манипулировать одним и тем же элементом формы несколько раз и делать разный выбор.
02:09 - Заполнение формы входа
Переход на страницу входа (the-internet.herokuapp.com/login) и заполнение поля имени пользователя «tomsmith» и поля пароля «SuperSecretPassword!». Демонстрирует автоматизацию заполнения форм.
02:28 - Отправка логина
Отправка учетных данных для входа и завершение процесса аутентификации. Демонстрирует способность Клода инициировать отправку форм и перемещаться по многоэтапным процессам.
02:36 — Выполнение API-запроса
Выполнение запроса GET к конечной точке API JSONPlaceholder. Демонстрирует способность Клода делать прямые вызовы API и обрабатывать возвращаемые данные посредством интеграции MCP.
Требования
Node.js 16 или выше
Клод Десктоп
Зависимости драматурга
Поддержка браузера
npm init playwright@latestЭтот пакет включает Playwright и необходимые зависимости для запуска автоматизации браузера. При запуске npm install будут установлены требуемые зависимости Playwright. Пакет поддерживает следующие браузеры:
Хром (по умолчанию)
Firefox
Microsoft Эдж
WebKit (движок Safari)
При первом использовании типа браузера Playwright автоматически установит соответствующие драйверы браузера по мере необходимости. Вы также можете установить их вручную с помощью следующих команд:
npx playwright install chrome
npx playwright install firefox
npx playwright install webkit
npx playwright install msedgeПримечание о Safari : Playwright не предоставляет прямую поддержку браузеру Safari. Вместо этого он использует WebKit, который является движком браузера, на котором работает Safari.
Примечание о Edge : при выборе Edge в качестве типа браузера агент фактически запустит Microsoft Edge (не Chromium). Технически, в Playwright Edge запускается с использованием экземпляра браузера Chromium с параметром канала 'msedge', поскольку Microsoft Edge основан на Chromium.
Установка
Установка вручную
Клонируйте или загрузите этот репозиторий:
git clone https://github.com/imprvhub/mcp-browser-agent
cd mcp-browser-agentУстановить зависимости:
npm installСоздайте проект:
npm run buildЗапуск сервера MCP
Существует два способа запуска сервера MCP:
Вариант 1: Запуск вручную
Откройте терминал или командную строку.
Перейдите в каталог проекта.
Запустите сервер напрямую:
node dist/index.jsДержите это окно терминала открытым при использовании Claude Desktop. Сервер будет работать, пока вы не закроете терминал.
Вариант 2: Автоматический запуск с помощью Claude Desktop (рекомендуется для регулярного использования)
Claude Desktop может автоматически запускать сервер MCP при необходимости. Чтобы настроить это:
Конфигурация
Файл конфигурации Claude Desktop находится по адресу:
macOS :
~/Library/Application Support/Claude/claude_desktop_config.jsonWindows :
%APPDATA%\Claude\claude_desktop_config.jsonLinux :
~/.config/Claude/claude_desktop_config.json
Отредактируйте этот файл, чтобы добавить конфигурацию Browser Agent MCP. Если файл не существует, создайте его:
{
"mcpServers": {
"browserAgent": {
"command": "node",
"args": ["ABSOLUTE_PATH_TO_DIRECTORY/mcp-browser-agent/dist/index.js",
"--browser",
"chrome"
]
}
}
}Важно : замените ABSOLUTE_PATH_TO_DIRECTORY на полный абсолютный путь , по которому вы установили MCP.
Пример для macOS/Linux:
/Users/username/mcp-browser-agentПример для Windows:
C:\\Users\\username\\mcp-browser-agent
Если у вас уже настроены другие MCP, просто добавьте раздел "browserAgent" внутри объекта "mcpServers". Вот пример конфигурации с несколькими MCP:
{
"mcpServers": {
"otherMcp1": {
"command": "...",
"args": ["..."]
},
"otherMcp2": {
"command": "...",
"args": ["..."]
},
"browserAgent": {
"command": "node",
"args": [
"ABSOLUTE_PATH_TO_DIRECTORY/mcp-browser-agent/dist/index.js",
"--browser",
"chrome"
]
}
}
}Выбор браузера
MCP Browser Agent поддерживает несколько типов браузеров. По умолчанию он использует Chrome, но вы можете указать другой браузер несколькими способами:
Вариант 1: Файл конфигурации
Создайте или отредактируйте файл .mcp_browser_agent_config.json в вашем домашнем каталоге:
{
"browserType": "chrome"
}Поддерживаемые значения для browserType :
chrome— использует установленный Chrome (по умолчанию)firefox- использует браузер Firefox «Nightly»webkit— использует движок WebKit (Примечание: это не сам Safari, а движок рендеринга WebKit, на котором работает Safari)edge— использует Microsoft Edge
Примечание о Safari : Playwright не обеспечивает прямую поддержку браузера Safari. Вместо этого он использует WebKit, который является движком браузера, на котором работает Safari. Реализация WebKit в Playwright обеспечивает схожую функциональность, но не идентична опыту браузера Safari.
Вариант 2: Аргумент командной строки
При ручном запуске сервера MCP можно указать тип браузера:
node dist/index.js --browser firefoxВариант 3: Переменная среды
Установите переменную среды MCP_BROWSER_TYPE :
MCP_BROWSER_TYPE=firefox node dist/index.jsВариант 4: Конфигурация рабочего стола Claude
При настройке MCP в claude_desktop_config.json Claude Desktop можно указать тип браузера:
{
"mcpServers": {
"browserAgent": {
"command": "node",
"args": [
"ABSOLUTE_PATH_TO_DIRECTORY/mcp-browser-agent/dist/index.js",
"--browser",
"chrome"
]
}
}
}Техническая реализация
MCP Browser Agent построен на Model Context Protocol, что позволяет Claude взаимодействовать с headful браузером через Playwright. Реализация состоит из четырех основных компонентов:
Сервер (index.ts)
Инициализирует сервер MCP с использованием стандартного протокола Model Context Protocol
Настраивает возможности сервера для инструментов и ресурсов
Устанавливает связь с Клодом через stdio-транспорт
Реестр инструментов (tools.ts)
Определяет схемы браузера и инструментов API
Указывает параметры, правила проверки и описания
Регистрирует инструменты на сервере MCP для обнаружения Клодом
Обработчики запросов (handlers.ts)
Управляет запросами протокола MCP на инструменты и ресурсы
Предоставляет журналы браузера и снимки экрана в качестве запрашиваемых ресурсов
Направляет запросы на выполнение инструмента соответствующим обработчикам
Исполнитель (executor.ts)
Управляет жизненным циклом браузера и API-клиента
Реализует функции автоматизации браузера с помощью Playwright
Обрабатывает запросы API с правильной обработкой ошибок и анализом ответов.
Поддерживает сеанс браузера с сохранением состояния между командами
Возможности агента
В отличие от базовых интеграций, MCP Browser Agent функционирует как настоящий агент ИИ:
Поддержание постоянного состояния браузера при выполнении нескольких команд
Сбор подробных журналов консоли для отладки
Хранение снимков экрана для справки и просмотра
Управление сложными последовательностями взаимодействия
Предоставление подробной информации об ошибке для восстановления
Поддержка цепочек операций для сложных рабочих процессов
Доступные инструменты
Инструменты браузера
Название инструмента | Описание | Параметры |
| Перейдите по URL-адресу |
|
| Сделать снимок экрана |
|
| Щелкните элемент |
|
| Заполните форму ввода |
|
| Выберите раскрывающийся список |
|
| Наведите курсор на элемент |
|
| Выполнить JavaScript |
|
API-инструменты
Название инструмента | Описание | Параметры |
| ПОЛУЧИТЬ запрос |
|
| POST-запрос |
|
| Запрос PUT |
|
| Запрос на исправление |
|
| Запрос на УДАЛИТЬ |
|
Доступ к ресурсам
Агент браузера MCP предоставляет следующие ресурсы:
browser://logs— доступ к журналам консоли браузераscreenshot://[name]— доступ к снимкам экрана по имени
Пример использования
Вот несколько реалистичных примеров использования MCP Browser Agent с Клодом:
Базовая навигация в браузере
Navigate to the Google homepage at https://www.google.comTake a screenshot of the current page and name it "google-homepage"Type "weather forecast" in the search boxПростые взаимодействия
Navigate to https://www.wikipedia.org and search for "Model Context Protocol"Go to https://the-internet.herokuapp.com/dropdown and select the option "Option 1" from the dropdownЗаполнение базовой формы
Navigate to https://the-internet.herokuapp.com/login and fill in the username field with "tomsmith" and the password field with "SuperSecretPassword!"Go to https://the-internet.herokuapp.com/login, fill in the username and password fields, then click the login buttonПростое выполнение JavaScript
Go to https://example.com and execute a JavaScript script to return the page titleNavigate to https://www.google.com and execute a JavaScript script to count the number of links on the pageБазовые API-запросы
Perform a GET request to https://jsonplaceholder.typicode.com/todos/1Make a POST request to https://jsonplaceholder.typicode.com/posts with appropriate JSON dataЭти примеры отражают реальные возможности агента браузера MCP и более реалистично отражают его возможности в текущем состоянии.
Поиск неисправностей
Ошибка «Сервер отключен»
Если вы видите ошибку «MCP Browser Agent: Server disconnected» в Claude Desktop:
Убедитесь, что сервер работает :
Откройте терминал и вручную запустите
node dist/index.jsиз каталога проекта.Если сервер запустится успешно, используйте Claude, оставив этот терминал открытым.
Проверьте вашу конфигурацию :
Убедитесь, что абсолютный путь в
claude_desktop_config.jsonправильный для вашей системы.Дважды проверьте, что вы использовали двойные обратные косые черты (
\\) для путей Windows.Убедитесь, что вы используете полный путь от корня вашей файловой системы.
Браузер не отображается
Если браузер не запускается или вы его не видите:
Проверьте, установлен ли указанный браузер.
Убедитесь, что в вашей системе установлен браузер (Chrome, Firefox, Edge или Safari/WebKit)
Драйверы браузера обрабатываются Playwright автоматически.
Перезагрузите сервер и Claude Desktop.
Завершите все существующие процессы узлов, которые могут запускать сервер.
Перезапустите Claude Desktop, чтобы установить новое соединение.
Процесс браузера не закрывается должным образом
Известны проблемы с браузерами Chromium и Chrome, когда процесс иногда не завершается должным образом после использования. Если у вас возникла эта проблема:
Закройте процесс браузера вручную :
Windows : нажмите Ctrl+Shift+Esc, чтобы открыть диспетчер задач, найдите процесс Chrome/Chromium и завершите его.
macOS : Откройте «Мониторинг системы» (Приложения > Утилиты > Мониторинг системы), найдите процесс Chrome/Chromium и нажмите X, чтобы завершить его.
Linux : выполните команду
ps aux | grep chromeилиps aux | grep chromiumчтобы найти процесс, затемkill <PID>, чтобы завершить его.
Примечание о совместимости браузеров :
Эта проблема наблюдалась в основном с Chromium и Chrome.
Встроенные браузеры Firefox и Playwright обычно не сталкиваются с этой проблемой.
[!ВНИМАНИЕ] Эта интеграция MCP построена на Playwright, который имеет известные проблемы и ошибки, которые могут повлиять на его работу. Пожалуйста, сообщайте о любых проблемах, с которыми вы сталкиваетесь при автоматизации браузера, в Playwright's GitHub issues . Команда Playwright постоянно работает над решением этих проблем, но этот агент обеспечивает основу для возможностей автоматизации браузера с Claude Desktop, несмотря на эти ограничения.
Разработка
Структура проекта
src/index.ts: Основная точка входа и инициализация сервера MCPsrc/tools.ts: Схемы инструментов и регистрацияsrc/handlers.ts: Обработчики запросов MCP для инструментов и ресурсовsrc/executor.ts: Логика реализации инструмента с использованием Playwright
Здание
npm run buildНаблюдение за изменениями
npm run watchТестирование
Проект включает в себя тесты для проверки основных функций и работы браузера.
npm test # Run tests
npm run test:watch # Watch mode
npm run test:coverage # Coverage reportТесты проверяют целостность конфигурации, функции автоматизации браузера, обработку ошибок и очистку процесса. Тестовый набор в частности фокусируется на обеспечении надлежащей обработки процессов браузера из-за известных проблем с завершением работы Chrome/Chromium.
Соображения безопасности
[!ВАЖНО] Эта интеграция MCP обеспечивает Claude возможностями автономного управления браузером. Ознакомьтесь с нашей Политикой безопасности для получения важной информации о запрещенных видах использования, последствиях для безопасности и передовых методах.
MCP Browser Agent разработан для законных задач автоматизации, но потенциально может быть использован не по назначению. Пользователи несут ответственность за обеспечение соответствия своего использования всем применимым законам, условиям обслуживания и этическим нормам. Для получения дополнительной информации см. нашу подробную Политику безопасности .
Внося вклад
Приветствуются вклады в MCP Browser Agent! Вот некоторые области, где вы можете помочь:
Добавление новых возможностей автоматизации браузера
Улучшение обработки ошибок и восстановления
Улучшение управления скриншотами и ресурсами
Создание полезных рабочих процессов и примеров
Оптимизация производительности для сложных операций
Лицензия
Данный проект лицензирован в соответствии с Mozilla Public License 2.0 — подробности см. в файле LICENSE .
Ссылки по теме
Available Tools
13 toolsapi_deleteB
Perform a DELETE request to an API endpoint
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | API endpoint URL | |
| headers | No | Request headers |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden but only states the action without disclosing side effects, idempotency, authentication needs, rate limits, or return format. This is minimal for an agent.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, front-loaded sentence with no wasted words. It is appropriately concise for a simple tool, though slightly more context could be included without losing efficiency.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's simplicity (2 params, no output schema), the description covers the basic action but omits expected return values, error conditions, or usage scope, leaving the description marginally adequate.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% with descriptive parameter names (url, headers). The description adds no additional meaning beyond the schema, so baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool performs a DELETE request to an API endpoint, using a specific verb and resource that distinguishes it from sibling tools (api_get, api_patch, etc.) which handle other HTTP methods.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is provided on when to use this tool versus alternatives like api_get or api_post. The description merely repeats the method, missing explicit when/when-not context or prerequisites.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
api_getC
Perform a GET request to an API endpoint
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | API endpoint URL | |
| headers | No | Request headers |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description must disclose behavioral traits. It fails to mention that GET is typically safe and idempotent, how errors are handled, or whether redirects are followed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single clear sentence with no unnecessary words. It efficiently conveys the core action.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool is simple but lacks information about return values, error handling, authentication requirements, or default behavior. The description is too minimal for an agent to use confidently.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% with descriptions for both parameters. The tool description adds no extra meaning beyond the schema, so baseline score of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states 'Perform a GET request to an API endpoint', which identifies the HTTP method and the action. It distinguishes from sibling tools like api_post and api_delete by the method name, but lacks mention of read-only nature.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance on when to use this tool versus siblings. The description does not specify that GET should be used for retrieving data, nor does it mention alternatives for modifying or deleting resources.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
api_patchB
Perform a PATCH request to an API endpoint
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | API endpoint URL | |
| data | Yes | Request body data (JSON string) | |
| headers | No | Request headers |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries full burden. It only states it performs a PATCH request, implying mutation, but does not disclose side effects, authentication needs, rate limits, or response behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single sentence with no unnecessary words. However, it is very brief and could benefit from additional context without becoming verbose.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description is incomplete given the tool's complexity. No output schema exists, and the description does not explain return values or error handling. Sibling tools suggest it is part of an HTTP client set, but the description lacks depth.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all parameters. The description does not add any additional meaning beyond what is in the input schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states it performs a PATCH request to an API endpoint, which is a specific verb and resource. This distinguishes it from sibling tools like api_get, api_post, etc.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is provided on when to use this tool versus other HTTP methods (e.g., POST, PUT) or alternatives. The context signals show siblings, but the description offers no differentiation advice.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
api_postB
Perform a POST request to an API endpoint
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | API endpoint URL | |
| data | Yes | Request body data (JSON string) | |
| headers | No | Request headers |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations provided, so description must disclose behavior. It fails to mention typical POST behavior (resource creation), data validation, or error handling.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Single sentence is concise and front-loaded, but it is too brief and lacks substance for a practical tool description.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple HTTP tool, description should mention typical use (e.g., 'sends data to URL'). Schema covers parameters but context is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% (all parameters described), so baseline is 3. Description adds no additional meaning beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Description states 'POST request' verb and 'API endpoint' resource, clearly distinguishing from sibling tools like api_get, api_put, etc.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance on when to use this tool versus alternatives, no mention of prerequisites or context such as authentication or data format.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
api_putC
Perform a PUT request to an API endpoint
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | API endpoint URL | |
| data | Yes | Request body data (JSON string) | |
| headers | No | Request headers |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations provided; description only states 'PUT request' implying mutation, but omits critical details like idempotency, side effects, authentication needs, or error behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Single sentence is concise and front-loaded, but does not add meaningful content beyond the tool name; minimal but not wasteful.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema and no behavioral details; for a tool with 3 parameters and nested objects, the description is insufficient to fully understand usage and return values.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema covers all 3 parameters with descriptions; description adds no extra meaning beyond what schema already provides, meeting the baseline for high coverage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Clearly states the action ('Perform a PUT request') and resource ('API endpoint'), but lacks differentiation from sibling tools like api_patch or api_post, which perform similar HTTP methods.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance on when to use PUT versus other HTTP methods (e.g., PATCH for partial updates, POST for creation). Does not mention idempotency or replacement semantics.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
browser_clickC
Click an element on the page
| Name | Required | Description | Default |
|---|---|---|---|
| selector | Yes | CSS selector for element to click |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description does not disclose behavioral traits beyond the basic action, such as potential side effects (e.g., navigation, page changes) or element visibility requirements. No annotations exist to compensate for this lack of detail.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, concise sentence with no unnecessary words. It is appropriately sized for a simple action, though it could benefit from slight elaboration without becoming verbose.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the sibling tools and the single parameter, the description is incomplete. It fails to mention crucial context like element visibility, clicks causing navigation, or waiting behavior, leaving the agent with insufficient information.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema fully describes the 'selector' parameter with a clear definition. The description adds no extra meaning beyond the schema, which is acceptable given 100% coverage, but it does not enhance understanding.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action ('Click an element') and the resource ('on the page'), distinguishing it from sibling tools like browser_fill or browser_hover. However, it could benefit from specifying that it operates within the current page context.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is provided on when to use this tool versus alternatives, such as when to click vs. hover, or prerequisites like page navigation. The description lacks context for proper selection.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
browser_evaluateB
Execute JavaScript in the browser context
| Name | Required | Description | Default |
|---|---|---|---|
| script | Yes | JavaScript code to execute |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description must fully disclose behavior. It mentions script execution but omits critical details: return value, side effects, permissions, sandboxing, or error handling. This is insufficient for a potentially powerful tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, clear sentence with no unnecessary words. It is optimally concise.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool has no output schema, so the description should explain return values or behavior. It does not. Additionally, it lacks details on execution context (e.g., async support, timeout). This leaves the agent underinformed.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The single required parameter 'script' has a description in the schema. The tool description adds no additional meaning beyond what the schema already provides, making it adequate but not additive.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description 'Execute JavaScript in the browser context' clearly states the action and resource. It is specific and distinct from sibling tools like browser_click or api_get.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use this tool versus alternatives (e.g., browser_navigate or API calls). An agent receives no context for decision-making.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
browser_fillB
Fill a form input with text
| Name | Required | Description | Default |
|---|---|---|---|
| value | Yes | Text to enter in the field | |
| selector | Yes | CSS selector for input field |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description should disclose behavioral details, but it only says 'fill' without specifying whether it overwrites existing text, waits for elements, or handles disabled fields. This leaves significant behavioral ambiguity.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single concise sentence, front-loading the action. However, it could be slightly expanded with useful context while remaining brief.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the simplicity of the tool (2 params, no output schema), the description is adequate but lacks details on behavior like clearing the field or submission. Sibling tools exist for other form actions, but no comparative guidance is provided.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Input schema covers 100% of parameters with descriptions (selector and value), so baseline is 3. The description adds no additional semantic information beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description 'Fill a form input with text' clearly states the action (fill) and the target (form input), distinguishing it from sibling tools like browser_click or browser_select which perform different actions.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is provided on when to use this tool versus alternatives such as browser_evaluate for setting values or browser_click for activation. There is no mention of prerequisites like element visibility or state.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
browser_hoverC
Hover over an element on the page
| Name | Required | Description | Default |
|---|---|---|---|
| selector | Yes | CSS selector for element to hover over |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description must disclose behavioral traits. It only states 'hover over an element' without explaining whether it triggers JavaScript events, waits for any transitions, or is safe. Essential behavioral context is missing.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is very concise at 6 words and front-loaded. It is not verbose, but the brevity may sacrifice necessary detail. It earns its place but could be slightly more informative.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's simplicity (1 parameter, no output schema), the description is incomplete. It does not mention the return value (likely void or success), side effects, or behavior after hovering. Sibling tools suggest a sequence of actions, but this tool's role is under-described.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% with a single parameter 'selector' described as 'CSS selector for element to hover over'. The description adds no additional meaning beyond what the schema already provides, so it meets the baseline but does not enhance understanding.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description 'Hover over an element on the page' clearly states the action (hover) and target (element on page). It differentiates from siblings like browser_click and browser_fill. However, it could be more specific about the effect (e.g., triggering hover state) but is not a tautology.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is provided on when to use this tool versus alternatives like browser_click or browser_evaluate. The description lacks any context about prerequisites, typical scenarios, or exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
browser_screenshotB
Capture a screenshot of the current page or a specific element
| Name | Required | Description | Default |
|---|---|---|---|
| mask | No | Selectors for elements to mask | |
| name | Yes | Identifier for the screenshot | |
| fullPage | No | Capture full page height | |
| savePath | No | Path to save screenshot (default: user's Downloads folder) | |
| selector | No | CSS selector for element to capture |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, and the description does not disclose behavioral traits such as whether the capture affects the page state, any authorization requirements, or rate limits. It only states the action, leaving the agent without important context about side effects or constraints.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single well-formed sentence that directly states the tool's purpose. It wastes no words, though it could benefit from slightly more detail without becoming verbose.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description is brief and lacks context about return values (e.g., image path), default behavior, or how the tool interacts with other browser tools. Given the five parameters and no output schema, the description should provide more operational context to be fully complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Since schema description coverage is 100%, the baseline is 3. The description does not add additional meaning beyond the parameter names and schema descriptions; for example, it doesn't explain when to use fullPage or mask options more concretely.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool captures a screenshot of the current page or a specific element, which is a specific verb-resource combination. It distinguishes itself from sibling tools like browser_navigate or browser_click by focusing on capture rather than navigation or element interaction.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies the tool is used when a screenshot is needed, but it provides no explicit guidance on when to use it versus alternative methods (e.g., browser_evaluate for custom captures) or any exclusions (e.g., not for video capture).
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
browser_selectC
Select an option from a dropdown menu
| Name | Required | Description | Default |
|---|---|---|---|
| value | Yes | Value or label to select | |
| selector | Yes | CSS selector for select element |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, and the description does not disclose behavioral traits such as whether it waits for options to load, supports custom dropdowns, triggers events, or requires scrolling. The description is too minimal to inform the agent of important behaviors.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single concise sentence that is front-loaded. However, it sacrifices completeness for brevity. It earns a 4 for being efficient, but could include more key details.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given that no annotations or output schema exist, the description is insufficiently complete. It does not explain return values, constraints, or behavior for complex dropdowns, leaving gaps for an AI agent.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema describes both parameters (selector and value) with 100% coverage. The description adds no additional meaning beyond the schema, so baseline score of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states it selects an option from a dropdown menu, using a specific verb and resource. It distinguishes from sibling tools like browser_fill (text input) and browser_click (clicking), though it could be more precise by specifying HTML <select> elements.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use this tool versus alternatives (e.g., browser_click for custom dropdowns) or any prerequisites. No exclusions or context are given.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
browser_set_viewportA
Change the browser's viewport size and scale factor
| Name | Required | Description | Default |
|---|---|---|---|
| width | No | Viewport width in pixels | |
| height | No | Viewport height in pixels | |
| deviceScaleFactor | No | Device scale factor (affects how content is scaled) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations exist, so the description should disclose behavioral traits. It only states what the tool changes, but not side effects (e.g., impact on screenshots, persistence across navigation). Lacks important context for a mutation tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Single sentence, no fluff. Every word is relevant. Excellent conciseness.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Adequately describes the core function, but lacks details on required fields, defaults, or behavioral context. Acceptable for a simple tool, but could be more complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, with meaningful descriptions for each parameter. The description adds no extra meaning beyond the schema, justifying the baseline score.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Description clearly states the action 'Change' and the target 'browser's viewport size and scale factor'. It is specific and distinct from sibling tools like browser_navigate or browser_click.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No explicit when-to-use guidance is provided. The description implies usage for adjusting viewport, but does not mention alternatives or exclusions. Minimal viable score.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
13 tool updates
v1.0.0- First observed
api_delete - First observed
api_get - First observed
api_patch - First observed
api_post - First observed
api_put - First observed
browser_click - First observed
browser_evaluate - First observed
browser_fill - First observed
browser_hover - First observed
browser_navigate - First observed
browser_screenshot - First observed
browser_select - First observed
browser_set_viewport
TDQS
Scored across 13 tools
Each tool has a clearly distinct purpose: API tools are differentiated by HTTP method, and browser tools cover unique interactions like clicking, hovering, filling, etc. No two tools overlap in function.
All tools follow a consistent '<domain>_<action>' pattern, with 'api_' prefix for HTTP methods and 'browser_' prefix for browser actions. Naming is unambiguous and predictable.
13 tools is well-scoped for a browser automation and API testing server. Each tool covers a fundamental operation without unnecessary bloat or gaps.
Core browser interactions (navigation, clicking, form filling, selecting, screenshot) and all major HTTP methods are covered. Minor omissions like file upload or wait-for-element are acceptable for this scope.
Maintenance
Related MCP Connectors
AI-powered web automation. Navigate websites using AI agents for one page or a thousand
AI-powered web automation. Navigate websites using AI agents for one page or a thousand
AI-powered browser automation — navigate, click, fill forms, and extract data from any website.
An agent-first office suite Claude & ChatGPT read and write over one MCP URL.
Related MCP Servers
- AlicenseBqualityDmaintenanceA browser automation agent that enables Claude to interact with web browsers through the Model Context Protocol, allowing for actions like navigating websites, manipulating elements, and managing browser state.29MIT
- AlicenseNot gradedqualityDmaintenanceEnables real browser automation as tools in Cursor, Claude Desktop, Windsurf, and any MCP-compatible client, allowing AI agents to interact with web pages through natural language.17 npmMIT
- FlicenseNot gradedqualityBmaintenanceEnables Claude Code to control a real browser using AI for web scraping, competitive intelligence, and UX auditing through the MCP protocol.-
- AlicenseNot gradedqualityDmaintenanceEnables browser automation through the Claude Chrome Extension, allowing agents to navigate websites, fill forms, take screenshots, and debug web apps via standard MCP protocols.1MIT