Scrapybara MCP
The Scrapybara MCP server allows clients to interact with virtual Ubuntu desktops for web browsing, code execution, and more.
Start an instance: Launch a Scrapybara Ubuntu desktop sandbox and get a stream URL for real-time viewing.
Get instances: Retrieve all running Scrapybara instances.
Stop an instance: Terminate a specific instance using its ID.
Run bash commands: Execute shell commands in a specified instance.
Act on instances: Control an instance through an agent that can perform mouse/keyboard actions and execute commands based on prompts, with optional structured output.
Click on "Install Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@Scrapybara MCPstart a new Ubuntu instance and open the browser to github.com"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
A Model Context Protocol server for Scrapybara. This server enables MCP clients such as Claude Desktop, Cursor, and Windsurf to interact with virtual Ubuntu desktops and take actions such as browsing the web, running code, and more.
Prerequisites
Node.js 18+
pnpm
Scrapybara API key (get one at scrapybara.com)
Related MCP server: taw-computer
Installation
Clone the repository:
git clone https://github.com/scrapybara/scrapybara-mcp.git
cd scrapybara-mcpInstall dependencies:
pnpm installBuild the project:
pnpm buildAdd the following to your MCP client config:
{
"mcpServers": {
"scrapybara-mcp": {
"command": "node",
"args": ["path/to/scrapybara-mcp/dist/index.js"],
"env": {
"SCRAPYBARA_API_KEY": "<YOUR_SCRAPYBARA_API_KEY>",
"ACT_MODEL": "<YOUR_ACT_MODEL>", // "anthropic" or "openai"
"AUTH_STATE_ID": "<YOUR_AUTH_STATE_ID>" // Optional, for authenticating the browser
}
}
}
}Restart your MCP client and you're good to go!
Tools
start_instance - Start a Scrapybara Ubuntu instance. Use it as a desktop sandbox to access the web or run code. Always present the stream URL to the user afterwards so they can watch the instance in real time.
get_instances - Get all running Scrapybara instances.
stop_instance - Stop a running Scrapybara instance.
bash - Run a bash command in a Scrapybara instance.
act - Take action on a Scrapybara instance through an agent. The agent can control the instance with mouse/keyboard and bash commands.
Contributing
Scrapybara MCP is a community-driven project. Whether you're submitting an idea, fixing a typo, adding a new tool, or improving an existing one, your contributions are greatly appreciated!
Before contributing, read through the existing issues and pull requests to see if someone else is already working on something similar. That way you can avoid duplicating efforts.
If there are more tools or features you'd like to see, feel free to suggest them on the issues page.
Available Tools
5 toolsactB
Take action on a Scrapybara instance through an agent. The agent can control the instance with mouse/keyboard and bash commands.
| Name | Required | Description | Default |
|---|---|---|---|
| instance_id | Yes | The ID of the instance to act on. | |
| prompt | Yes | The prompt to act on. <EXAMPLES> - Go to https://ycombinator.com/companies, set batch filter to W25, and extract all company names. - Find the best way to contact Scrapybara. - Order a Big Mac from McDonald's on Doordash. </EXAMPLES> | |
| schema | No | Optional schema if you want to extract structured output. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries full burden. It mentions the agent can control the instance with mouse/keyboard and bash commands, which hints at behavioral traits like interactivity and command execution. However, it doesn't disclose critical details such as whether this is a read/write operation, potential side effects (e.g., data changes), authentication needs, rate limits, or error handling. For a tool with no annotations, this leaves significant gaps in transparency.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is concise with two sentences that efficiently state the tool's purpose and capabilities. It's front-loaded with the main action and avoids unnecessary details. However, it could be slightly more structured by explicitly separating the agent's roles or adding brief context, but overall it's well-sized and wastes no words.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the complexity of controlling an instance with an agent (implying interactive or automated tasks), no annotations, and no output schema, the description is incomplete. It doesn't explain what the tool returns (e.g., success status, extracted data), potential errors, or behavioral nuances. For a tool with 3 parameters and significant implied functionality, more context is needed to make it fully usable by an AI agent.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all parameters (instance_id, prompt, schema). The description adds minimal semantic value beyond the schema: it implies 'prompt' is for action instructions and 'instance_id' targets a specific instance, but doesn't elaborate on format or constraints. With high schema coverage, the baseline is 3, and the description doesn't significantly enhance parameter understanding.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose: 'Take action on a Scrapybara instance through an agent' with specific verbs like 'control' and actions like 'mouse/keyboard and bash commands'. It distinguishes from siblings like 'bash' (which might be standalone) and 'get_instances' (which is read-only), though it doesn't explicitly name alternatives. The purpose is specific but could be more precise about the agent's role.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage by mentioning controlling an instance with an agent, suggesting it's for interactive or automated tasks on Scrapybara. However, it lacks explicit guidance on when to use this vs. siblings like 'bash' (e.g., for direct commands) or 'start_instance' (for setup). No exclusions or prerequisites are stated, leaving usage context inferred rather than clearly defined.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
bashC
Run a bash command in a Scrapybara instance.
| Name | Required | Description | Default |
|---|---|---|---|
| instance_id | Yes | The ID of the instance to run the command on. | |
| command | Yes | The command to run in the instance shell. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It states the action but doesn't reveal critical traits like whether this is a read-only or destructive operation, potential security implications, rate limits, or error handling. For a tool that runs arbitrary commands, this lack of transparency is a significant gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, efficient sentence that directly states the tool's purpose without unnecessary words. It is front-loaded and wastes no space, making it easy to parse quickly.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the complexity of running arbitrary bash commands and the absence of annotations and output schema, the description is incomplete. It doesn't cover behavioral aspects like safety, permissions, or output format, which are crucial for such a tool. This leaves the agent with insufficient context for reliable use.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema has 100% description coverage, clearly documenting both parameters ('instance_id' and 'command'). The description adds no additional meaning beyond what the schema provides, such as command syntax examples or instance state requirements. Baseline 3 is appropriate since the schema does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action ('Run') and target ('bash command in a Scrapybara instance'), making the purpose understandable. However, it doesn't differentiate from sibling tools like 'act' or 'start_instance', which might also involve instance operations, leaving some ambiguity about when this specific tool is appropriate versus others.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use this tool versus alternatives like 'act' or 'start_instance'. It lacks context about prerequisites, such as whether the instance must be running, or exclusions, such as not using it for non-bash commands. This leaves the agent without clear usage instructions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_instancesB
Get all running Scrapybara instances.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It states what the tool does but lacks behavioral details such as whether it requires authentication, how it handles errors, or what the return format looks like. This is a significant gap for a tool with zero annotation coverage.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, efficient sentence with no wasted words. It's front-loaded and directly states the tool's purpose without unnecessary elaboration.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given no annotations and no output schema, the description is incomplete. It lacks details on behavioral traits, return values, or error handling, which are crucial for a tool that likely interacts with system resources like instances.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has 0 parameters, and schema description coverage is 100%, so no parameter information is needed. The description doesn't add param details, but that's appropriate here, meeting the baseline for zero parameters.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the verb ('Get') and resource ('all running Scrapybara instances'), making the purpose unambiguous. It doesn't explicitly differentiate from siblings like 'start_instance' or 'stop_instance', but the action is distinct enough to imply differentiation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is provided on when to use this tool versus alternatives like 'start_instance' or 'stop_instance'. The description implies usage for listing instances but doesn't specify prerequisites, timing, or exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
start_instanceA
Start a Scrapybara Ubuntu instance. Use it as a desktop sandbox to access the web or run code. Always present the stream URL to the user afterwards so they can watch the instance in real time.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden. It discloses key behavioral traits: the tool creates a desktop sandbox for web access or code execution, and it generates a stream URL for real-time viewing. However, it lacks details on permissions, rate limits, or error handling.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with the core action, followed by usage context and a critical post-action step. Every sentence adds essential information with zero waste, making it highly efficient.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity (starting an instance with streaming) and lack of annotations or output schema, the description is mostly complete. It covers purpose, usage, and output behavior, but could benefit from details on instance specifications or error cases.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema has 0 parameters with 100% coverage, so no parameter documentation is needed. The description adds value by explaining the tool's purpose and output behavior, justifying a baseline score above the minimum.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the specific action ('Start') and resource ('a Scrapybara Ubuntu instance'), and distinguishes it from siblings like 'stop_instance' and 'get_instances' by focusing on initiation rather than termination or querying.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explicitly states when to use this tool ('Use it as a desktop sandbox to access the web or run code') and provides a clear post-action directive ('Always present the stream URL to the user afterwards'), offering strong guidance without mentioning alternatives.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
stop_instanceC
Stop a running Scrapybara instance.
| Name | Required | Description | Default |
|---|---|---|---|
| instance_id | Yes | The ID of the instance to stop. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It states the action ('Stop') but lacks details on effects (e.g., whether data is preserved, permissions required, or if the operation is reversible). This is a significant gap for a mutation tool with zero annotation coverage.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, efficient sentence with zero waste, clearly front-loading the core action. It's appropriately sized for a simple tool with one parameter, earning its place without unnecessary elaboration.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity as a mutation operation with no annotations and no output schema, the description is incomplete. It lacks information on behavioral outcomes, error conditions, or return values, which is inadequate for guiding an agent in safe and effective use.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, with the single parameter 'instance_id' documented in the schema. The description adds no additional parameter details beyond what the schema provides, so it meets the baseline for high schema coverage without compensating value.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action ('Stop') and target resource ('a running Scrapybara instance'), making the purpose immediately understandable. However, it doesn't differentiate from sibling tools like 'start_instance' or 'get_instances' beyond the obvious verb difference, missing explicit comparison.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is provided on when to use this tool versus alternatives. It doesn't mention prerequisites (e.g., the instance must be running), exclusions, or comparisons to siblings like 'start_instance' or 'act', leaving usage context implied but unspecified.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
TDQS
Each tool has a clearly distinct purpose: act controls the instance via agent, bash runs commands, get_instances lists instances, start_instance creates one, and stop_instance terminates one. There is no overlap or ambiguity between these functions.
Four tools follow a consistent verb_noun pattern (get_instances, start_instance, stop_instance, run_command would fit but bash is used instead). The tool 'act' deviates as a single verb, and 'bash' is a noun, but overall the naming is mostly consistent and readable.
With 5 tools, this is well-scoped for managing Scrapybara instances. It covers core lifecycle operations (start, stop, list) and usage actions (act, bash), with each tool earning its place without bloat or thinness.
The toolset provides good coverage for instance management and interaction, including lifecycle (start, stop, list) and control (act, bash). A minor gap is the lack of a tool for configuring or modifying instances, but agents can work around this using bash or act.
Maintenance
Resources
Unclaimed servers have limited discoverability.
Looking for Admin?
If you are the server author, to access and configure the admin panel.
Related MCP Connectors
A comprehensive Model Context Protocol (MCP) server that enables AI assistants to interact with yo…
A comprehensive Model Context Protocol (MCP) server that enables AI assistants to control Unreal E…
A Model Context Protocol server for Wix AI tools
The Telnyx MCP server is an official implementation of the Model Context Protocol that enables AI clients (like Claude Desktop, Cursor, and OpenAI Agents) to interact with Telnyx's telephony, messaging, and AI assistant APIs. It provides comprehensive capabilities including making and managing phone calls, sending SMS/MMS messages, purchasing and configuring phone numbers, creating AI assistants with custom instructions, managing cloud storage buckets, scraping and embedding website content, and handling integration secrets. The server exists as both a local implementation and a remotely hosted version, allowing developers to integrate real-world communication infrastructure directly into AI applications.
Related MCP Servers
- AlicenseNot gradedqualityDmaintenanceA desktop automation MCP server that enables AI agents to interact with Linux environments through screenshots, window inspection, and input simulation. It provides tools for mouse control, keyboard input, and screen capture using xdotool and XDG Desktop Portals.MIT
- AlicenseBqualityDmaintenanceAn MCP server that provides AI agents with a full Ubuntu desktop environment inside Docker, enabling them to perform complex computer tasks like browsing, coding, testing, and GUI automation.368MIT
- AlicenseAqualityDmaintenanceA Model Context Protocol server that enables AI assistants to create, control, and interact with virtual desktop environments through E2B's secure cloud sandboxes.6224MIT
- AlicenseNot gradedqualityDmaintenanceA Model Context Protocol server that gives any AI full control over a Windows PC, enabling application control, mouse/keyboard automation, screen capture, browser automation, and more.4MIT
Latest Blog Posts
- Who's Calling? MCP Hosts Are an Identity Blind Spot (And the Spec Knows It)By Om-Shree-0709 on .mcpAgent IdentityOAuth 2.1
- Your AI Chatbot Just Exposed Your CEO's Salary to an InternBy Om-Shree-0709 on .Agent IdentityMCP SecurityOAuth Delegation
- Why MCP Servers Need Execution Sandboxing (And Why Your Current Stack Isn't Enough)By Om-Shree-0709 on .Agentic AiPrompt InjectionWebAssembly
MCP directory API
We provide all the information about MCP servers via our MCP API.
curl -X GET 'https://glama.ai/api/mcp/v1/servers/Scrapybara/scrapybara-mcp'
If you have feedback or need assistance with the MCP directory API, please join our Discord server