capskip-mcp
OfficialThis MCP server lets AI agents solve CAPTCHAs locally through CapSkip, returning tokens or answers so the agent can continue browser automation.
Check whether the CapSkip desktop app is running and reachable with
capskip_status.Solve image/text CAPTCHAs from a file path, URL, data URI, or base64 string.
Solve reCAPTCHA v2 (checkbox and invisible), v3, and Enterprise; returns a
g-recaptcha-responsetoken.Solve Cloudflare Turnstile widgets and interstitial pages; returns a token plus the exact
userAgentto submit it with.Solve GeeTest v3; returns
challenge,validate, andseccodefor form submission.Solve ALTCHA proof-of-work challenges from either
challenge_urlorchallenge_json; returns the token for thealtchafield.Also supports Capy Puzzle, CaptchaFox, and Friendly Captcha (per the README).
Proxies are supported on most solve tools except image captchas.
Returns readable error messages instead of raw protocol failures, so an agent can diagnose and retry.
Provides tools to solve Cloudflare Turnstile challenges, returning a token for form submission.
Provides tools to solve Google reCAPTCHA v2/v3/Enterprise challenges, returning a token for form submission.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@capskip-mcpsolve the reCAPTCHA on this page and submit the form"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
CapSkip MCP Server — Unlimited Captcha Solver for AI Agents
A captcha solver MCP server that lets AI agents solve reCAPTCHA, Cloudflare Turnstile, GeeTest, ALTCHA, Capy Puzzle, CaptchaFox, Friendly Captcha and image captchas instead of stalling on them.
Works with Claude Desktop, Claude Code, Cursor, VS Code, and any Model Context Protocol client. Powered by CapSkip — a local captcha solver that runs on your own machine, licensed once rather than billed per solve.
npx -y capskip-mcpWhat this solves
An AI agent driving a browser hits a captcha and stops. This server gives it nine tools so it can read the sitekey, solve the challenge, and carry on — without a human stepping in and without a per-solve API bill.
CapSkip runs as a desktop app exposing a captcha-solving HTTP API on 127.0.0.1:8080. capskip-mcp is a thin translation layer over that API: the fifth official CapSkip client, alongside the Python, Node.js, PHP and .NET SDKs.
Related MCP server: ma-browser
Supported captcha types
Captcha | Tool | Notes |
reCAPTCHA v2 solver (checkbox) |
| Returns a |
reCAPTCHA v2 invisible solver |
| Pass |
reCAPTCHA Enterprise solver |
| Pass |
reCAPTCHA v3 solver |
| Pass |
Cloudflare Turnstile solver |
| Widget and interstitial challenge pages |
GeeTest v3 solver |
| Slide puzzle; returns challenge/validate/seccode |
ALTCHA solver |
| Proof-of-work; returns the token for the |
Capy Puzzle solver |
| Slide puzzle; returns three form values, not a token |
CaptchaFox solver |
| Returns the token for the |
Friendly Captcha solver |
| Proof-of-work; say which version (v1 or v2) |
Image captcha solver (text/OCR) |
| File path, URL, data URI, or base64 |
Not supported: hCaptcha and FunCaptcha/Arkose. There is no tool for them and capskip_solve_recaptcha will not work on one. hCaptcha is the easiest to misidentify since it also carries a data-sitekey — check for class="h-captcha" or a js.hcaptcha.com script first.
Try them against live widgets on the captcha demo pages.
Quick start (5 minutes)
1. Install the CapSkip captcha solver
Download and run the CapSkip desktop app from capskip.com. Leave it running in the background.
In CapSkip settings, note the API port (default 8080) and API key (optional — if key validation is disabled, any string works).
2. Add capskip-mcp to your MCP client
No install step — npx fetches and runs it on demand.
{
"mcpServers": {
"capskip": {
"command": "npx",
"args": ["-y", "capskip-mcp"],
"env": {
"CAPSKIP_HOST": "127.0.0.1",
"CAPSKIP_PORT": "8080",
"CAPSKIP_API_KEY": "capskip"
}
}
}
}MCP client | Config file | Example |
Claude Desktop |
| |
Claude Code |
| |
Cursor |
| |
VS Code |
|
3. Restart your client
It should list nine tools, all prefixed capskip_.
4. Ask your agent to solve a captcha
"Call capskip_status to confirm CapSkip is running, then solve the reCAPTCHA on this page and submit the form."
The agent reads the sitekey off the page, calls capskip_solve_recaptcha, and places the returned token in the page's g-recaptcha-response field.
Why a local captcha solver
Cloud captcha APIs bill per solve, so an agent that retries is an agent that costs money, and every page URL and sitekey you solve leaves your network.
CapSkip runs on your machine:
Unlimited captcha solving — licensed once, no per-solve fees, no credit balance to top up
Local by default — the solver talks to
127.0.0.1; nothing is proxied through a third-party queueNo rate limit per key — throughput is bounded by your machine, not a vendor's plan tier
Fast — image captchas return in well under a second; typical reCAPTCHA v2 solves land in 30–45s
Using an existing 2captcha or Anti-Captcha integration?
CapSkip exposes the familiar in.php / res.php endpoints, so it works as a 2captcha API alternative — point your existing client at 127.0.0.1:8080 and keep your code. See the migration notes. This MCP server is the equivalent for AI agents rather than scripts.
Tools
Tool | Purpose | Required arguments |
| Check whether the CapSkip desktop app is running and reachable | none |
| Read the text out of a distorted-text captcha image |
|
| Solve reCAPTCHA v2 or v3, including invisible and Enterprise |
|
| Solve a Cloudflare Turnstile widget or challenge page |
|
| Solve a GeeTest v3 slide-puzzle captcha |
|
| Solve an ALTCHA proof-of-work challenge |
|
| Solve a Capy Puzzle captcha |
|
| Solve a CaptchaFox challenge |
|
| Solve a Friendly Captcha proof-of-work challenge |
|
There is no
min_scoreparameter oncapskip_solve_recaptcha. reCAPTCHA v3 scores are assigned by Google from signals no solver has access to — local or cloud, none can raise a score after the fact. Amin_scoreoption would promise control that does not exist, so it is deliberately left out. Passing it anyway is rejected as an unrecognized key, not silently ignored.
Full parameter tables and worked examples: API Reference.
Guide | Description |
Every captcha type — how to recognize it, what to read off the page, what call to make, where the answer goes | |
Full setup: CapSkip app, client config, first solve | |
Every tool, parameter, and return shape | |
Connection errors, timeouts, rejected tokens |
Browser automation: Playwright, Puppeteer and Selenium
This server solves the captcha and hands back a token; your agent's existing browser tooling does the driving. The pattern is the same whichever you use:
Read the sitekey from the page (
data-sitekey, or the widget's config object).Call the matching
capskip_solve_*tool with that sitekey and the page URL.Write the token into the response field and submit.
// The agent does this via its browser tool after capskip_solve_recaptcha returns
document.querySelector('#g-recaptcha-response').value = TOKEN;For non-agent scripts, use the language SDKs directly — see the Playwright, Puppeteer and Selenium guides.
Configuration
Variable | Default | Meaning |
|
| Any string when key validation is off |
|
| CapSkip host |
|
| API port from CapSkip settings |
|
| Default |
|
| Default |
ALTCHA and Capy use CAPSKIP_DEFAULT_TIMEOUT, not the reCAPTCHA one — neither
is a browser solve. ALTCHA is CPU proof-of-work measured in milliseconds; a Capy
solve is one HTTP fetch plus pixel math, held back to roughly two seconds because
Capy refuses answers that arrive faster than a human could have produced them.
Friendly Captcha is proof-of-work too but takes the longer timeout: the service sets the difficulty per request, and v2 always solves in a browser.
| CAPSKIP_POLLING_INTERVAL | 5 | Max seconds between polls |
CLI flags override environment variables, which override the defaults:
capskip-mcp --api-key <key> --host <host> --port <port> --timeout <seconds> \
--recaptcha-timeout <seconds> --polling-interval <seconds>An invalid value (non-numeric port, port outside 1–65535, a negative or out-of-range timeout, an unknown flag) fails at startup naming the offending flag or variable, rather than surfacing later as a confusing solve failure.
What you get back
Every solve tool returns a human-readable text block and structuredContent matching its declared output schema:
{
"captchaId": "12345",
"code": "03AGdBq26f...",
"solveSeconds": 11.8
}capskip_solve_turnstileaddsuserAgent. Submit the token with this exact User-Agent — Cloudflare rejects a token replayed under a different one.capskip_solve_geetestaddschallenge,validateandseccode, to post back exactly as the site's own front-end would;codekeeps the raw JSON string CapSkip returns.capskip_solve_altchaaddstoken, the base64 payload to submit verbatim in the site's form field namedaltcha, andnumber, the counter that solved the challenge;codeholds the same string astoken.capskip_solve_capyaddscaptchakey,challengekeyandanswer— not a token. All three go into the target form, in fields namedcapy_captchakey,capy_challengekeyandcapy_answer;codekeeps the raw answer as JSON. Theansweris the drag path the widget would have recorded, so submit it verbatim, and promptly: the challenge key is single-use and short-lived.capskip_solve_captchafoxaddstoken, for the form field namedcf-captcha-response, anduserAgentwhen the solve reported one — the UA the browser minted the token under, not any you sent. Submit under that one.capskip_solve_friendly_captchaaddstoken. The field it goes into differs by version:frc-captcha-solutionon v1,frc-captcha-responseon v2. The text reply names the right one for the version that was solved.
Long solves emit MCP progress notifications, so a 45-second reCAPTCHA does not trip your client's tool-call timeout.
Errors
Every tool call returns isError: true with readable text on failure — never a stack trace or a bare protocol error — so the model can read the message and correct course.
Cause | Message |
CapSkip unreachable |
|
Another process holds the port |
|
Wrong API key |
|
|
|
|
|
|
|
|
|
|
|
Solve exceeded |
|
Unknown or misspelled parameter | Rejected before the call reaches CapSkip, naming the key, e.g. |
See Troubleshooting for fixes.
FAQ
Can AI agents solve captchas?
Not on their own — a model cannot produce a valid reCAPTCHA or Turnstile token. It needs a solver. This MCP server connects your agent to CapSkip so it can request a real token and continue the task.
Does this work with Claude, Cursor and VS Code?
Yes. It is a standard MCP stdio server, so it works with any Model Context Protocol client, including Claude Desktop, Claude Code, Cursor and VS Code. Config examples for each are in examples/.
Is there a free captcha solver here?
The MCP server is MIT-licensed and free. It requires the CapSkip desktop app, which is licensed once and then solves without per-solve charges — unlike cloud APIs that bill per captcha.
Which captchas can it solve?
reCAPTCHA v2 (checkbox and invisible), reCAPTCHA v3, reCAPTCHA Enterprise, Cloudflare Turnstile, GeeTest v3, ALTCHA, Capy Puzzle, CaptchaFox, Friendly Captcha, and image/text captchas. hCaptcha and FunCaptcha/Arkose are not supported.
Why does my reCAPTCHA v3 token get a low score?
Google assigns v3 scores from signals such as IP reputation and browsing history. A solver returns a valid token, but cannot raise the score. If a site enforces a high threshold, solve from a cleaner IP — a proxy is supported on every solve tool except the image one (for ALTCHA it applies only to the challenge fetch, and for Capy only to the puzzle-image fetch).
Does it need my captcha to be on a public page?
Yes for widget captchas — CapSkip loads the page URL you pass. Image captchas need only the image, which can be a local file.
Can I use it with Playwright or Puppeteer?
Yes. The agent drives the browser; this server supplies the token. See the browser automation section.
Requirements
Node.js 18 or newer
The CapSkip desktop app installed and running (download)
Links
Captcha demo pages — live reCAPTCHA, Turnstile, GeeTest and image widgets
License
MIT — see LICENSE.
Available Tools
9 toolscapskip_solve_altchaSolve ALTCHAA
Solve an ALTCHA proof-of-work challenge. ALTCHA is not a recognition captcha — there is nothing to read; the client brute-forces a number that satisfies a challenge, so a solve is deterministic and takes milliseconds. Give it either challenge_url (the endpoint the fetches from, which CapSkip will fetch) or challenge_json (the challenge document itself). Returns a token to put in the page's form field named altcha, verbatim. IMPORTANT: challenges expire quickly — some sites inside two minutes — so read the challenge immediately before calling and submit the token promptly. An expired challenge is rejected with a bare "verification failed" that looks exactly like a wrong answer.
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | Full URL of the page the captcha appears on, including scheme. | |
| proxy | No | Fetch the challenge through this proxy. Used ONLY for the challenge_url fetch — an inline challenge_json never touches the network. | |
| timeout | No | Seconds to wait before giving up. Maximum 600. | |
| challenge_url | No | The endpoint the ALTCHA widget fetches its challenge from, e.g. 'https://site.com/altcha/challenge'. CapSkip fetches it for you. Pass this or challenge_json. | |
| challenge_json | No | The challenge document itself, as a JSON string, when you already have it. Solved locally with no network request. Pass this or challenge_url. |
Output Schema
| Name | Required | Description |
|---|---|---|
| code | Yes | |
| token | Yes | |
| number | No | |
| captchaId | Yes | |
| solveSeconds | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description adds substantial behavior beyond the annotations: challenges expire quickly, an expired challenge produces a bare 'verification failed' indistinguishable from a wrong answer, and challenge_url causes CapSkip to fetch the endpoint while challenge_json is solved locally. It also characterizes the solve as deterministic and fast, which is exactly the operational context an agent needs.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The definition is front-loaded with the core purpose, then states the two input modes, the return-value placement, and a targeted expiration warning. Every sentence earns its place; the important operational warning is placed clearly and does not bloat the text.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Combined with a 100%-covered input schema and an output schema, the description covers the full workflow: what to read, which inputs to supply, what to do with the token, and the critical failure mode. Nothing needed to invoke the tool correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema already documents all five parameters with descriptions, and the description mostly restates the challenge_url/challenge_json choice that the schema already explains. It adds workflow-level color (fetch versus local solve) but no new constraints or formats beyond the schema, so the schema-coverage baseline of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with 'Solve an ALTCHA proof-of-work challenge', naming a specific verb, resource, and task. It differentiates the tool from recognition-based captcha solvers by explicitly stating there is nothing to read and that the solve is deterministic, so an agent can distinguish it from sibling captcha tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives clear when-to-use context: for an ALTCHA challenge, pass either challenge_url or challenge_json, with the trade-off that challenge_json is solved locally with no network request. It does not explicitly name sibling tools or state when not to use this tool, but the proof-of-work vs recognition contrast plus the two alternative inputs make the intended use clear.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
capskip_solve_captchafoxSolve CaptchaFoxA
Solve a CaptchaFox challenge. CaptchaFox scores the browser itself rather than asking the visitor to read anything, so most solves draw no puzzle at all. Returns a token to put in the form field named 'cf-captcha-response', verbatim — it is verified server side against the session that produced it, so any edit invalidates it. IMPORTANT: the url must be the page the widget actually runs on. CaptchaFox checks it against the domains the key is registered for and refuses a mismatch permanently, not intermittently — so a sitekey that fails immediately and consistently usually means the wrong page URL, not a bad key. The result also carries the User-Agent the token was minted under; submit under that one, not your own.
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | Full URL of the page the captcha appears on, including scheme. | |
| proxy | No | Solve from this network. CaptchaFox scores the address a widget runs on as well as the browser, so repeated solves from one address drift toward interactive challenges and then refusals. | |
| sitekey | Yes | The public key the widget renders with, conventionally prefixed 'sk_'. Read it from data-sitekey on the widget container, from the captchafox.render(...) call, or — if the page builds the widget at runtime — from the path segment after /captcha/ in the request to api.captchafox.com. | |
| timeout | No | Seconds to wait before giving up. Maximum 600. | |
| api_server | No | The widget entry point the target page loads. Defaults to 'https://cdn.captchafox.com/'. Send the MAM package path instead when the page loads that build — it returns a MAM_ prefixed token, and the two are not interchangeable. |
Output Schema
| Name | Required | Description |
|---|---|---|
| code | Yes | |
| token | Yes | |
| captchaId | Yes | |
| userAgent | No | |
| solveSeconds | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description discloses key behaviors beyond annotations: token is verified server-side and any edit invalidates it, URL mismatches cause permanent refusals (not intermittent), and the result carries a User-Agent that must be used for submission. These are critical operational traits not present in the schema or annotations, enriching the agent's understanding of side effects and constraints.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single paragraph but well-structured with an 'IMPORTANT' callout for the URL requirement. Every sentence contributes actionable information—challenge behavior, token handling, and troubleshooting. It is slightly long but not padded; the density justifies the length.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a captcha-solving tool with two required parameters, an output schema, and a nested proxy object, the description covers the essential operational details: token usage, server-side verification, URL accuracy, and User-Agent alignment. It does not detail proxy behavior beyond the schema, but given that the schema already documents proxy semantics, the description is complete enough for an agent to solve correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% with rich per-parameter descriptions, so the baseline is 3. The description adds extra context for the 'url' parameter (must be the actual page, mismatch is permanent) and introduces the User-Agent alignment that affects how 'proxy' and 'url' are used. This goes beyond schema, providing meaningful semantic guidance.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: 'Solve a CaptchaFox challenge.' It further specifies the nature of the challenge (browser scoring, often no puzzle) and the output format (token for 'cf-captcha-response'). This clearly distinguishes it from other captcha-solving siblings by naming the provider and the token's placement.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives explicit operational guidance: the URL must be the exact widget page, the User-Agent must match the one used at solve time, and errors are often due to URL mismatch rather than key failure. It does not explicitly compare with alternatives (e.g., 'use ReCaptcha tool for other providers'), but the tool name and context make the target provider obvious, so the guidance is sufficient for selection.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
capskip_solve_capySolve Capy PuzzleA
Solve a Capy Puzzle captcha — a slide puzzle where a piece is dragged into a hole cut out of a photograph. IMPORTANT: unlike every other captcha tool here, the answer is NOT a single token. It is three values — captchakey, challengekey and answer — which all go into the target form, in fields named 'capy_captchakey', 'capy_challengekey' and 'capy_answer'. Submit the answer verbatim: it is the drag path the widget would have recorded, so trimming or re-encoding it invalidates the solve. The challenge key is single-use and short-lived, so submit promptly rather than caching the three values. A solve takes about two seconds — Capy refuses answers that arrive faster than a human could have produced them, so CapSkip holds the result back deliberately.
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | Full URL of the page the captcha appears on, including scheme. | |
| proxy | No | Draw the puzzle image through this proxy. That fetch is the only request a Capy solve makes. | |
| sitekey | Yes | The site's public Capy key, conventionally prefixed 'PUZZLE_'. Read it from 'capy_captchakey' in the page source, or from the widget script URL: <script src='https://jp.api.capy.me/puzzle/get_js/?k=PUZZLE_XXXX'>. | |
| timeout | No | Seconds to wait before giving up. Maximum 600. | |
| api_server | No | The root of the Capy API the key lives behind, taken from the widget script URL. Defaults to 'https://jp.api.capy.me'. Note that 'api.capy.me' no longer resolves, although several solver services still document it. |
Output Schema
| Name | Required | Description |
|---|---|---|
| code | Yes | |
| answer | No | |
| respKey | No | |
| captchaId | Yes | |
| captchakey | No | |
| challengekey | No | |
| solveSeconds | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description goes well beyond the annotations, disclosing that the answer must be submitted verbatim, that trimming or re-encoding invalidates it, that the challenge key is single-use and short-lived, and that CapSkip deliberately holds the result for about two seconds to avoid fast-solve rejection. This is rich behavioral context that annotations alone do not provide.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is appropriately sized for an unusually complex captcha solver. It front-loads the core purpose, then explains the non-standard answer formatituation, verbatim requirement, and timing behavior in compact sentences. Each sentence carries necessary information without filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the 100% schema coverage, existing output schema, and detailed annotations, the description covers all the unique behavioral requirements an agent must know to invoke and consume the result correctly. Nothing essential is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3. The tool description does not add parameter-level meaning beyond the schema, but it does not need to because every parameter already has a detailed schema description. No contradiction or missing compensation is present.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific verb and resource: 'Solve a Capy Puzzle captcha' and explains what kind of puzzle it is. It also distinguishes the tool from siblings by explicitly stating that, unlike every other captcha tool here, the answer is not a single token.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear context for when to use this tool: when dealing with a Capy Puzzle captcha, and contrasts it with 'every other captcha tool here' by explaining the three-value answer. It does not explicitly list when not to use it or mention sibling tools by name, but the unique answer format and field names provide strong usage direction.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
capskip_solve_friendly_captchaSolve Friendly CaptchaA
Solve a Friendly Captcha proof-of-work challenge. Nothing is shown on screen — the widget does arithmetic instead of asking the visitor to do anything. IMPORTANT: two entirely different protocols ship under this one name, sharing a brand and a sitekey namespace and nothing else, and a sitekey does not tell you which a site uses. Send version (v1 or v2), or send module_script and let CapSkip read the version off the build the site loads. Solving the wrong one returns a well-formed token the site rejects, with nothing to indicate the version was the problem. The form field also differs by version: v1 uses 'frc-captcha-solution', v2 uses 'frc-captcha-response'. Solve time is not a constant — the service sets the difficulty per request — so budget for that rather than assuming a fixed duration.
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | Full URL of the page the captcha appears on, including scheme. | |
| proxy | No | Solve from this network. Worth setting sooner here than almost anywhere else: the service raises the difficulty for addresses it has already seen a lot of. | |
| sitekey | Yes | The data-sitekey attribute of the widget element — the one carrying class='frc-captcha'. | |
| timeout | No | Seconds to wait before giving up. Maximum 600. | |
| version | No | The protocol version the site uses. Two entirely different protocols ship under this one name and a sitekey does not say which, so send this or module_script. Defaults to v1 when neither is given. | |
| api_server | No | The data-residency tenant the sitekey belongs to: 'global' (the default), 'eu', or a full URL. Both tenants mint a token for the same sitekey, so the wrong one is only caught by the site's own check. | |
| module_script | No | The src of the widget script tag carrying type='module'. The most reliable version signal there is, because it is the build the site actually loads: 'widget.module.min.js' means v1, 'site.min.js' v2. | |
| nomodule_script | No | The src of the widget script tag carrying 'nomodule'. Read for the same reason. |
Output Schema
| Name | Required | Description |
|---|---|---|
| code | Yes | |
| token | Yes | |
| captchaId | Yes | |
| solveSeconds | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only declare readOnly=false, openWorld=true, and destructive=false, so the description carries the burden of behavioral disclosure. It adds substantial context: nothing is shown on screen, two unrelated protocols share the name and sitekey namespace, a wrong solution returns a well-formed but rejected token with no diagnostic, the form field differs by version, and solve time is not constant. This goes far beyond the structured annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is information-dense with no filler. It front-loads the core purpose, then adds the critical protocol warning, version-selection guidance, field-name differences, and latency caveat in a logical order. Every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the 8-parameter complexity, nested proxy object, enums, and output schema, the description plus schema together cover protocol selection, failure modes, version-specific form fields, and variable solve time. There is no critical missing context for an agent to invoke this tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% with rich parameter descriptions, so a baseline of 3 applies. The description adds extra operational meaning beyond the schema by explaining how to choose between version and module_script, why sitekey alone is not a reliable version signal, and which form field names correspond to each version.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description begins with a specific verb, resource, and type: 'Solve a Friendly Captcha proof-of-work challenge.' It also clarifies the unusual v1/v2 protocol split, which distinguishes this tool from the sibling captcha solvers without relying on the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explicitly tells the agent to send version or module_script, warns that solving the wrong protocol produces a token the site rejects, and notes that difficulty varies per request. It does not name sibling alternatives for exclusion, but the captcha-family context and protocol-specific guidance are strong enough to guide correct invocation.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
capskip_solve_geetestSolve GeeTest v3A
Solve a GeeTest v3 slide-puzzle captcha. Returns geetest_challenge, geetest_validate, and geetest_seccode to post back exactly as the site's own front-end would. IMPORTANT: the challenge value is single-use and expires in roughly a minute, so fetch gt and challenge from the page immediately before calling. A stale challenge is the most common failure.
| Name | Required | Description | Default |
|---|---|---|---|
| gt | Yes | The gt value. Static per site, so it can be reused. | |
| url | Yes | Full URL of the page the captcha appears on, including scheme. | |
| proxy | No | Solve through this proxy so the answer is produced from its IP. | |
| timeout | No | Seconds to wait before giving up. Maximum 600. | |
| challenge | Yes | The challenge value. Single-use and expires in about a minute — fetch a fresh one immediately before calling this. | |
| api_server | No | A non-default GeeTest API domain, e.g. 'api-na.geetest.com'. |
Output Schema
| Name | Required | Description |
|---|---|---|
| code | Yes | |
| seccode | No | |
| validate | No | |
| captchaId | Yes | |
| challenge | No | |
| solveSeconds | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No contradiction with annotations (readOnlyHint=false, destructiveHint=false). Description adds valuable behavioral context beyond annotations, such as single-use challenge expiry, the need for immediate fetching, and that results should be posted back like the site's own front-end. This enhances transparency about failure modes and usage constraints.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Description is three sentences, front-loaded with the core action and return values. It includes an important usage warning without unnecessary elaboration, making every sentence purposeful.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the six parameters, output schema, and annotations, the description covers the key operational context: what the tool solves, what returns, and the critical freshness requirement. It does not discuss proxy usage or timeout, but those are well-documented in the schema, so the description is sufficiently complete for typical use.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, with detailed descriptions for all 6 parameters. The description reinforces the challenge freshness warning and gt reuse but adds minimal new meaning beyond the schema. Baseline of 3 is appropriate since the schema does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Description clearly states 'Solve a GeeTest v3 slide-puzzle captcha', naming the specific captcha type and resource. It also specifies return values (geetest_challenge, geetest_validate, geetest_seccode), distinguishing it from sibling captcha-solving tools like reCAPTCHA and Turnstile.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides explicit operational guidance: 'fetch gt and challenge from the page immediately before calling' and warns that a stale challenge is the most common failure. It does not explicitly contrast with alternative siblings, but the captcha-type-specific name and description give clear context for when this tool is appropriate.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
capskip_solve_image_captchaSolve an image captchaA
Read the text out of a distorted-text captcha image. Returns the recognized text, which you type into the page's captcha field. Proxies are not supported for image captchas.
| Name | Required | Description | Default |
|---|---|---|---|
| image | Yes | The captcha image: a local file path, an http(s) URL, a data: URI, or a raw base64 string. | |
| timeout | No | Seconds to wait before giving up. Maximum 600. |
Output Schema
| Name | Required | Description |
|---|---|---|
| code | Yes | |
| captchaId | Yes | |
| solveSeconds | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With annotations already indicating readOnlyHint=false and destructiveHint=false, the description adds useful behavioral context: it returns recognized text and notes that proxies are unsupported. This goes beyond the structured annotations without contradicting them, providing operational expectations for the agent.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences, front-loaded with the primary action and immediate use case. The second sentence provides a necessary limitation (proxy support) without any filler or redundancy. Every word contributes to understanding the tool's function and constraints.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the simple parameter set (2 params, 1 required), full schema coverage, presence of an output schema, and annotations, the description covers the essential behavior, return value, and a key limitation. It could explicitly contrast with sibling captcha-solving tools, but the tool name and sibling list sufficiently fill that gap, making the description complete for the tool's scope.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema covers 100% of parameters with detailed descriptions for 'image' and 'timeout'. The description only reiterates that the image is a captcha, which is already evident from the schema. With full schema coverage, the baseline of 3 is appropriate; the description adds no significant semantic value beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action: 'Read the text out of a distorted-text captcha image' with a specific resource and outcome. It distinguishes this tool from siblings by explicitly focusing on 'image captcha' versus Turnstile, reCAPTCHA, and Geetest, making the purpose unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies when to use it: when a distorted-text image captcha needs solving, and explains the result should be typed into the page's captcha field. It also provides a clear exclusion: 'Proxies are not supported for image captchas.' It does not explicitly name alternatives, but the sibling tool names and the 'image captcha' qualifier offer sufficient differentiation.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
capskip_solve_recaptchaSolve reCAPTCHAA
Solve a Google reCAPTCHA v2 or v3 widget, including invisible and Enterprise variants. Returns a token to place in the page's "g-recaptcha-response" field before submitting the form. Read the sitekey from the page first — a guessed sitekey fails. Note that reCAPTCHA v3 returns a score assigned by Google; no solver can raise it.
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | Full URL of the page the captcha appears on, including scheme. | |
| proxy | No | Solve through this proxy so the token is issued against its IP. | |
| action | No | v3 only. The action passed to grecaptcha.execute(), e.g. 'login'. | |
| data_s | No | v2 only. The data-s value, used by Google's own services. Rarely needed — CapSkip rejects it on a v3 submit. | |
| sitekey | Yes | The site key, from the widget's data-sitekey attribute or the grecaptcha config. | |
| timeout | No | Seconds to wait before giving up. Maximum 600. | |
| version | No | Which reCAPTCHA generation the page uses. Defaults to v2. | |
| invisible | No | v2 only. True when the widget renders with size=invisible. | |
| enterprise | No | True for reCAPTCHA Enterprise. Works with both v2 and v3. |
Output Schema
| Name | Required | Description |
|---|---|---|
| code | Yes | |
| captchaId | Yes | |
| solveSeconds | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already mark the tool as non-readonly, non-destructive, and open-world. The description adds useful behavioral context: the token is placed in g-recaptcha-response before form submission, guessed sitekeys fail, and v3 scores are immutable. It does not describe potential external service requests or rate limits, but the openWorldHint and schema largely cover the interaction model.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is three sentences and about 70 words, with no filler or repetition of schema details. It front-loads the main purpose, then packs in the key practical warnings. Every sentence contributes new information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity (9 parameters, nested proxy object, output schema) and the exhaustive input schema, the description covers all essential operational context: what it solves, what it returns, a critical prerequisite, and a fundamental limitation. It does not need to explain return values because an output schema exists, and it complements the rich schema and annotations well.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so a baseline of 3 applies. The description adds meaningful parameter guidance beyond the schema by stressing that the sitekey must be read from the page (not guessed) and by noting v3's score cannot be raised, which affects expectations for the version/action parameters. This extra context justifies a score above baseline.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'Solve a Google reCAPTCHA v2 or v3 widget, including invisible and Enterprise variants.' It clearly distinguishes itself from sibling tools (Turnstile, image captcha, Geetest) by naming Google reCAPTCHA and describing what it produces. The returned token's placement in g-recaptcha-response further clarifies its purpose.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides a concrete prerequisite: 'Read the sitekey from the page first — a guessed sitekey fails.' It also warns about the v3 score limitation, helping agents know not to expect a higher score. It lacks an explicit 'use this instead of X' statement, but the sibling tool names and captcha-specific language make the context clear enough.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
capskip_solve_turnstileSolve Cloudflare TurnstileA
Solve a Cloudflare Turnstile widget or interstitial challenge page. Returns a token for the page's "cf-turnstile-response" field. IMPORTANT: submit the token using the returned userAgent — Cloudflare rejects a token replayed under a different User-Agent. For an interstitial challenge page, also pass cdata and pagedata read from the page.
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | Full URL of the page the captcha appears on, including scheme. | |
| cdata | No | The cData value. Interstitial challenge pages only, not ordinary widgets. | |
| proxy | No | Solve through this proxy so the token is issued against its IP. | |
| action | No | The action from data-action or turnstile.render(). | |
| sitekey | Yes | The Turnstile site key, from the widget's data-sitekey attribute. | |
| timeout | No | Seconds to wait before giving up. Maximum 600. | |
| pagedata | No | The chlPageData value. Interstitial challenge pages only. |
Output Schema
| Name | Required | Description |
|---|---|---|
| code | Yes | |
| captchaId | Yes | |
| userAgent | No | |
| solveSeconds | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
It discloses a critical behavioral constraint: the returned token must be submitted with the same userAgent or Cloudflare rejects it. It also clarifies when cdata/pagedata are required. These go beyond the annotations (readOnlyHint=false, openWorldHint=true) and are essential for correct use.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tightly written sentences immediately convey purpose, output, and a high-priority caveat. No filler or repetition of schema fields; every sentence adds value.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 7-parameter solver with a nested proxy object and an output schema, the description covers the essential behavioral nuances (widget vs interstitial, token-userAgent binding). The schema documents parameter details, and the output schema removes the need to explain return values, making the description complete for this complexity.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema already describes all seven parameters with 100% coverage, giving a solid baseline. The description adds extra meaning by tying cdata/pagedata to the interstitial challenge page scenario and explaining the userAgent token requirement, though it does not detail proxy or timeout semantics beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb ('Solve') and resource ('Cloudflare Turnstile widget or interstitial challenge page'), and explicitly states the output (a token for cf-turnstile-response). This clearly distinguishes it from sibling captcha-solving tools (image, recaptcha, geetest).
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides clear usage context by distinguishing between widget and interstitial challenge page forms, and instructs passing cdata/pagedata for interstitial pages. It does not explicitly name alternatives or exclusions, but the sibling names make the intended use obvious.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
capskip_statusCheck CapSkip statusARead-only
Check whether the CapSkip desktop app is running and reachable. Call this first when a solve fails unexpectedly, to tell "CapSkip is not running" apart from "the sitekey was wrong". Takes no arguments.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
| host | Yes | |
| port | Yes | |
| detail | Yes | |
| latencyMs | No | |
| reachable | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and openWorldHint=true, so the description adds value by specifying what is being checked (running and reachable) and the diagnostic purpose. It does not introduce any surprising behavior, and there is no contradiction with annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences, front-loaded with the core purpose, and includes essential usage guidance. Every sentence contributes value with no redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple status check with no parameters, an output schema, and read-only annotations, the description fully covers what the tool does, when to use it, and that it takes no arguments. No gaps remain.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes no parameters, and the description explicitly states 'Takes no arguments,' which matches the empty schema. Baseline for 0 params is 4, and since no additional semantic detail is needed, this is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool checks whether the CapSkip desktop app is running and reachable, using a specific verb and resource. It also distinguishes itself from sibling solving tools by framing this as a diagnostic pre-check.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicit usage guidance is provided: 'Call this first when a solve fails unexpectedly' and it distinguishes between 'CapSkip is not running' and 'the sitekey was wrong'. This gives clear when-to-use context without needing to mention alternatives.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
3 tool updates
v1.2.0- Added
capskip_solve_captchafox - Added
capskip_solve_capy - Added
capskip_solve_friendly_captcha
1 tool update
v1.1.0- Added
capskip_solve_altcha
5 tool updates
v1.0.0- First observed
capskip_solve_geetest - First observed
capskip_solve_image_captcha - First observed
capskip_solve_recaptcha - First observed
capskip_solve_turnstile - First observed
capskip_status
TDQS
Scored across 9 tools
Each tool targets a distinct captcha provider or protocol, so there is little overlap in purpose. The only potential confusion is among the proof-of-work tools (Friendly, ALTCHA) since both are invisible challenges, but the descriptions clearly distinguish them by protocol and form field.
All tools follow a consistent capskip_solve_<provider> pattern, with capskip_status as the lone exception, which is still clearly named and fits the prefix convention. The verb_noun structure is uniform and predictable.
Nine tools is well-scoped for a captcha-solving server: eight provider-specific solvers plus one status check. Each tool covers a distinct captcha type, and the count is appropriate for the domain.
The tool surface covers all major captcha providers (reCAPTCHA, Turnstile, hCaptcha is missing but not listed, GeeTest, etc.) and includes a status check for operational failures. Minor gaps exist—such as no hCaptcha solver or no batch/retry support—but the core lifecycle of solving and submitting is complete.
Maintenance
Related MCP Connectors
Official MCP server for Agentwork — delegate tasks to AI agents with human-in-the-loop
Hosted MCP server connecting claude.ai, ChatGPT and other AI apps to your own computer
An MCP server that integrates with Discord to provide AI-powered features.
9 remote MCP servers on Cloudflare Workers for AI agents. Free tier + Pro API keys.
Related MCP Servers
AlicenseAqualityDmaintenanceA local MCP server that lets AI agents bypass bot detection, geo-restrictions, and JavaScript rendering challenges when scraping the web, backed by ScraperAPI's services285MIT- AlicenseNot gradedqualityAmaintenanceMCP server that lets AI agents use your real browser as an API, accessing any website with your login state, no keys or scrapers needed.1MIT
- AlicenseCqualityDmaintenanceThis MCP server enables automated captcha solving for Cloudflare challenges and provides a comprehensive JSVMP offline replay pipeline for signature recovery and analysis.453MIT
- AlicenseNot gradedqualityDmaintenanceAn MCP server that provides an undetectable Playwright browser to bypass Cloudflare and other bot detection systems, enabling AI agents to navigate and scrape web pages without being blocked.9 npmMIT