Skip to main content
Glama

Site-Shot MCP server

让 Claude、Cursor 和其他 AI 代理能够看到任意网页 —— 通过 Model Context Protocol 使用 Site-Shot 截取网站截图。

真实的 Chromium 渲染 · 整页截取 · 国家代理 · 自动移除广告和 Cookie 横幅(图像更干净、视觉令牌更少)。

快速开始(Claude Desktop)

  1. 在 https://www.site-shot.com/start/ 获取 Site-Shot API 密钥。

  2. 将其添加到 Claude Desktop 配置(claude_desktop_config.json)中:

{
  "mcpServers": {
    "site-shot": {
      "command": "npx",
      "args": ["-y", "site-shot-mcp"],
      "env": { "SITESHOT_API_KEY": "YOUR_API_KEY" }
    }
  }
}
  1. 重启 Claude Desktop。让它 “对 https://news.ycombinator.com 截取整页截图”,它就会调用服务器并把图像展示给你。

在任何 MCP 客户端(Cursor、Cline、VS Code、LangChain、CrewAI)中的用法都一样 —— 只需要让客户端在环境中携带 SITESHOT_API_KEY 指向 npx -y site-shot-mcp 即可。

Related MCP server: Webpage Screenshot MCP Server

工具

capture_screenshot

截取一个网页的截图(默认按视口截取)。

参数

类型

默认值

说明

url

string(必填)

—

要截取的页面

full_page

boolean

false

截取整个可滚动页面

width / height

number

API 默认值

视口 / 设备尺寸

format

"png" | "jpeg"

png

图像格式

block_ads

boolean

true

移除广告

block_cookie_banners

boolean

true

移除 Cookie 同意弹窗

country

string

—

代理国家 / 地区,使用两位 ISO 3166-1 alpha-2 代码,例如 "DE"(自动确定 IP / 语言 / 时区 / 地理位置)

strict_country

boolean

true

如果该国没有代理则报错,而不是回退到美国

language / time_zone / geolocation

string

—

手动覆盖项

wait_ms

number

API 默认值

截取前的额外等待(SPA / 动画)

max_height

number

20000(整页)

限制截取的高度上限

以 MCP 图像形式返回截图。

“API 默认值”并不是这个包能够声明的某个数值。 width、height 和 wait_ms 只会在你传入时才被转发,因此当你不传参数时实际生效的值由 Site-Shot API 决定,并且可能在本包没有任何发版变动的情况下发生变化。1.1.0 及更早版本为 width / height 打印了 API 并不使用的像素尺寸 —— 省略这些参数去“采用默认值”的代理会得到 不同的视口,而返回的图像中没有任何信息能揭示这一点。只要尺寸很重要,就请传入显式值。

国家代码是 ISO 代码,绝不是名称。 请传 "DE",而不是 "Germany"。API 会精确匹配代码, 否则就会在不告知你的情况下通过美国代理渲染。因此,服务器会在消耗一次渲染之前拒绝完整名称。strict_country (默认开启)同样会把不可用的国家变成错误,而不是静默地截出一张美国截图 —— 传 false 即可重新启用回退。 支持的国家 →

capture_full_page

与 capture_screenshot 相同,但已启用了整页截取。

为什么要调用这个服务器,而不是用代理自带的浏览器?

如果你的代理驱动浏览器,它完全可以自己截取页面 —— 对于那些需要登录或需要逐步走完整流程的页面,那才是正确的工具。对于公开 URL,把截取工作委托给这个服务器通常是更好的工程实践:每次截取都运行相同的流水线(多次运行之间无需重新规划),可以从特定国家 / 地区截取并匹配相应的语言环境、时区(country + strict_country),返回前还会由图像分类器评分,并带有逐级升级的重试机制作为后盾,而且每次查看的成本只是一美分的一小部分,而不是一个浏览器会话加视觉令牌。完整的两侧对比、两种方向都如实说明: AI 代理 vs. 截图 API —— 谁应当截取页面。

配置

环境变量

必填

说明

SITESHOT_API_KEY

是

你的 Site-Shot API 密钥(用作 userkey)。

该服务器是现有 Site-Shot HTTP API(https://api.site-shot.com/)的薄封装 —— 没有独立的后端。

本地开发

npm install
npm run check   # syntax check
npm run smoke   # offline tests (stubbed fetch, no API key needed)
SITESHOT_API_KEY=yourkey npm start   # run the server on stdio

环境要求

Node.js ≥ 18(使用内置的 fetch)。

许可证

MIT

Available Tools

2 tools
capture_full_pageCapture full-page website screenshotA

Take a full-page (entire scrollable height) screenshot of a web page with Site-Shot and return it as an image. Convenience wrapper around capture_screenshot with full-page capture enabled.

ParametersJSON Schema
NameRequiredDescriptionDefault
urlYesThe URL of the web page to capture. A bare domain like example.com is accepted (https:// is assumed).
widthNoViewport width in pixels (default 1280).
heightNoViewport height in pixels (default 1024).
formatNoImage format. Default: png.
block_adsNoRemove ads for a cleaner screenshot. Default: true.
block_cookie_bannersNoRemove cookie-consent banners/popups. Default: true.
countryNoRender through a proxy in this country, e.g. "Germany" (auto-sets IP, language, time zone, geolocation).
languageNoOverride browser language, e.g. "de".
time_zoneNoOverride time zone, e.g. "Europe/Berlin".
geolocationNoOverride geolocation as "lat,lng".
wait_msNoMilliseconds to wait after load before capturing (for SPAs/animations).
max_heightNoCap the captured height in pixels (max 20000).

TDQS

A3.9/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It describes the action as a wrapper but does not disclose side effects, output format details, or limitations beyond what the schema parameters cover.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences with no wasted words; the core purpose is front-loaded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool has 12 parameters (all well-described in schema) and no output schema, the description is adequate as a summary but lacks details on return format and additional behavioral context.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% with detailed parameter descriptions, so the description adds little extra meaning. Baseline score of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool takes a full-page screenshot and mentions it's a convenience wrapper around capture_screenshot with full-page capture enabled, distinguishing it from the sibling tool.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage for full-page screenshots and references the sibling tool, but does not explicitly state when not to use it or provide alternative scenarios.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

capture_screenshotCapture website screenshotB

Take a screenshot of a web page with Site-Shot and return it as an image. Renders in a real Chromium browser. Supports viewport/device sizing, full-page capture, country proxies, and automatic ad & cookie-banner removal (cleaner image, fewer vision tokens).

ParametersJSON Schema
NameRequiredDescriptionDefault
urlYesThe URL of the web page to capture. A bare domain like example.com is accepted (https:// is assumed).
widthNoViewport width in pixels (default 1280).
heightNoViewport height in pixels (default 1024).
formatNoImage format. Default: png.
block_adsNoRemove ads for a cleaner screenshot. Default: true.
block_cookie_bannersNoRemove cookie-consent banners/popups. Default: true.
countryNoRender through a proxy in this country, e.g. "Germany" (auto-sets IP, language, time zone, geolocation).
languageNoOverride browser language, e.g. "de".
time_zoneNoOverride time zone, e.g. "Europe/Berlin".
geolocationNoOverride geolocation as "lat,lng".
wait_msNoMilliseconds to wait after load before capturing (for SPAs/animations).
max_heightNoCap the captured height in pixels (max 20000).
full_pageNoCapture the entire scrollable page instead of just the viewport. Default: false.

TDQS

B3.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the burden of behavioral disclosure. It states the tool renders in a real Chromium browser and automatically removes ads and cookie banners, which is helpful. However, it does not mention potential side effects, rate limits, execution time, or authentication requirements. It also does not clarify whether the screenshot is destructive or what happens to the browser instance after capture.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two sentences long, making it relatively concise. The first sentence states the core action, and the second lists major features. It avoids extraneous details but could be slightly more compact by combining the two sentences or trimming the feature list slightly.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the complexity of the tool (13 parameters, no output schema), the description provides a high-level overview of capabilities but lacks detail on return format (e.g., image type, resolution), error handling, and how features like 'full_page' work in practice. It is adequate for an experienced user but incomplete for a novice.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema covers all 13 parameters with descriptions, achieving 100% coverage. The tool description reiterates some schema concepts (viewport sizing, full-page capture, country proxies) but does not add significant new meaning beyond what the schema already provides. For example, 'country' parameter is explained in the schema; the description only mentions 'country proxies' generically. Baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states it takes a screenshot of a web page using Site-Shot and returns an image. It mentions features like viewport sizing, full-page capture, and ad removal. However, it does not explicitly distinguish itself from the sibling tool 'capture_full_page', which may cause confusion.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description lists features but provides no guidance on when to use this tool versus alternatives like 'capture_full_page'. It does not mention any prerequisites or conditions for use, nor does it explain when to use the 'full_page' parameter or when to prefer a different tool.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 2 tool updatesv1.0.1
    • Changedcapture_full_page3 fields changed
      • changedInput schema / properties / url / description
        Previous value: -"The URL of the web page to capture."New value: +"The URL of the web page to capture. A bare domain like example.com is accepted (https:// is assumed)."
      • removedInput schema / properties / url / format
        Removed value: -"uri"
      • addedInput schema / properties / url / minLength
        Added value: +1
    • Changedcapture_screenshot3 fields changed
      • changedInput schema / properties / url / description
        Previous value: -"The URL of the web page to capture."New value: +"The URL of the web page to capture. A bare domain like example.com is accepted (https:// is assumed)."
      • removedInput schema / properties / url / format
        Removed value: -"uri"
      • addedInput schema / properties / url / minLength
        Added value: +1
  2. 2 tool updatesv0.1.1
    • First observedcapture_full_page
    • First observedcapture_screenshot

TDQS

B3.2/5.0

Scored across 2 tools

Disambiguation2/5

The two tools are nearly identical; capture_full_page is explicitly a wrapper for capture_screenshot with full-page enabled. An agent would likely misuse them, as the difference is only a parameter.

Naming Consistency3/5

Both use verb_noun pattern ('capture_screenshot', 'capture_full_page'), but 'full_page' is a qualifier while 'screenshot' is the resource; inconsistent because one tool name specifies a parameter in the name itself.

Tool Count3/5

Two tools is minimal but arguably sufficient for a simple screenshot service. However, the duplication suggests one tool could have been omitted, making the surface slightly too heavy for the scope.

Completeness3/5

The set covers basic screenshot needs with features like viewport sizing, proxies, and ad removal. However, it lacks tools for specific device emulation or batch processing, which are common in screenshot services.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers