Skip to main content
Glama
evalstate
by evalstate

mcp-ウェブカメラ

Web カメラを使用して、ライブ画像を Claude Desktop (またはその他の MCP クライアント) に送信します。

Claude がウェブカメラからフレームを取得したり、スクリーンショットの撮影を開始したりできるようにするための"capture"および"screenshot"ツールを提供します。

current view from the webcamも提供します。

インストール

NPM パッケージは@llmindset/mcp-webcamです。

お使いのプラットフォーム用の最新バージョンのNodeJSをインストールし、 claude_desktop_config.jsonファイルのmcpServersセクションに次のコードを追加します。

    "webcam": {
      "command": "npx",
      "args": [
        "-y",
        "@llmindset/mcp-webcam"
      ]
    }

Claude Desktop 0.78 以上を使用している限り、Windows と MacOS の両方で動作します。

組み込み Express サーバーのポートを設定するために、単一の引数を取ります。

デフォルトのポートは3333です (Inspector で使用する場合の競合を回避するため)。

Related MCP server: Webcam MCP

使用法

Claude Desktopを起動し、 http://localhost:3333に接続します。すると、Claudeに「 get the latest picture from my webcam 」と頼んだり、「 Claude, take a look at what I'm holdingとかwhat colour top am i wearing?と頼んだりできます。現在の画像を「フリーズ」すると、ライブキャプチャではなく、Claudeに返されます。

スクリーンショットをリクエストできます。リクエストが届いたらブラウザを開いて、キャプチャ領域を指定してください。スクリーンショットはClaudeが操作しやすいように自動的にサイズ調整されます(4K画面をお持ちの場合に便利です)。このボタンは、プラットフォーム固有のスクリーンショットUXをテストするためのもので、Claudeからのリクエストに備えるためのものです。Safariではこの機能は動作しません。これは、ユーザーが操作を開始する必要があるためです。

MCPサンプリング

「何を持っているの?」ボタンを押すと、画像とWhat is the User holding?質問を含むサンプリング要求がクライアントに送信されます。

[!TIP] Claude Desktopは現在サンプリングをサポートしていません。マルチモーダルサンプリングリクエストに対応できるクライアントが必要な場合は、 https://github.com/evalstate/fast-agent/をお試しください。

その他の注意事項

本当にそれです。

この MCP サーバーは、MCP サーバー上でユーザー インターフェイスを公開し、ライブ リソースを Claude Desktop に返す方法を示すために構築されました。

このプロジェクトは、ローカルでインタラクティブな MCP サーバーを構築する場合に役立つ可能性があります。

テストとセットアップに協力してくれたhttps://github.com/tadasantに感謝します。

LLM / MCP チャット アプリケーションでのファイルとリソースの処理方法と、これを行う理由の詳細については、 https://llmindset.co.uk/posts/2025/01/resouce-handling-mcpの記事をお読みください。

サードパーティのMCPサービス

Available Tools

2 tools
captureA
Read-only

Gets the latest picture from the webcam. You can use this if the human asks questions about their immediate environment, if you want to see the human or to examine an object they may be referring to or showing you.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4.6/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations provide readOnlyHint=true and openWorldHint=true, indicating safe, non-destructive operation with potential for varied outcomes. The description adds valuable context by specifying it captures from 'the webcam' and returns 'the latest picture,' clarifying the source and immediacy of the data, which goes beyond what annotations alone convey.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is front-loaded with the core purpose in the first sentence, followed by usage guidelines in a clear, efficient manner. Every sentence adds value without redundancy, making it appropriately sized and well-structured for quick comprehension.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's low complexity (0 parameters, no output schema), the description is complete enough for effective use. It covers purpose, usage guidelines, and behavioral context. The absence of an output schema is mitigated by the description's clarity on what is returned ('the latest picture'), though more detail on output format could enhance completeness.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema has 0 parameters with 100% coverage, so no parameter documentation is needed. The description appropriately does not discuss parameters, maintaining focus on tool functionality. A baseline of 4 is applied since there are no parameters to document.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the specific action ('Gets the latest picture from the webcam') and resource ('webcam'), distinguishing it from the sibling tool 'screenshot' which likely captures screen content rather than camera input. The verb 'Gets' is precise and the resource is unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly provides when-to-use guidance with concrete examples: 'if the human asks questions about their immediate environment,' 'if you want to see the human,' or 'to examine an object they may be referring to or showing you.' This gives clear context for selecting this tool over alternatives like 'screenshot'.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

screenshotB
Read-only

Gets a screenshot of the current screen or window

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already indicate readOnlyHint=true and openWorldHint=true, so the agent knows this is a safe, non-destructive operation with open-world assumptions. The description adds minimal behavioral context beyond this, such as specifying it captures the 'current screen or window', but doesn't detail aspects like format, size, or potential limitations. No contradiction with annotations exists.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, clear sentence that efficiently conveys the core functionality without any wasted words. It is front-loaded with the essential information, making it easy for an agent to parse and understand quickly.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's simplicity (0 parameters, no output schema) and rich annotations (readOnlyHint, openWorldHint), the description is adequate but minimal. It covers the basic purpose but lacks details on output format or behavioral nuances that could aid the agent, such as whether it returns an image file or data. It meets minimum viability but has gaps in completeness.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema has 0 parameters with 100% coverage, so no parameter documentation is needed. The description appropriately doesn't discuss parameters, focusing instead on the tool's action. A baseline of 4 is applied since it avoids redundancy and adds value by explaining the tool's purpose without unnecessary details.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose with a specific verb ('Gets') and resource ('screenshot of the current screen or window'), making it immediately understandable. However, it doesn't explicitly differentiate from the sibling tool 'capture', which might have overlapping functionality, preventing a perfect score.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no guidance on when to use this tool versus the sibling tool 'capture', nor does it mention any prerequisites, context, or exclusions. It merely states what the tool does without offering usage instructions, leaving the agent to infer when it's appropriate.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections. Dates show when Glama detected each change.

  1. 2 tool updatesv1.0.0
    • First observedcapture
    • First observedscreenshot

TDQS

B3.4/5.0
Disambiguation2/5

The tools have overlapping purposes—both capture visual data from the user's environment—with 'capture' targeting the webcam and 'screenshot' targeting the screen, but descriptions could lead to confusion as 'capture' mentions examining objects the human shows, which might overlap with screen content. The boundaries are somewhat unclear, especially for agents interpreting use cases.

Naming Consistency4/5

Tool names follow a consistent verb-based pattern ('capture' and 'screenshot'), both being single words describing the action. There are no deviations in style or casing, making them readable and predictable, though 'screenshot' is more specific than 'capture' in terms of naming convention.

Tool Count2/5

With only 2 tools, the server feels under-scoped for a webcam domain, as it lacks operations like video capture, settings adjustment, or multi-camera support. This minimal set may limit agent functionality, making it borderline too few for comprehensive visual input handling.

Completeness2/5

The tool surface is significantly incomplete for a webcam server; it covers basic image capture from webcam and screen but misses essential operations such as starting/stopping video, configuring camera settings, or handling multiple inputs. This creates gaps that could lead to agent failures in more complex visual tasks.

Maintenance

ActivityInactive
ResponsivenessSyncing

Resources

Unclaimed servers have limited discoverability.

Looking for Admin?

If you are the server author, to access and configure the admin panel.

Related MCP Connectors

Related MCP Servers

  • -
    license
    Not graded
    quality
    Not graded
    maintenance
    Provides real-time screen capture streaming in base64 format via WebSocket connections, enabling Claude to view and analyze user screens through natural language requests.
    1
    -
  • A
    license
    Not graded
    quality
    C
    maintenance
    An MCP server that provides LLM agents with direct access to webcam hardware for capturing high-resolution photos and recording video sequences. It enables autonomous agents to monitor environments and interact with the physical world through standard Model Context Protocol tools.
    3
    MIT
  • A
    license
    Not graded
    quality
    D
    maintenance
    An MCP server that enables interaction with local camera devices to capture and process images. It allows LLMs to access video devices with configurable settings such as resolution, orientation, and image format.
    1
    MIT
  • F
    license
    A
    quality
    D
    maintenance
    Enables LLMs to capture screenshots and screen recordings through MCP with chunked session-based transfers for reliable image consumption. Supports multi-monitor selection, timeline capture, and compatibility with both vision and non-vision language models.
    11
    1
    -

Latest Blog Posts

MCP directory API

We provide all the information about MCP servers via our MCP API.

curl -X GET 'https://glama.ai/api/mcp/v1/servers/evalstate/mcp-webcam'

If you have feedback or need assistance with the MCP directory API, please join our Discord server