eufy-cam-mcp
View-only access to an allowlisted set of local Eufy cameras through MediaMTX.
List known cameras (
eufy_cam1–eufy_cam4) with live ready state.Check one camera's ready state, tracks, and error counters.
Capture a single JPEG snapshot from a camera and return the file path.
Get recording info: segment count, disk use, oldest/newest recording.
List newest motion clips for a camera (
limit1–50).Extract one JPEG frame at
tseconds from a named motion clip.All camera tools reject names outside the allowlist and path traversal;
clip_framealso rejects non-bare.mp4clip names.No PTZ, talkback, cloud API, or recording control.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@eufy-cam-mcpshow me which eufy cameras are online and grab a snapshot from eufy_cam1"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
eufy-camera-mcp
A Python FastMCP server that gives an AI agent view-only access to an allowlisted set of local Eufy cameras through MediaMTX: live status, a single JPEG snapshot, and recorded motion clips. It has no PTZ, no talkback, and no cloud API.
Why it exists
An agent that can see a camera should not be able to reach cameras it was not given, call out to the internet, or move the camera. This server keeps the surface small: four known camera names, localhost MediaMTX only, and one frame at a time.
Related MCP server: Frigate MCP Server
Tools
Tool | What it returns |
| Each known camera with its ready state from MediaMTX |
| Ready state, tracks, and error counters for one camera |
| Path to one JPEG frame grabbed with |
| Segment count, disk use, and oldest and newest recording |
| Newest motion recordings for one camera, |
| Path to one JPEG frame at |
Every per-camera tool rejects names outside the allowlist, including path traversal attempts such as ../../outside, before any MediaMTX call or ffmpeg process. clip_frame also rejects clip names that are not a bare *.mp4 file name.
Requirements
Python 3.12 or newer and uv
ffmpegonPATHA MediaMTX instance on the same host that publishes your cameras as
eufy_cam1toeufy_cam4A recorder that writes
<name>_*.mp4segments, if you want the recording tools (not part of this repo)
Install and run
git clone https://github.com/thefiredev-cloud/eufy-camera-mcp.git
cd eufy-camera-mcp
uv sync --frozen
uv run eufy-cam-mcpThis starts the MCP server on stdio. Register it with your MCP client as a stdio command, for example:
{
"mcpServers": {
"eufy-cam": {
"command": "uv",
"args": ["--directory", "/path/to/eufy-camera-mcp", "run", "eufy-cam-mcp"]
}
}
}Configuration
Settings are constants at the top of src/eufy_cam_mcp/server.py:
Constant | Default |
|
|
|
|
|
|
|
|
|
|
The ~/meshvault/... folders are the defaults used on the author's machines. Edit the constants to match your layout; there are no environment variable overrides yet.
Check it without cameras
The boundary check calls camera_status and snapshot_camera with bad camera names and confirms that no MediaMTX call, network connection, ffmpeg process, or snapshot folder was created:
uv run python scripts/check-camera-boundaries.py
uv run pytest -vCI runs Ruff, the boundary check, and the test suite on pushes and pull requests to main.
Scope
In scope: allowlisted names, MediaMTX ready state from /v3/paths/list, one-frame JPEG capture, and listing and sampling local motion clips.
Out of scope: PTZ, talkback, Eufy cloud APIs, recording, and retention. Recording belongs to a separate recorder service; this server only reports files already on disk.
Status
Working, version 0.1.0. Not published to PyPI and not deployed as a hosted service.
License
MIT. See LICENSE.
Available Tools
6 toolscamera_statusC
Ready state, tracks, and error counters for one camera.
| Name | Required | Description | Default |
|---|---|---|---|
| name | No | eufy_cam1 |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full behavioral burden. It implies a read-only status read but says nothing about freshness (live query vs cached), permission requirements, or behavior for an unknown/nonexistent camera name.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single short sentence with no padding, and the payload contents are front-loaded. It is terse to the point of being fragmentary, but nothing is wasted.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return structure need not be explained, and the description usefully hints at the payload contents. However, for a tool with no annotations and an undocumented parameter, it leaves open how the camera is selected and whether the status is live or cached.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% for the single 'name' parameter, and the description only refers vaguely to 'one camera' without explaining that the camera is identified by the name argument or noting the eufy_cam1 default. It does not compensate for the coverage gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names the resource ('one camera') and enumerates the payload ('ready state, tracks, and error counters'), which lets an agent distinguish this from snapshot_camera or list_cameras. It is a noun phrase rather than a verb+resource statement, so it is clear but slightly under-specified about the action being a retrieval.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to reach for this tool versus siblings like list_cameras or snapshot_camera, and no conditions or exclusions are stated. The agent must infer usage from the payload description alone.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
clip_frameB
Extract one JPEG frame at t seconds into a motion clip. Returns the file path.
| Name | Required | Description | Default |
|---|---|---|---|
| t | No | ||
| clip | No | ||
| name | No | eufy_cam1 |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the load, and it does disclose the output form ('Returns the file path') and that extraction is read-only in effect. However, it omits permission requirements, behavior when `clip` is empty, and whether the frame is written to disk or streamed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences, front-loaded with the action and the key parameter, with the return value noted briefly. Nothing wasted.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so explaining return values is optional, yet the description does so succinctly. The gap is that odd parameter defaults (t=5, clip='', name='eufy_cam1') and the meaning of an unset clip are never addressed for a tool with 0% schema description coverage.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must compensate for 3 undocumented params. It only clarifies `t` (seconds); `clip` and `name` (default 'eufy_cam1') are left entirely unexplained, including what an empty-string clip default means.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: 'Extract one JPEG frame at `t` seconds into a motion clip.' The 'motion clip' scoping implicitly separates it from snapshot_camera (live capture), but no sibling is named explicitly. Clear purpose, weak sibling differentiation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance on when to use this versus snapshot_camera or motion_clips. The phrase 'into a motion clip' hints that a clip reference is needed, but no prerequisites, exclusions, or alternative-routing statement is given.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_camerasB
List known Eufy cameras with live ready state from MediaMTX.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the behavioral burden. It usefully discloses the data source (MediaMTX) and the state nature (live ready state), which hints at a read-only enumeration, but it does not explicitly state read-only semantics, live-vs-cached behavior, or whether results can be stale.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single tight sentence with no wasted words, and the resource is front-loaded. It is efficient but so terse that it leaves sibling differentiation unaddressed.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be explained, and with zero parameters the schema burden is minimal. Still, for an enumeration tool sitting among five related siblings, the description says nothing about scope, freshness, or when this list is the right call, leaving a clear gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so there is nothing for the description to clarify beyond what the empty schema already conveys. Baseline 4 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('List') and resource ('known Eufy cameras') and adds the returned state qualifier ('live ready state from MediaMTX'). It does not, however, distinguish itself from the sibling camera_status, which plausibly covers overlapping ground.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no when-to-use or when-not-to-use guidance, and no mention of the sibling tools such as camera_status or snapshot_camera that an agent must choose between. Usage is only inferable from the name.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
motion_clipsB
Newest motion-event recordings for one camera (newest first, capped).
| Name | Required | Description | Default |
|---|---|---|---|
| name | No | eufy_cam1 | |
| limit | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It usefully discloses ordering (newest first) and that results are capped, but does not say it is a read-only operation, how large the cap is, or whether any auth/scoping applies.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single dense fragment with no waste and the scope constraints front-loaded. It is effective, though its telegraphic style skips the explicit verb that would make it unmistakable.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need no explanation, and the core scope is stated. Still, the cap size, parameter meanings, and the read-only nature are absent, leaving real gaps for a two-parameter listing tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must compensate. 'One camera' gestures at the name parameter and 'capped' gestures at limit, but neither parameter is named, defaulted, or bounded, so format and valid ranges remain unclear.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The fragment names the resource (motion-event recordings) and the scope (one camera, newest first, capped), so an agent can tell what it returns. It does not, however, distinguish itself from siblings like recordings_info or clip_frame, leaving overlap ambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no when-to-use or when-not-to-use guidance and no mention of the alternative siblings (recordings_info, clip_frame, snapshot_camera). The agent must infer selection criteria entirely from the names.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
recordings_infoB
Segment count, disk use, and oldest/newest recording (retention visibility).
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden; it discloses the shape of what is returned (aggregate counts and timestamps) which signals this is a non-destructive read. It says nothing about permissions, cost, or whether the scan is expensive, so the behavioral picture remains thin.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single tight phrase with the most informative element (segment count) front-loaded and zero filler. It is under-specified as a sentence rather than overly verbose.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter read tool whose return fields are already declared in the output schema, describing the returned values is somewhat redundant but harmless, and nothing an agent needs to invoke it is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so the baseline is 4; there are no parameter semantics to document and the description correctly does not invent any.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Names the concrete data it reports (segment count, disk use, oldest/newest recording), so an agent knows this is a recording-storage summary tool rather than a clip lister like motion_clips. However, it uses a bare noun phrase with no verb, so the operation itself ('reports'/'summarizes') is only implied.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The parenthetical '(retention visibility)' hints at the use case, but there is no explicit statement of when to call this versus siblings such as motion_clips or clip_frame, and no exclusions or prerequisites.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
snapshot_cameraC
Capture a single JPEG frame from a camera. Returns the file path.
| Name | Required | Description | Default |
|---|---|---|---|
| name | No | eufy_cam1 |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden, but it only notes the output is a file path. It says nothing about permissions, whether the file is persisted to disk (a side effect), capture latency, or whether a camera must be online/selected.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences with the purpose front-loaded and no wasted words. Size is appropriate for a simple single-parameter tool.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so explaining the return value is not required, but the description omits any handling of the one input parameter and offers no usage context against the many sibling tools. Adequate for a trivial tool but with clear gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
There is one parameter ('name') with a default of 'eufy_cam1' and 0% schema description coverage, yet the description never mentions how to target a camera or what values 'name' accepts. The description does not compensate for the schema gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Capture') and resource ('a single JPEG frame from a camera'), which is clear on its own. However, it does not differentiate itself from the sibling 'clip_frame', which sounds like a closely related capture operation, so an agent has to guess which one applies.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance on when to use this tool versus 'clip_frame', 'motion_clips', or 'recordings_info'. The agent is given no conditions or exclusions to route between the several capture/clip siblings.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
6 tool updates
v0.1.0- First observed
camera_status - First observed
clip_frame - First observed
list_cameras - First observed
motion_clips - First observed
recordings_info - First observed
snapshot_camera
TDQS
Scored across 6 tools
Tools mostly target distinct resources: list all cameras, get one camera status, snapshot live, get recordings overview, list motion clips, and extract a clip frame. However, snapshot_camera and clip_frame both return a JPEG file path, and camera_status partially overlaps list_cameras' ready state, so minor confusion is possible.
All names use snake_case, but the patterns are mixed: verbs (list_cameras, snapshot_camera, clip_frame) alternate with noun phrases (camera_status, recordings_info, motion_clips). Still readable, but not a consistent verb_noun convention.
Six tools is well within the ideal 3–15 range for a camera monitoring server, and each tool appears to serve a distinct purpose.
Core monitoring and frame-extraction workflows are covered, but there is no direct tool to retrieve or download a full motion clip (only frame extraction), and no camera control or configuration operations. These are minor gaps for the apparent monitoring-focused purpose.
Maintenance
Related MCP Connectors
Ask your security cameras anything and set up alert rules in a sentence. Built into Agent DVR.
Video, audio, and image processing for AI agents: convert, transcribe, upscale - 150+ operations.
FFmpeg as a service for AI agents: typed video editing tools, async jobs, downloadable outputs.
Zero-setup MCP gateway securely connecting AI to your tools with authentication and workflows
Related MCP Servers
- AlicenseBqualityDmaintenanceEnables PTZ camera control with gimbal positioning, snapshots, and AI visual analysis for OBSBOT and UVC cameras. Supports autonomous scanning patterns and integrates with vision-language models for real-time camera analysis.71MIT
- FlicenseAqualityAmaintenanceEnables AI assistants to interact with Frigate NVR security camera systems, supporting camera management, event detection, snapshots, recordings, and system stats via natural language.710-
- AlicenseNot gradedqualityBmaintenanceEnables LLMs to list, inspect, snapshot, and control Scrypted home-automation/NVR devices such as switches, locks, and cameras through natural language.34 npmMIT
- AlicenseAqualityBmaintenanceEnables interacting with a self-hosted UniFi Protect console through natural language, covering cameras, recorded events, smart detections, snapshots, footage export, and connected devices, with optional write access.2234 npm1MIT