Skip to main content
Glama

uvc-ptz-camera-mcp

python license

An MCP server for USB (UVC) pan/tilt/zoom cameras: aim them, nudge them, sweep them smoothly, zoom, and look through them — with every move confirmed by comparing the picture before and after.

No account, no cloud, no vendor SDK. One device, one USB cable.

camera_status       the camera, its axes and ranges, and whether it is real
aim                 point an axis at an absolute value, confirmed from the picture
nudge               move relative to where the device says it is
sweep               a smooth timed move, streamed at 15 Hz
zoom                set zoom as a multiplier (1x .. the camera's maximum)
recentre            every axis back to its default
look                one frame, returned as an image
aim_learn           remember "this direction is the desk"
aim_list            what has been recorded for this camera
go_to               point at a recorded direction
run_shot            a multi-step move, verified after every waypoint
mark_view           store the current picture under a label
check_view          has the picture changed since that label?

Why this exists

This class of camera reports positions it never moved to. Measured on the reference device (a DJI Osmo Pocket 4P in webcam mode, over its standard UVC controls):

  • a command to pan to 180 read back 180 while the video proved the camera had not moved at all;

  • a whole 32-second run read back 0 on every sample while the frame demonstrably changed;

  • pan 120 read back 119, pan 200 read back 199.

An agent acting on those values reports moves that never happened. So this server treats every device-reported value as a hint — returned, labelled, never trusted — and decides moved by comparing frames.

The threshold is not a guess either. It was calibrated on labelled hardware frames: every "same view" pair scored ≤ 0.104 (a 96-second quiet baseline, and two commands the hardware ignored), every "view changed" pair scored ≥ 0.624 (a pan, a tilt, a 12× zoom, two aim steps). The shipped threshold is the midpoint, 0.364 — a 6× separation — and the test suite asserts it still separates them.

Related MCP server: OBSBOT Camera MCP Server

What keeps the agent honest

  • moved comes from the picture. device_reported is returned beside it, labelled a hint.

  • A move that produced no change is an error, not a quiet success — after a bounded retry, the tool fails naming what was requested and what was observed.

  • Already being at the target is success without movement: moved: false, no error.

  • Blocking work never runs on the event loop (COM calls, ffmpeg captures), pinned by a test that records which thread the camera was touched on.

  • Observation is explicit: nothing takes a picture except a tool you called.

  • A refused write aborts a shot rather than continuing to drive a camera that stopped listening.

Install

uvx uvc-ptz-camera-mcp            # run without installing
pip install uvc-ptz-camera-mcp    # or install it

For a real camera on Windows you also need the DirectShow extras:

pip install "uvc-ptz-camera-mcp[dshow]"

Then point your host at it — --print-config emits the snippet with the right interpreter:

uvc-ptz-mcp --print-config
uvc-ptz-mcp --list-devices        # what to pass to --device
uvc-ptz-mcp --backend simulator   # try it with no camera attached

Where to put it

Claude Desktop, Cursor, VS Code and friends all take the same shape — this is the whole configuration, because there is nothing to authenticate:

{
  "mcpServers": {
    "ptz-camera": {
      "command": "uvx",
      "args": ["uvc-ptz-camera-mcp"]
    }
  }
}

Add "--device", "Osmo" (any substring of the name --list-devices prints) when the machine has more than one camera, and "--backend", "simulator" to work with none. Both can also be set in the environment instead of the args — UVC_PTZ_DEVICE, UVC_PTZ_BACKEND and UVC_PTZ_STATE_DIR — which is what the Claude Desktop bundle uses, and which survives a host restart without editing its config again. A flag wins over the environment, and an empty value means "not set" (a host that leaves an option blank writes "", not nothing). Claude Desktop also accepts the .mcpb bundle attached to each release as a one-click install.

Modes, and how you can tell which one you are in

Backend

What it is

auto (default)

Prefer a real camera; fall back to the simulator, saying why

dshow

Windows DirectShow: the real device path

simulator

No hardware at all: a model calibrated from the measurements below

Every result carries "simulated": true or false, and camera_status reports the reason when the simulator is standing in. Nothing quietly pretends to be hardware: a caller must never believe it is driving a camera when it is driving a model.

The simulator is calibrated from the reference device, not invented: ~0.4 s from write to first motion, a pan completing in about a second, a 4× zoom in about 2.5 s, silently dropped writes, and a read-back that echoes a request the hardware never applied. It renders frames by cropping a wide panorama, so a simulated pan genuinely changes the pixels — a simulator with a static picture would let verification pass vacuously.

Honest limitations

  • The real-device path is Windows-only (DirectShow). A Linux backend would use v4l2; the interface is in place, the implementation is not written.

  • No exposure, focus or white balance. The reference camera exposes no Processing Unit at all — only pan/tilt/roll/zoom.

  • The vendor Extension Unit is unreachable on Windows, so features that live there (on the reference camera, its built-in subject tracking) cannot be driven from software. Measured: IKsControl is refused on the device filter, IKsTopologyInfo lists three nodes and no vendor node, and CreateNodeInstance fails on all of them. Linux can reach it through uvcvideo's UVCIOC_CTRL_QUERY; that is a separate backend.

  • A moving subject looks like a moving camera. Frame comparison cannot tell the two apart; the metric is calibrated so that a static scene behaves, and check_view is there for judging a view rather than a move.

  • One unusable device must not hide the others. A registered-but-unavailable virtual camera raised when its DirectShow moniker was bound, which aborted a listing that should have returned two working cameras. Enumeration now skips what it cannot load and says so, and the same rule applies inside the backend when it looks for the camera you named.

  • Measured latency sets the ceiling: ~0.4 s from command to motion, so control loops run at a couple of hertz, not tens. That is the hardware's floor, not the software's.

  • The DirectShow path is tested but dormant. The reference device was returned partway through development, so the hardware path is exercised up to its interface and typed correctly against it, while the simulator carries the test load. Treat the first run against a real camera as the true acceptance test.

Development

pip install -e ".[dshow]" numpy pytest pytest-asyncio ruff
python -m pytest -q                     # 39 tests, three layers
python -m ruff check . && python -m ruff format --check .

The suite is layered deliberately: the tool surface in-process against the simulator (mapping, validation, honest failure), pure unit tests for the compiler and the metric, and a real stdio handshake as a subprocess — including one run with a device name that cannot exist, because a server that dies at startup is invisible to every host and directory that lists it.

Reference device

Built against a DJI Osmo Pocket 4P in webcam mode: VID_2CA3 / PID_0023, exposing pan −38…215°, tilt −33…105°, roll ±35°, zoom 100…1200 (1×–12×) as standard UVC camera controls. Any UVC PTZ camera exposes the same surface; the ranges are read from the device at startup rather than assumed, and nothing vendor-specific is hard-coded.

Not affiliated with, endorsed by, or supported by DJI.

Available Tools

13 tools
aimA

Point one axis at an absolute value, and return what the picture confirms.

axis is one of pan, tilt, roll, zoom. Gimbal axes are in degrees; zoom is the device's own ratio x100, so 300 means 3x. The value is clamped to the range the camera advertises and the clamped value is what gets sent.

This moves the camera and waits for it: allow roughly a second for a gimbal axis and three for a zoom. moved is decided by comparing frames, never by the device's report -- the report is returned as device_reported and labelled a hint, because this hardware was measured claiming positions it had not reached. If the picture shows no change and the camera was not already at the target, the call fails rather than reporting success.

ParametersJSON Schema
NameRequiredDescriptionDefault
toYes
axisYes

TDQS

A4.2/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does so richly: latency expectations (about one second for gimbal, three for zoom), value clamping to the camera's advertised range, synchronous waiting, and the crucial disclosure that success is decided by frame comparison rather than the device's unreliable self-report.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core action, then layered with units, timing, and verification semantics. Dense but each sentence carries distinct operational value; the parenthetical about clamping is well placed. Slightly heavy for a two-parameter tool, but nothing is filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description usefully names the return signals (`moved` from frame comparison, `device_reported` as an untrusted hint) and spells out the failure condition when the picture shows no change. A brief note on the full response shape or the `axis` value constraints at the schema level would make it complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0% and the schema is bare (`to`, `axis` with no descriptions or enums), so the description must compensate fully — and it does. It enumerates the axis domain (pan, tilt, roll, zoom), gives units (degrees; zoom as ratio x100, so 300 = 3x), and explains that `to` is clamped to the camera's advertised range.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ("Point one axis at an absolute value") and tightens the scope with the axis domain and units. The phrase "one axis" and "absolute value" implicitly separates it from relative siblings like `nudge`, but no sibling is named explicitly, so it stops short of full differentiation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is implied rather than stated: the agent can infer this is for absolute one-axis moves, and the mention that zoom is handled via axis=zoom partially disambiguates from the `zoom` sibling. However, there is no explicit when-to-use vs `nudge`, `recentre`, `go_to`, or `zoom` guidance and no stated prerequisites.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

aim_learnA

Remember the current direction under a name, for later use.

Pan is an absolute axis whose meaning depends on where the camera is standing, so "point at the door" is unanswerable without recording it once. Call this with the camera already pointed where you want, and the label becomes usable with go_to.

ParametersJSON Schema
NameRequiredDescriptionDefault
labelYes

TDQS

A3.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden but only partially fulfills it. It usefully explains the rationale (pan is an absolute axis, so a label must be recorded once) and that the label is later consumed by go_to, but it says nothing about whether an existing label is overwritten, whether labels persist across sessions, or what the call returns.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The core action is front-loaded in the first sentence, and the second paragraph's explanation of the pan/absolute-axis problem fits in two lines. Slightly verbose in its justification, but every sentence supports the reader's decision to call it.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a single-parameter tool with no output schema or annotations, the description covers purpose, precondition, and consumer, which is largely sufficient. It remains incomplete on label lifecycle semantics (overwrite behavior, persistence) that an agent would need to call it safely.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

One parameter with 0% schema description coverage, so the description must compensate. It implies the 'label' is a reusable name for the recorded direction, but adds no constraints on format, uniqueness, or whether reusing a name replaces a prior recording.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a concrete verb+resource ('Remember the current direction under a name'), which is far more informative than the cryptic name 'aim_learn'. It also names the sibling that consumes the result ('go_to'), though it does not distinguish itself from other state-recording siblings like mark_view or check_view.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

'Call this with the camera already pointed where you want' gives an explicit precondition for invoking the tool, and 'the label becomes usable with go_to' routes the agent to the correct follow-up tool. No when-not guidance or contrast with the other camera-state siblings is provided.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

aim_listA

List the directions recorded for this camera, and its calibrated default pose.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the burden of behavioral disclosure. It does communicate that this is a read-only list operation and describes the returned content (recorded directions plus the calibrated default pose), but it omits any detail about permissions, side effects, or output format. For a simple no-parameter query tool, this is adequate but not rich.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The entire definition is a single front-loaded sentence with no filler. Every clause contributes specific information about what is listed and what the camera context is. Nothing could be removed without losing meaning.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the low complexity (no parameters, no annotations) and the absence of an output schema, the description does the essential work of naming what is returned: recorded directions and the calibrated default pose. It does not describe the output format or list ordering, but for a simple enumeration tool it is nearly complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so there is no parameter syntax to explain. The baseline for a zero-parameter tool is 4, and the description correctly adds no misleading parameter claims.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('List') and a clearly bounded resource ('the directions recorded for this camera, and its calibrated default pose'). It is easy to distinguish from sibling tools like aim or aim_learn, which likely set or learn directions rather than enumerate them. It stops short of naming an explicit sibling alternative, which keeps it from a full 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description states what the tool returns but gives no guidance about when to use it versus siblings such as camera_status, check_view, or aim_learn. There are no explicit exclusions or alternative-routing instructions. An agent must infer usage from the tool name and purpose alone.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

camera_statusA

Report the camera, its axes and ranges, and whether it is a real device.

Call this first. It says which camera was found, the range each axis accepts, where the camera believes it is pointing (a hint, not proof), what has been calibrated, and -- if no real device was found -- the reason. simulated: true in any result means nothing physical moved and no picture was ever taken.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does real work: it warns the pointing value is 'a hint, not proof', explains that a failed device returns a reason, and defines the crucial `simulated: true` semantics ('nothing physical moved and no picture was ever taken'). It does not explicitly state that the call is side-effect-free/idempotent, which is the one remaining gap for a zero-annotation tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The 'Call this first' directive is front-loaded ahead of the longer enumeration of return contents, so the most decision-relevant information comes first. The second sentence is dense but each clause corresponds to a distinct output field, so little is wasted.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

There is no output schema, so the description must enumerate what comes back, and it does: camera identity, per-axis ranges, estimated pointing, calibration state, and failure reason. Combined with the simulated-flag caveat, an agent has everything needed to interpret and trust the result.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes no parameters, so the schema has nothing to describe and there is no parameter ambiguity to resolve. The baseline of 4 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The first sentence gives a specific verb and resource ('Report the camera, its axes and ranges') plus the real-vs-simulated distinction, which no sibling (aim, go_to, run_shot, etc.) provides. An agent can tell this apart from the motion and vision siblings without opening any schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

'Call this first' is an explicit, actionable ordering directive that tells the agent when to reach for this tool relative to the other camera tools. It stops short of naming exclusions or explicitly contrasting with a specific sibling, so it lands just below the top band.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

check_viewA

Compare the current picture with one stored by mark_view.

Answers "is the camera still looking at what it was looking at?", which is the question a scene change can otherwise hide: a person walking through the shot changes the picture as much as a pan does, so treat a small difference as "still there" and a large one as "worth looking at".

ParametersJSON Schema
NameRequiredDescriptionDefault
labelYes

TDQS

A3.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It does add real behavioral value by describing the nature of the result (a difference judgment where small = still there, large = worth looking at) and warns about a false-positive source (a person walking through the shot). However, it never states the concrete return form (boolean, numeric diff, image) or any permission/rate considerations, which is a notable gap given the absence of an output schema.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The core comparison is front-loaded in the first sentence, which is good. The second sentence is long and somewhat discursive, running the motivating question, the scene-change caveat, and an interpretation heuristic together; the interpretation guidance earns its place but could be tightened.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with no annotations, no output schema, and an entirely undocumented parameter, the description does partially compensate by explaining how to read the comparison qualitatively. But it leaves the concrete output shape, the meaning of `label`, and any failure modes unspecified, so an agent still lacks what it needs to invoke and consume the result confidently.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% for the single required `label` parameter, so the description must compensate and largely does not. It refers to 'one stored by mark_view', which only weakly hints that `label` identifies a saved view and must match a prior `mark_view` call; the label's format and constraints remain undocumented anywhere.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (compare) and resources (the current picture vs. one stored by `mark_view`), and explicitly names the sibling tool responsible for the stored state. An agent can immediately distinguish this from `mark_view` and the other camera-control siblings without opening any schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It conveys the operating context clearly: use it to answer whether the camera is still pointed at what it was, and it implies a prerequisite (a view previously stored via `mark_view`). It offers no explicit when-not or alternative tools, but the trigger condition is unambiguous.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

go_toA

Point the camera at a direction recorded earlier with aim_learn.

Each axis is moved and confirmed from the picture, the same as aim. A label is resolved to the angles recorded for it, which may drift if the camera has been physically moved since; the picture check still proves that each axis moved, not that the label is right.

ParametersJSON Schema
NameRequiredDescriptionDefault
labelYes

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and largely meets it: it discloses that each axis is moved and visually confirmed (same as `aim`), that labels resolve to interpolated angles, and that a physically moved camera can cause drift. The key caveat that the picture check proves axis motion, not label correctness, is genuinely useful. It omits error/failure handling for unknown labels.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Purpose is front-loaded in the first sentence and the caveat paragraph is short and information-dense. Both paragraphs earn their place, though the second reads slightly densely.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No annotations and no output schema, so the description must cover behavior on its own; it covers purpose, mechanism, and the main correctness caveat. The remaining gap is what happens on an unresolved label or on visual-check failure.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, so the description must compensate, and it does: 'A label is resolved to the angles recorded for it' explains that `label` is a previously learned identifier rather than an arbitrary string. It does not specify naming format or case sensitivity, keeping it below a 5.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Point the camera at a direction recorded earlier with aim_learn') and explicitly ties the mechanism to sibling `aim` while sourcing the target from `aim_learn`. An agent can distinguish this from `aim` (explicit direction) and `aim_learn` (recording) without opening any schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The phrase 'recorded earlier with `aim_learn`' makes the precondition and the alternative tool clear: this tool consumes labels that aim_learn produced. It stops short of stating failure behavior when a label does not exist, so it is clear context without explicit exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

lookA

Capture one frame of what the camera currently sees, as a PNG image.

This observes the camera's view. Call it when the user asks to see the picture, not on your own initiative, and be aware that a still only tells you what is in front of the lens -- not where the camera is, and not whether a move happened.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does real work: 'observes' establishes it is non-mutating, and it discloses output limitations ('a still only tells you what is in front of the lens -- not where the camera is, and not whether a move happened'). It omits anything about permissions, latency, or failure modes, but the interpretive caveat is genuinely useful.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The core action and output format are front-loaded in the first sentence, followed by usage and limitation notes. It is efficient overall, though the second sentence runs long and the caveat clause is slightly more verbose than needed.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema and no annotations, the description supplies what a caller needs: what is returned (a PNG frame of the current view) and what that frame cannot tell the agent. Nothing critical about invocation is missing, though return-handling details (e.g., where the image lands) are left implicit.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so the baseline is 4. The description correctly implies no inputs are needed to trigger a capture, and there is nothing for it to disambiguate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource: 'Capture one frame of what the camera currently sees, as a PNG image.' It is unambiguous against siblings like go_to or sweep, though it never names which sibling to use instead for related tasks (e.g., check_view), so it stops short of explicit sibling differentiation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives an explicit trigger ('Call it when the user asks to see the picture') and an explicit exclusion ('not on your own initiative'). That is strong when/when-not guidance, but it does not point to any alternative tool for cases the agent might confuse with this one.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

mark_viewA

Store the current picture under a label, so it can be compared later.

Useful before a move whose result you will judge later, or to establish what "the wide shot" looks like. Marks live in memory for this session only.

ParametersJSON Schema
NameRequiredDescriptionDefault
labelYes

TDQS

A4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden, and it does disclose one important trait: marks are ephemeral, 'live in memory for this session only'. It says nothing about overwrite behavior for a reused label, any cap on the number of marks, or whether the snapshot captures the framing/pose versus the rendered image.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short sentences, front-loaded with the core action, followed by motivation and lifetime note. Every sentence contributes and none restate the name or the schema.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a one-parameter, no-output-schema tool the description covers purpose, typical triggers, and persistence scope, which is most of what an agent needs. It omits the label-collision/overwrite rule and how marks are later retrieved, which are the remaining gaps given no annotations exist to fill them.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% for the single 'label' parameter, so the description must compensate. It reveals the label's role as the key used for later comparison, but gives no format, uniqueness, or length guidance; the agent still has to guess what makes a valid label.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description gives a concrete verb and object ('Store the current picture under a label') plus the reason it matters ('so it can be compared later'). An agent can distinguish this snapshot-taking tool from siblings like check_view, run_shot, or aim without inspecting any schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It states concrete contexts where the tool is appropriate: before a move whose result you will judge later, or to establish a known reference framing like 'the wide shot'. It does not name a sibling alternative (e.g. check_view for retrieval) or any when-not condition, so it stops short of a full routing rule.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

nudgeB

Move one axis relative to where the device says it is now.

Relative moves inherit the device's reported position, which is a hint: after a series of nudges the accumulated position may drift from belief. The picture check still tells you whether this call moved the camera; it cannot tell you it is at an absolute angle.

ParametersJSON Schema
NameRequiredDescriptionDefault
byYes
axisYes

TDQS

B3.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does disclose genuine behavioral traits: relative positioning, that accumulated position can drift from the device's belief, and that the 'picture check' only verifies this call moved the camera rather than the absolute angle. These are meaningful caveats an agent can act on, though valid axis values and units remain undisclosed.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The core purpose is front-loaded in the first sentence, but the rest is somewhat rambling and includes a trailing fragment about the 'picture check' that is never defined. Every sentence is not fully earning its place in this form.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No annotations and no output schema mean the description must stand alone, and it covers the key behavioral caveat well. However, for a movement tool it omits parameter meaning and axis enumeration, leaving gaps an agent needs to invoke it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0% – neither 'axis' (string) nor 'by' (integer) has any schema description, and the description only hints at 'one axis' without listing valid values or explaining what the integer 'by' signifies or its units. It does not compensate for the near-total absence of parameter documentation.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Move one axis') and clarifies it is relative to the device's reported position, which distinguishes it conceptually from absolute-positioning siblings like go_to. It does not name any sibling explicitly, so the differentiation is implied rather than stated.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies when this is appropriate (relative movement, and its cumulative use 'after a series of nudges') and warns about drift, but it never states when to prefer this over siblings like aim, go_to, or sweep. Usage context is inferable but no explicit alternatives or exclusions are given.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

recentreA

Return every axis to the camera's own default: centre, level, unzoomed.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden, and it does disclose the concrete state change: all axes returned to the camera's own defaults (centre, level, unzoomed). What it does not address is whether the camera's default is per-camera or session-defined, or whether any prior aim/mark state is lost. That leaves a meaningful behavioral gap for a mutating camera tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with a colon-delimited list of the resulting states. No padding, no restatement of the tool name, every clause carries information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-parameter, zero-annotation tool with no output schema, the description supplies the essential effect and target state. The only omission is what the call returns or whether it can fail, which is a minor gap at this complexity level.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so there are no parameter semantics to explain. Baseline 4 applies; the description correctly needs to say nothing further about inputs.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific action (reset) and enumerates the resulting state on three axes: centred, level, unzoomed. An agent can tell this is a global camera-reset rather than a per-axis move. It does not explicitly contrast itself with siblings like zoom or aim, but the purpose is unmistakable.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance, no prerequisite or trigger condition, and no mention of alternatives such as zoom or nudge that also adjust the view. The reset semantics imply a recovery/intent, but the agent must infer that on its own.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

run_shotA

Execute a multi-step camera move, verifying after every waypoint.

Each step is an object: {"axis": "pan", "to": 60, "seconds": 2.5, "ease": "in_out"}. Steps run in sequence and every axis that moves in a step is written on each tick, so a step naming two axes produces a genuine diagonal rather than a staircase. Add {"hold": 1.0} to pause after a step.

The whole plan is validated against the camera's advertised ranges before anything moves, so an impossible shot fails without touching the hardware. Afterwards the report gives the picture difference at each waypoint, which is how you tell a shot that happened from one that did not. A long shot takes its full duration; do not set a host timeout shorter.

ParametersJSON Schema
NameRequiredDescriptionDefault
stepsYes

TDQS

A4.5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden and does so richly: it discloses that the plan is validated against camera ranges before motion, that every axis in a step is written on each tick (avoiding staircase artifacts), that holds can be added, that a report shows picture difference at each waypoint, and that long shots need sufficient host timeout. This is exactly the behavioral context an agent needs.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is front-loaded with the core action, then efficiently covers step structure, sequencing, validation, return report, and timing caveat. Every sentence provides a distinct, useful detail for correct invocation without padding.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Despite no output schema and no annotations, the description explains what the tool does, how steps are structured, how validation and reporting work, and a critical timing constraint. For a complex multi-step camera tool with an opaque nested schema, this is complete enough for an agent to call it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% and the single parameter is an untyped array of objects, so the description must compensate. It supplies the step object shape with example fields (axis, to, seconds, ease) and explains hold semantics, sequences, and multi-axis behavior. It does not enumerate all axis or ease values, but it adds substantial meaning beyond the empty schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The first sentence states a specific verb and resource: 'Execute a multi-step camera move, verifying after every waypoint.' This clearly distinguishes it from single-move siblings such as go_to, aim, or nudge by emphasizing the multi-step, verified workflow.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies when to use this tool (for multi-step camera moves) and includes a timing constraint ('do not set a host timeout shorter'), but it does not explicitly name alternatives or state when not to use it versus siblings like go_to or sweep. Usage guidance is present but inferential.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

sweepA

Move an axis smoothly to a value over a chosen duration.

Use this for any move that should read as deliberate: a slow pan across a scene, a gradual push in. The axis is driven by streaming absolute targets at 15 Hz rather than one jump, which is how this hardware was measured moving smoothly. ease is one of linear, in_out, out, in; in_out starts and stops gently, which is what makes a move look intentional. Long sweeps take as long as seconds, so expect to wait.

ParametersJSON Schema
NameRequiredDescriptionDefault
toYes
axisYes
easeNoin_out
secondsNo

TDQS

A3.9/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does real work: it discloses the control mechanism ('streaming absolute targets at 15 Hz rather than one jump, which is how this hardware was measured moving smoothly') and the blocking cost ('Long sweeps take as long as seconds, so expect to wait'). It omits permissions, axis range limits, and whether a sweep can be interrupted or superseded.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the action, then usage, mechanism, ease semantics, and a timing warning, in that order. Every sentence contributes something an agent could not get from the schema. It runs slightly long, with the 'which is what makes a move look intentional' clause repeating the earlier 'read as deliberate' framing.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema and no annotations exist, so the description is the only source of behavior, and it covers mechanism, timing, and enum values well. The gap is the undocumented required inputs (axis identity and target value), which an agent needs before it can call this correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate, and it only does so for the optional parameters: it supplies the full ease enum (linear, in_out, out, in) with semantics for in_out, and explains seconds as both duration and wait time. The two REQUIRED parameters, axis and to, get no semantics at all — no valid axis names, no value units or range.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

Specific verb and resource: 'Move an axis smoothly to a value over a chosen duration.' An agent immediately knows this is a timed, interpolated motion primitive rather than an instant jump. It does not explicitly contrast itself with siblings like go_to or nudge, which would have made the distinction airtight.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicit positive guidance: 'Use this for any move that should read as deliberate: a slow pan across a scene, a gradual push in.' That is a clear selection criterion with concrete examples. There is no negative guidance naming an alternative (e.g. 'for instant repositioning use go_to'), so an agent must infer the counterpart from sibling names.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

zoomA

Set optical zoom as a multiplier: 1.0 is wide, up to the camera's maximum.

Convenience over the raw zoom axis, which counts in hundredths. The requested ratio is clamped to what the camera supports and the clamped value is returned.

ParametersJSON Schema
NameRequiredDescriptionDefault
ratioYes

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does disclose a real behavioral trait: the requested ratio is clamped to hardware limits and the clamped value is returned, which tells the caller the value may not be honored as requested. It stops short of saying whether zoom persists across other operations, what permissions or connection state are needed, or how it interacts with concurrent camera tools.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short sentences, front-loaded with the action and the value semantics before the contrasting/raw-axis note. No filler, no restatement of the tool name or the schema title.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a one-parameter, no-output-schema tool, the description covers value semantics, hardware limits, clamping, and the return of the resolved value – everything needed to call it correctly. The remaining gap is environmental context (does it require an active camera session, does zoom persist), which is minor.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0% for the single 'ratio' parameter, so the description must compensate and does: it defines the unit (multiplier, not hundredths), the anchor point (1.0 is wide), the upper bound (camera maximum), and the clamping behavior with returned resolved value. That is more meaning than the bare name 'ratio' conveys by a wide margin.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description gives a specific verb+resource ('Set optical zoom') and immediately defines the value domain: a multiplier where 1.0 is wide, up to the camera's max. It also positions itself as a convenience layer over the raw zoom axis. It does not distinguish itself from any of the listed siblings (aim, nudge, sweep), but the intent is unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is only implied: the reader learns this is the 'convenience' alternative to the raw hundredths-based zoom axis, but that raw tool isn't among the siblings and no when-to-use/when-not guidance is given relative to aim, look, or the other camera tools. An agent can infer the appropriate context but is not told explicitly.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 13 tool updatesv0.1.1
    • First observedaim
    • First observedaim_learn
    • First observedaim_list
    • First observedcamera_status
    • First observedcheck_view
    • First observedgo_to
    • First observedlook
    • First observedmark_view
    • First observednudge
    • First observedrecentre
    • First observedrun_shot
    • First observedsweep
    • First observedzoom

TDQS

A3.8/5.0

Scored across 13 tools

Disambiguation4/5

Most tools are clearly distinct by resource or action: status, capture, view comparison, labeling, and axis movement. The main overlap is among the movement tools (aim, nudge, sweep, zoom, go_to, run_shot), which all move the camera but differ by absolute/relative/smooth/sequence semantics; descriptions help, though an agent could still pause over which move primitive to choose.

Naming Consistency3/5

All names use snake_case, but the structural conventions are mixed: single verbs (aim, nudge, sweep, zoom, recentre, look), verb_noun forms (mark_view, check_view, run_shot), noun_noun (camera_status), and less predictable orderings (aim_list, aim_learn). The set remains readable, but there is no single predictable pattern.

Tool Count5/5

13 tools is within the well-scoped range and each tool has a plausible role: movement, scripting, view memory, calibration labels, status, and capture. The set does not feel bloated, and the extra tools support verification and repeatability rather than duplicating obvious functionality.

Completeness4/5

Core PTZ workflows are covered: status, capture, absolute/relative/smooth movement, zoom convenience, recentre, position labels, view comparison, and multi-step shots. Minor gaps remain, such as no explicit abort/stop for long-running moves and no delete/list operations for stored marks or aim labels, but these are workable within session-only memory.

Maintenance

ActivityNo data
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers