obsbot-mcp
Controls OBSBOT Tiny 2 USB cameras and OBSBOT Tail 2 network cameras via MCP, covering gimbal/PTZ, zoom, AI tracking, image, capture, and preset operations.
Discover/select Tiny 2 cameras, wake/sleep, read status, and handle multi-camera via serial
camera.Move/recenter/read gimbal pose; speed moves where supported; save/recall/update/rename/delete presets.
Zoom (UVC/vendor), aim at a pixel, zoom-to-fit a region from a snapshot.
Control AI tracking modes/speed, face focus, FOV, HDR, focus auto/manual, exposure auto/manual, white balance, image adjustments.
Capture snapshot, record MP4, open preview, stop/list sessions.
Tail 2 network control: scan/register/list devices, status/info/live status, zoom/hybrid zoom, gimbal jog/move/recenter/position/invert, focus/exposure/image/HDR/WB/audio/stream/encoder, portrait/roll bias, AI tracking/tap-to-track/focus point, presets, snapshot via SRT, export log, UDP zone/auto-zoom/custom tracking controls.
Debug raw XU probe available with
--debug.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@obsbot-mcprecenter the camera gimbal"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
obsbot-mcp
A cross-platform Model Context Protocol server that controls an OBSBOT Tiny 2 camera over its standard UVC/USB interface — pan/tilt/roll the gimbal, zoom, AI subject tracking, focus/exposure/white-balance/image controls, HDR and field-of-view, plus snapshot, preview, and recording — without any vendor SDK. It also controls the OBSBOT Tail 2 network camera over its HTTP/WS API (see OBSBOT Tail 2 below).
Install
npm install obsbot-mcpRelated MCP server: Robot MCP Server
MCP client configuration
Add a stdio server entry pointing at the installed binary (or directly at dist/index.js):
{
"mcpServers": {
"obsbot": {
"command": "obsbot-mcp"
}
}
}If you're running from a local checkout instead of an npm install, point command/args at
node and the built entry point instead:
{
"mcpServers": {
"obsbot": {
"command": "node",
"args": ["path/to/obsbot-mcp/dist/index.js"]
}
}
}Debug / diagnostics tools
By default the server advertises only the normal control surface. Pass --debug to additionally
expose the diagnostics surface — the obsbot_debug_probe tool (raw XU byte get/set/query) and the
raw 60-byte status block on obsbot_status:
{
"mcpServers": {
"obsbot": {
"command": "node",
"args": ["path/to/obsbot-mcp/dist/index.js", "--debug"]
}
}
}With the installed binary, use "command": "obsbot-mcp" and "args": ["--debug"].
Tools
35 tools on Windows and macOS, 34 on Linux (obsbot_gimbal_move_speed is unavailable there — see
limitations). --debug adds obsbot_debug_probe for
one more. Every Tail 2 camera on the network adds the 32 obsbot_tail2_* tools in
OBSBOT Tail 2 below (all platforms — they are pure TypeScript).
All names below are current as of v0.4.0 — every tool was renamed in this
release and there is no backward-compatible alias; see CHANGELOG.md for the
full old→new mapping if you're updating a caller.
The camera selector
Every camera-addressing tool accepts an optional camera parameter: the target camera's serial
number. Omit it with a single camera attached and nothing changes — this matches the server's
pre-v0.4.0, single-camera behaviour exactly. With more than one camera attached, a call that omits
camera fails with an error naming every attached serial, so you always know what to pass next.
Exempt (no camera parameter, ever): obsbot_devices (enumerates the whole fleet),
obsbot_capture_stop / obsbot_capture_list (address a sessionId, not a device), and
obsbot_debug_probe (operates on the current diagnostics transport). Two more tools honor it only
partially — see Capture below.
Multi-camera support is new in v0.4.0. It's exercised by the unit test suite against fakes; running two physical Tiny 2s at once has not yet been hardware-verified (see Known limitations).
obsbot_devices is the way to discover the serials you pass as camera: it reports each attached
camera's serial (where obtainable — reading it requires briefly opening the camera), name, and
status (available | bound | busy). A camera another process already holds comes back busy
with no serial, since it can't be opened to read one.
Device & power
Tool | Parameters | Description |
| — | List attached OBSBOT cameras with each one's serial (where obtainable), name, and status ( |
|
| Wake the camera/gimbal (sends |
|
| Sleep the camera/gimbal (sends |
|
| Read the live status block: |
Gimbal (PTZ)
Tool | Parameters | Description |
|
| Move the gimbal to an absolute angle. Positive yaw pans to the camera's left, positive pitch tilts down. Yaw clamped to |
|
| Drive the gimbal at a speed, then auto-stop after |
|
| Recenter the gimbal — drives it to yaw |
|
| Read the gimbal's current absolute |
|
| Point the camera at a pixel from a frame you just captured. Reads the camera's magnification from its own reported state — a discrete FOV mode or a continuous zoom alike — so it needs no FOV or zoom argument and works at any zoom. Refuses while AI tracking is active, when the FOV mode can't be decoded, or when a corrupt zoom reading would resolve to an implausible magnification, and refuses if the camera had to be woken (waking moves the gimbal, invalidating the frame you measured) or if the zoom is still ramping (the frame was captured at a magnification the camera has already left, so wait for the zoom and take a fresh snapshot). |
|
| Frame a region of a frame you just captured: centre the gimbal on it and zoom so the region fills the frame. Same refusal conditions as |
Aiming at what you can see
obsbot_capture_snapshot returns the frame as an image plus its width/height, so a model can
locate something in the picture and then point the camera at it:
obsbot_capture_snapshot— look at the frameobsbot_aim_at_pixel— pass the target's pixel and that frame's dimensionsobsbot_capture_snapshotagain — confirm it landed, and repeat if needed
Pass the frameWidth/frameHeight from the same snapshot the pixel came from. Mixing a pixel from
one frame with dimensions from another aims at the wrong place, and nothing can detect it.
The snapshot must be source: "device". obsbot_capture_snapshot can also read from
source: "virtual" or "ndi", which come from OBSBOT Center's own output rather than the camera's
raw stream. Those are framed and cropped by OBSBOT Center, not by this camera's optics, so the
measured field-of-view constants this tool relies on don't describe them — aiming from a virtual or
NDI frame lands in the wrong place with no way to detect it. Only use a device-source snapshot's
pixel and dimensions here.
The tool reads the camera's magnification itself — a discrete FOV mode or a continuous zoom alike
(m = 3*ratio-2, measured on hardware to better than 0.05%) — so there is no FOV or zoom argument to
get wrong, and it works at any zoom. It refuses rather than guessing when AI tracking is on (tracking
drives the gimbal and would fight the aim), when the FOV mode can't be decoded, when a corrupt zoom
reading would resolve to an implausible magnification, or when the camera had to
be woken from sleep (waking moves the gimbal, so the frame you measured no longer matches where the
camera is pointing — take a fresh snapshot and retry).
Framing what you can see
obsbot_zoom_to_fit extends the same idea from a point to a region: instead of just centring on a
pixel, it also zooms so that region fills the frame.
obsbot_capture_snapshot— look at the framePick a bounding box around whatever should fill the frame (a face, a whiteboard, ...)
obsbot_zoom_to_fit— pass the box (x,y,width,height) and that frame's dimensionsobsbot_capture_snapshotagain — confirm the framing, and repeat if needed
It shares obsbot_aim_at_pixel's refusals (AI tracking, undecodable FOV/zoom, a woken camera, a zoom still ramping, a
non-16:9 frame), and adds one of its own: the region must lie within the frame — edges included, so a
region that already IS the full frame is valid — with a positive width and height, or the call refuses
rather than guess what a negative width or an off-frame box was supposed to mean.
margin (default 0.1, i.e. 10%) backs the requested zoom off by that fraction so the region isn't
framed exactly edge-to-edge — some breathing room around it survives small aim/zoom error. The
region's two axes rarely need the same zoom to fill the frame; the tool always picks the smaller of
the two required magnifications, because zooming to the larger one would fill one axis by cropping the
other. The result is clamped to the camera's [1x, 4x] magnification range (reported via clamped) —
a region demanding more zoom than the camera has still gets the closest fit available, rather than
being refused outright.
The gimbal moves before the zoom is commanded. Zoom re-centres what's already in frame but does not keep a specific pixel under the crosshair as it changes — zooming first can push the region's centre out of frame entirely, which would make the subsequent move aim at a pixel that no longer means what it did when the caller measured it.
Zoom is not instantaneous: on this hardware it ramps toward the commanded value rather than jumping to
it, so a status read taken immediately after commanding it can catch it mid-transit (observed:
commanding ratio 1.5 read back partway there before settling). obsbot_zoom_to_fit polls for up to 3
seconds waiting for the zoom to arrive and returns settled:false — not an error — if it didn't. A
frame captured while the zoom is still moving is at an unknown magnification, so check settled
before trusting a follow-up snapshot; a false just means the camera was moving slower than expected,
not that anything failed.
Gimbal presets
Three on-device preset slots (1–3). Slots are create-once: obsbot_preset_save requires an
empty slot (delete first to reuse one); every other preset tool requires the slot to already be
occupied. Each tool re-reads the slot list after writing and returns a structured { ok:false }
failure if the device didn't land the change.
Tool | Parameters | Description |
|
| Read the three preset slots: occupied/empty, name, and pose in degrees. |
|
| Save the gimbal's current live pose into an empty slot. |
|
| Recall an occupied slot, driving the gimbal to its saved pose. |
|
| Overwrite an occupied slot with the gimbal's current live pose. |
|
| Rename an occupied slot (names over 40 bytes are truncated). |
|
| Delete an occupied slot, freeing it for |
Zoom
Two tools, not one — they ride different transports (standard UVC vs. the vendor command frame)
and produce different physical zoom at the same commanded ratio, so merging them would silently
change what ratio means. Pick by which behaviour you need.
Tool | Parameters | Description |
|
| Standard UVC zoom: set an absolute zoom ratio, clamped to |
|
| Vendor zoom path with adjustable speed: zoom to a ratio at a chosen speed ( |
AI tracking
Tool | Parameters | Description |
|
| Enable/disable AI tracking and choose the mode: a human framing ( |
|
| Set the tracking-speed preset (Center's Standard/Sport): |
|
| Enable or disable face-priority autofocus. |
Image & lens
Focus, white balance, and exposure each split into a dedicated _auto and _manual tool in
v0.4.0 (previously one tool with a mode parameter) — auto and manual take different parameters, so
splitting them lets each schema say exactly what it needs.
Tool | Parameters | Description |
|
| Set the field of view: wide (86°), medium (78°), narrow (65°). |
|
| Toggle HDR/WDR imaging on or off. |
|
| Enable continuous autofocus. |
|
| Set the focus motor to |
|
| Enable auto-exposure; optional |
|
| Set exposure |
|
| Enable auto white balance. |
|
| Set a colour temperature (clamped to device range). |
|
| Adjust |
Capture
obsbot_capture_record and obsbot_capture_preview do not take camera. They select a device
by source (device/virtual/ndi) through ffmpeg/ffplay, not by serial — there is no
serial-to-ffmpeg-device mapping yet. obsbot_capture_snapshot honors camera only for
source:"device"; for source:"virtual"/"ndi" the pixel source is still resolved by device
name, independent of camera.
Tool | Parameters | Description |
|
| Grab one still frame and return it as an image (for framing/lighting/exposure checks). |
|
| Start recording to MP4. Open-ended recordings auto-stop after 60 min; audio uses the OBSBOT mic; defaults to |
|
| Open a live preview window. Returns a |
|
| Stop a recording or preview session (recordings are finalized gracefully). No |
| — | List active recording/preview sessions. No |
Diagnostics (--debug only)
Tool | Parameters | Description |
|
| RE/diagnostics only — raw XU byte get/set and framed table queries. Advertised only under |
¹ record/preview shell out to ffmpeg/ffplay (install: winget install Gyan.FFmpeg
on Windows, brew install ffmpeg on macOS, apt install ffmpeg on Linux). snapshot does not
need ffmpeg — it grabs the frame through the native helper.
OBSBOT Tail 2 (network cameras)
The Tail 2 is a different animal: a network PTZ camera (NDI/RTSP/SRT/RTMP output, ethernet + WiFi,
and a USB-C port that is either MTP for footage offload or a UVC webcam). The tools below use its
network control plane — an HTTP REST API plus a WebSocket status push, documented in
TAIL2-PROTOCOL.md from on-device reverse engineering. That makes the whole
module pure TypeScript with zero platform-specific code: no native helper, no helper build,
identical behavior on Windows/Linux/macOS.
In UVC mode the camera also answers the standard UVC controls over USB (pan, tilt, zoom, pan/tilt speed, focus, exposure, white balance), and its vendor Extension Unit gives tracking on/off, human tracking mode, tracking speed and Auto Zoom. This server has no Tail 2 USB transport yet, so none of that is exposed as a tool; presets, portrait rotation and roll trim are network only. See TAIL2-PROTOCOL.md §11 for what was measured.
Getting started:
Put the Tail 2 on your network (it defaults to DHCP on ethernet).
Call
obsbot_tail2_scanonce — it listens ~5 s for the camera's own mDNS announcements (it multicasts its MAC, name and IPs every few seconds, the same channel OBSBOT Center discovers by) and falls back to an HTTP subnet sweep only on multicast-filtered networks. Or setOBSBOT_TAIL2_HOSTS(comma-separated addresses) in the server's environment, or just pass any Tail 2 tool acameraparameter holding the camera's IP.
The camera selector for these tools is the camera's MAC (its stable identity), or its
device_name, or its host address — case-insensitive. Omitted with one Tail 2 registered it
resolves to that one.
Tool | Parameters | Description |
| — | List registered Tail 2 cameras (mac, name, host(s)). Registration is not liveness. |
| — | Discover Tail 2 cameras: ~5 s mDNS listen (the camera's own announcements), HTTP subnet sweep fallback. Registers what it found. |
|
| One WebSocket status push: the full live block — power, rec, portrait, AI mode + tracking settings, zoom ratio, roll bias, focus modes, NDI/RTSP/SRT/RTMP flags, SD card, presets with poses, and per-subsystem health. No live yaw/pitch — Tail 2 gimbal moves are open-loop. |
|
| Identity + static config in one call: device_info, range (zoom 1.0–12.0, focus 1–100, WB 2000–10000 K), networkconfig (NDI/stream settings). |
|
| Absolute zoom on the camera's own ratio scale. |
|
| Unlock (or re-lock) the 5–12x digital zoom region — the switch OBSBOT Center owns, which does not exist in the REST API (zoom above 5.0 is acknowledged but pinned until this runs). Speaks Center's private UDP-9999 protocol directly, both checksums decoded; verified by readback from the live status snapshot ( |
|
| Gimbal recenter (yaw/pitch/roll 0). Open-loop: returns on ack. Can drop AI tracking to |
|
| Live gimbal pose in degrees (yaw/pitch/roll) + zoom ratio, via save-scratch-preset → read → delete (~2 s). Needs one empty preset slot (probe slots are self-cleaning). |
|
| Jog the gimbal (the joystick primitive) with automatic stop. Measured: command 20 ≈ 7.9°/s; positive yaw command DECREASES recorded yaw. Disable AI tracking first. |
|
| Absolute move via closed loop (jog → pose read → correct, ≤5 rounds, ±1.5° tolerance). Slow by design (~1 s per small move, more for large). Refuses while AI tracking is active. Hardware-verified 2026-09-30: 4° move converged in one iteration. |
|
| Read (bare call) or set (verified) control-direction inversion. |
|
| Read/control the recording switch. Needs storage — check |
|
| Trigger a still photo (to storage). |
|
| Read or set focus. Position only in |
|
| Read (mode-appropriate) or set exposure: auto (face-priority AE + EV bias) or manual (ISO 100–6400, shutter |
|
| Read or set image style (brightness/contrast/hue/saturation/sharpness). Control writes are gated to |
|
| Read/set HDR. |
|
| Read/set white balance (6 modes + Kelvin in manual). |
|
| Read/set the ONE active network output. The programmatic way to arm SRT for |
|
| Read/set OnlyMe human-tracking (track one person, not everyone). |
|
| Read/set audio input. |
|
| Read/set the auto-zoom framing pattern (Center's tracking slider). Writes gated: requires human tracking armed. |
|
| Read/set the preset switching speed. |
|
| Read/set the USB-C function mode. |
|
| Read/set anti-flicker. |
|
| Read/set the autofocus track mode. |
|
| Read/set the auto-ISO range (double slider; omitted bounds keep current). |
|
| Read (bare) or set the gesture-control switches and zoom factor. |
|
| Read (bare, incl. RTSP URLs) or set the stream encoder config. Resolution writes gated on output off. |
|
| Download the full diagnostic bundle (the Export Log archive: ust.json, status.json, factory records, kernel logs) to |
|
| Fresh runtime snapshot distilled: live gimbal pose (euler+joint), zoom internals incl. |
|
| Toggle zone tracking (UDP 9999 key 03). No readback exists — effect shows in tracking behavior. |
|
| Set the AI auto-zoom framing speed (UDP key 0x17; distinct from manual zoom speed). No readback. |
|
| Read (bare) or configure custom tracking speed: per-axis speeds, Auto buttons (readback-verified via tracking_settings). Axis locks shown read-only — no known write path. |
|
| Tap-to-track: engage AI tracking on the subject at a normalized frame coordinate (snapshot pixel ÷ frame size). Undocumented endpoint, hardware-measured. |
|
| Tap-to-focus: move the focus window to a normalized frame coordinate and start a point focus (afc/afs only). Bare call reads the window. |
|
| Motorized 90° barrel rotation for portrait framing — no Tiny 2 equivalent. Verified on orientation feedback; retry if |
|
| Roll trim in degrees (horizon correction / deliberate tilt). |
|
| Enable/disable AI tracking. Modes are the camera's own enum: human (single/group), animal (normal/close-up), object tracking ( |
|
| Tracking follow speed — the Tail 2's own six-speed enum (NOT the Tiny 2's standard/sport pair). |
|
| The three preset slots: occupied/empty, decoded name, pose in degrees + zoom ratio. |
|
| Save the CURRENT live pose into a slot — aim first, then save. Overwrites an occupied slot (unlike the Tiny 2's create-once), and there is no pose-by-value write in this API at all. |
|
| Drive gimbal + zoom to a saved pose. Refuses an empty slot. Arrival verified via the zoom readback (gimbal axes are open-loop — the Tail 2 reports no live pose). Disable AI tracking first or it fights the move. |
|
| Free a slot. |
|
| Rename a slot (base64 handled transparently). |
|
| Grab one still frame and return it as an image. Pulls from the camera's SRT output (ffmpeg SRT caller, port 5000) — requires SRT listener mode enabled in OBSBOT Center first (it disables NDI while active and is cleared by a camera reboot; Center only allows changing streaming settings while the output is off, so a failure usually means SRT is off — re-enable and retry). Hardware-verified 2026-09-26. |
Every write is verified by readback (settled in the result): the Tail 2 acknowledges commands
it may drop while an actuator is mid-motion — {"code":200} means acknowledged, not applied
(measured 2026-09-26; see TAIL2-PROTOCOL.md §8). settled:false means the ladder ran out before
the readback agreed; the command was still sent — retry it.
Known Tail 2 gaps (details in TAIL2-PROTOCOL.md §9): image-control writes via REST are unprobed, and snapshot depends on the camera's SRT output being enabled from OBSBOT Center (exclusive with NDI, cleared by reboots — the tool explains this when SRT is off; on this firmware the RTSP output never serves despite its enable flag). The camera's control API is also unauthenticated — anyone on your LAN can drive it, including poweroff.
Supported platforms
Windows x64 — supported today. The native helper is built from source in
native/windows/(CMake + MSVC); the published npm package ships a prebuilt binary so end users need no toolchain.Linux x64 — supported from v0.2. The native helper is in
native/linux/(CMake + GCC); it uses V4L2 for standard UVC controls (zoom, focus, exposure, pan/tilt, white balance, image controls) andUVCIOC_CTRL_QUERYfor vendor Extension Unit commands (gimbal speed/AI tracking, wake/sleep, HDR, FOV). Snapshots capture a MJPEG or YUYV frame via V4L2 mmap streaming and encode to JPEG using libjpeg. Thelinux-x64prebuilt binary ships with the published npm package. Build dependencies:build-essential cmake libjpeg-dev libv4l-dev.Gimbal position reads (
obsbot_gimbal_position) reflect the last-commanded value, not live in-flight position — see "Linux gimbal position feedback" below for why, and what would fix it.macOS 14+ (Apple Silicon and Intel) — supported. The native helper is in
native/macos/(Objective-C + IOKit/AVFoundation). It uses IOKit USB control transfers for both standard UVC controls and vendor Extension Unit commands, and AVFoundation for enumeration and snapshots. Bothdarwin-arm64anddarwin-x64prebuilt binaries ship with the published npm package (darwin-x64also covers Apple Silicon running Node under Rosetta, whereprocess.archreportsx64). macOS 14 is the floor because the helper usesAVCaptureDeviceTypeExternal; the build pins-mmacosx-version-minso the binary does not inherit the build machine's OS as its minimum.Note on macOS specifically:
UVCAssistant(a DriverKit system extension) owns the camera's UVC interfaces exclusively, soUSBInterfaceOpen— and evenUSBInterfaceOpenSeize— fail withkIOReturnExclusiveAccess. The helper therefore opens the USB device, which is not locked, and issues UVC control requests on its default control endpoint. This coexists withUVCAssistant: the camera keeps working as a normal webcam while under control, so no driver-replacement step is needed.
Building the native helper (Linux)
cd native/linux
mkdir build && cd build
cmake ..
make -j$(nproc)
make install # copies to native/prebuilt/linux-x64/Linux gimbal position feedback is not live
obsbot_gimbal_position on Linux reports the last position obsbot_gimbal_move/
obsbot_gimbal_recenter commanded — not a live, in-flight reading. Hardware testing (2026-07-21)
confirmed the OBSBOT Tiny 2's CT_PANTILT_ABSOLUTE control genuinely tracks live position — a raw
USB read of that same control, bypassing the kernel, showed a real slew progressing in real time.
The reason plain V4L2 (VIDIOC_G_CTRL) never sees that is that uvcvideo caches the control's
value and serves the cache instead of re-querying the device (confirmed via
VIDIOC_QUERY_EXT_CTRL, which reports no V4L2_CTRL_FLAG_VOLATILE). The driver invalidates that
cache when the device sends a UVC Control Change interrupt — which this camera's firmware never
does, and never advertises support for.
Getting a genuinely live reading through V4L2 requires briefly detaching the kernel driver from the camera's control interface and reading the control directly over raw USB — but detaching that interface (even briefly, even without writing anything) breaks any concurrent video capture on this device: streaming and control share one kernel-managed USB function, so pulling the driver off one takes both down together. That makes a libusb-based workaround incompatible with anything actually using the camera as a webcam at the same time, which ruled it out as a shipped default.
A kernel patch has been submitted upstream (media: uvcvideo: query pan/tilt position from the device on every read,
July 2026 — awaiting review, not merged). It marks CT_PANTILT_ABSOLUTE volatile so the driver
queries the device on every read; verified on this hardware to track a live slew through plain
VIDIOC_G_CTRL, concurrently with streaming. If it is accepted, obsbot_gimbal_position becomes
live on Linux with no code changes needed here. Until it ships in a kernel near you:
obsbot_gimbal_moveandobsbot_gimbal_recenterwork normally — hardware-verified, repeatedly, via direct V4L2VIDIOC_S_CTRLwrites. Their target values are known and clamped before being sent, so they can't exceed the gimbal's mechanical range regardless of the missing feedback.obsbot_gimbal_move_speedis not available on Linux (hidden from the tool list entirely, not just refused at runtime). A speed×duration burst has no target position to clamp — without a live reading to confirm where the gimbal actually is, there's no way to bound it against the mechanical limits before it gets there. It remains available on Windows/macOS.
Building the native helper (macOS)
make -C native/macos # -> native/prebuilt/darwin-arm64/obsbot-helperRequires the Xcode command line tools. CMake works too (cmake -S native/macos -B native/macos/build && cmake --build native/macos/build), which is what CI uses.
Known limitations
What has actually been exercised against hardware, and what hasn't:
Platform | Status |
| Hardware-verified — mid-session disconnect recovery ( |
| Hardware-verified — gimbal absolute moves and recenter via V4L2 (20 consecutive moves with a live preview running), and the arc-second scaling fix confirmed by physical swing. Gimbal position is not live and |
| Hardware-verified — control, gimbal movement and per-axis position readback, zoom, snapshot, USB vid/pid candidacy, serial readback and serial-keyed binding, single-owner IPC coordination, helper-death recovery, and unaided recovery from an unplug/replug, on a real Tiny 2 |
| Build-verified only — compiles with the right architecture and deployment target, never executed |
The Intel (
darwin-x64) helper has never been run. No Intel Mac was available to test it. It cross-compiles cleanly and is packaged, but nothing has confirmed it talks to a camera. It also covers Apple Silicon running Node under Rosetta, whereprocess.archreportsx64— likewise untested. Reports from Intel users are welcome.macOS 14 or newer is required, and macOS runtime is verified on 26.5 only. The helper uses
AVCaptureDeviceTypeExternal(macOS 14+), so the build pins-mmacosx-version-min=14.0. The binary will load on 14 through 25, but behavior there is untested — in particular the UVC control path relies onUVCAssistantholding the camera's UVC interfaces while leaving the USB device itself openable. That is how current macOS behaves; older releases are unconfirmed.The first snapshot on macOS raises a camera permission prompt. The helper is a plain CLI tool with no bundle identifier, so macOS attributes camera access to whichever app spawned it — your MCP client — and that app is named in the prompt and holds the grant. Approve once; the grant survives helper updates, since it is keyed to the client rather than to the helper's signature.
AI tracking overrides manual gimbal moves. When AI tracking is active (the Tiny 2's default on wake), a commanded pan/tilt executes and is then pulled back to the tracked subject —
obsbot_gimbal_positionshows the yaw/pitch move out and decay back to rest. This is the camera's behaviour, not a bug: turn tracking off for unopposed manual control.The camera may not enumerate through a USB hub or dock. A Tiny 2 connected through a USB-C dock was invisible to
ioregandsystem_profilerentirely — not just to this server. Ifobsbot_devicescomes back empty, try a direct connection before assuming a software fault.Only the OBSBOT Tiny 2 is supported. On Windows and macOS candidacy is gated on the Remo USB vendor ID plus a known-model product ID (
0x3564/0xFEF8), so no other model is detected at all — and a name-matching software source, such as the "OBSBOT Virtual Camera" that OBSBOT Center registers, is rejected because it reports no vid/pid. Linux still matches by name, because its helper does not report vid/pid yet, so a different OBSBOT may be found there — but the vendor command set is Tiny 2 specific either way. (On macOS the virtual camera cannot appear at all: the helper enumerates USB devices through the IORegistry, which a software camera never enters.)The vendor reply mailbox is unreliable for several seconds after a replug. On 2026-07-21 a Tiny 2 was seen returning only the host's own echoed request frame from the vendor reply mailbox (XU selector 2) — magic byte
0xaacleared to0x00, every other byte identical — for a continuous 3.2 s.readSerial()threw,bind()found no serial, and every tool needing a bound camera failed with "no OBSBOT camera found" while the device was plainly healthy: correct vid/pid, opened fine, XU node 2, live status block on selector 6.That was unexplained for a while. It is now reproducible: immediately after a USB re-enumeration. Polling
readSerialevery 50 ms across a replug failed 22 times in 80 attempts spread over the first 14 s, against 0 in 120 in steady state; the first read after arrival showed exactly the echoed-request signature above, and later failures showed the reply slot populated but with its magic byte still zeroed. Ruled out as causes: stale per-process device state (the same long-lived helper read a brand-new uniqueID cleanly at t+49 ms), re-opening the device (0/40 either way), and the per-transport sequence counter restarting at 1 (0/80).Consequence for callers: a bind attempted in the first seconds after a replug can fail even though the camera is fine. Retrying works. The arrival-driven re-bind now retries on a bounded ladder for this reason, and a
readSerialfailure reports what the mailbox actually held (echoed request / unparseable / a reply to another request) rather than only "no valid reply". Ruled out earlier and still ruled out: reply latency (polled 3.2 s), the wrong extension unit (the VideoControl interface exposes exactly one,bUnitID 2), the wrongwLength(every XU selector is 60 bytes byGET_LEN), the reply arriving on another selector (1–19 swept), camera sleep state, and contention from OBSBOT Center.Recovery after a replug is proactive, but not in every case, and it differs by platform. The server subscribes to OS camera arrival/removal events, so in the common case a replugged camera re-binds itself with no tool call —
obsbot_devicesreports itboundagain on its own. Every cell below is hardware-measured:scenario
macOS
Windows
Linux
same-port replug
proactive
proactive
next tool call
different-port replug
proactive
next tool call
next tool call
Where it says "next tool call", nothing is stranded — the call that follows detects the stale binding, prunes it and re-binds. It costs one failed call, which is exactly how every platform behaved before these events existed. The Windows difference comes from its arrival filter requiring a path it has already enumerated, which is also what stops the Tiny 2's audio interface from being reported as a second camera; macOS has no equivalent problem because it re-binds by serial and ignores the path. Linux emits no bus events at all yet.
Note the interaction with the mailbox entry above: a re-bind attempted immediately after a replug can still lose the first attempt, so the server retries on a short bounded ladder.
Two-camera operation is not yet hardware-verified. The
cameraselector and the per-camera device registry are covered by the unit test suite against fake transports; running two physical Tiny 2s attached at once has not been confirmed on real hardware (a second unit wasn't available). Single-camera use is unaffected either way.Linux gimbal position feedback is not live, and
obsbot_gimbal_move_speedis unavailable there as a result. See "Linux gimbal position feedback is not live" above — a kernel patch fixing this at the source has been submitted upstream (July 2026, awaiting review).obsbot_gimbal_moveandobsbot_gimbal_recenterare unaffected; both are hardware-verified to work normally.obsbot_aim_at_pixelis affected — it depends on a live pose reading to compute the aim, the same wayobsbot_gimbal_positiondoes.obsbot_zoom_vendor's ratio scale doesn't matchobsbot_zoom_uvc's at the sameratio. A hardware snapshot comparison atratio: 2.0showed the vendor path framed tighter than the UVC path. Whether the vendor-side ratio encoding is off by a scale factor, or the two zoom controls simply have different physical ranges, isn't determined yet — one comparison isn't enough to tell. Tracked separately; useobsbot_zoom_uvcif you need the ratio to land exactly.
No proprietary SDK
This project speaks the camera's USB protocol directly through the OS's standard UVC driver stack and
does not use, link, bundle, or ship any vendor SDK. See PROTOCOL.md for the
protocol reference (frame format, checksum, command table).
How it works
The camera exposes two independent control surfaces, both reachable through the OS's standard UVC (USB Video Class) driver stack — this project never talks to the USB device directly, so the OS keeps mediating access and the camera remains usable as a normal webcam at the same time commands are sent:
Standard UVC controls — zoom (
CT_ZOOM_ABSOLUTE), focus and exposure (IAMCameraControl), gimbal position readback (UVC Pan/Tilt), and the image controls plus white balance (IAMVideoProcAmp) — are the camera's built-in UVC properties, driven via DirectShow on Windows.Vendor commands — gimbal moves, recenter, wake/sleep, AI tracking, HDR, and field of view — are sent through the camera's UVC Extension Unit, driven via
IKsControl::KsPropertyagainst the XU's topology node on Windows.
Both are issued through a small native helper process (obsbot-helper.exe on Windows, obsbot-helper
on Linux) that the Node server spawns and talks to over a line-delimited JSON-RPC protocol on
stdin/stdout. The helper is the only platform-specific piece; the codec (frame encoding, CRC-16/USB
checksum, command table), transport abstraction, device manager, and MCP tool definitions are all pure
TypeScript/JavaScript and shared across platforms.
Verifying against real hardware
scripts/e2e.mjs drives the built stack (dist/) against a physically connected camera: it wakes the
device, zooms in, pans the gimbal, recenters, zooms back out, and puts the camera to sleep, with a short
pause and console log before each step so a human can watch it happen. This moves the physical gimbal —
only run it under supervision:
npm run build
node scripts/e2e.mjsTesting changes through the live MCP tools
Two traps make it easy to test the wrong thing and believe the result. Both cost real time on 2026-07-25.
Rebuild, then start one server from the new build. The MCP server runs from dist/, so a
source change is invisible until npm run build. This project coordinates concurrent clients by
electing a single owner process (see IPC-DESIGN.md), and the owner executes every tool call,
whichever session made it. The owner is the newest build alive: when a server starts from a newer
build than the owner's, the owner finishes the call in flight, releases the camera, and hands the
endpoint over. Reloading the server in one MCP client is therefore enough. Other sessions stay open
and their calls are served by the new build.
Three things to know:
The tool list of other sessions does not refresh. A session that started on an older build keeps the list it had; its calls run the new code. Reload that session's server to get new tools.
A server that predates the handover cannot hand over (0.7.0 and earlier). A newer server will not run calls on it, and says so:
an older instance owns the camera endpoint and cannot hand it over. Stop the old process once; after that, rebuilds take over on their own.Use
npm run buildandnpm run build:helper. They stamp the build with its identity (dist/build-info.json). A baretsc, or a helper copied into place by hand, is not stamped and does not count as a newer build.
To see which build will run your next call, read the last ipc role= line in the server's log.
role=owner means this process; role=client owner-pid=… owner-build=… names the one that will.
Claude Code keeps the log under ~/Library/Caches/claude-cli-nodejs/<project>/mcp-logs-obsbot/ on
macOS.
# Linux / macOS — every server process, with when it started
ps -o pid,lstart,command -p $(pgrep -f "obsbot.*dist/index.js" | paste -sd, -)# Windows
Get-CimInstance Win32_Process -Filter "Name='node.exe'" |
Where-Object { $_.CommandLine -like "*Obsbot*" } |
Select-Object ProcessId, CreationDateSet OBSBOT_IPC_NAME to give a server a rendezvous endpoint of its own (1–64 characters from
A-Z a-z 0-9 . _ -). Test harnesses do this so they leave a live session alone. Two owners on one
machine can contend for the camera, so it is not for everyday use.
Frame rate selects the field of view, so the preview shows less than snapshots. At 1920×1080 this
camera has two different windows onto the sensor and the frame rate picks between them — not the
codec. Measured at one pose and one zoom: MJPEG@30 vs YUYV@30 came out at scale 1.00001 over 2382
inliers, i.e. the same field to within a fifth of a pixel, while MJPEG@60 is a 1.214× crop of both.
obsbot_capture_preview pins 60fps for smooth motion, so it shows ~21% less of the room than
obsbot_capture_snapshot, which negotiates 1080p30. Framing by eye in the preview and then aiming at a
pixel from a snapshot will not agree. The geometry constants describe the 30fps field. Any measurement
you make must state pixel format and frame rate; neither a resolution nor a codec alone identifies
the field.
A preview holds the camera stream, so snapshots fail while one is open. obsbot_capture_preview
and obsbot_capture_snapshot both need the device stream, and on Windows the second one gets
Camera is in use by another application. This matters for the aim loop above, which is
snapshot → aim → snapshot: stop the preview around each snapshot, or work without one. Gimbal control
is unaffected — it uses control transfers, not the stream — so obsbot_aim_at_pixel itself works fine
with a preview running. The error text suggests source: "virtual" or "ndi" as a workaround. That is safe for looking,
and safe for aiming only if the feed is an unmodified pass-through of the camera — declare it with
the source parameter on obsbot_aim_at_pixel / obsbot_zoom_to_fit and read the note it returns.
A compositor that rescales or letterboxes the frame silently invalidates the geometry.
Available Tools
80 toolsobsbot_aim_at_pixelA
Point the camera at a specific pixel in a frame you just captured. Give the pixel's x/y and the frameWidth/frameHeight from THE SAME obsbot_capture_snapshot result — mixing a pixel from one frame with dimensions from another aims at the wrong place and cannot be detected. source DECLARES which feed the frame came from (default device); it cannot be inferred from the pixels. A virtual or ndi frame is accepted, but only aims correctly if that feed is an unmodified pass-through of the camera — a compositor's rescale or letterbox is invisible in the picture and silently wrong here — so a non-device declaration comes back with that assumption stated. Takes no field-of-view or zoom argument: it reads the camera's magnification from its reported state, a discrete FOV mode or a continuous zoom alike, so it works at any zoom. Refuses when AI tracking is active (tracking moves the gimbal itself and would fight the aim), when the FOV mode can't be decoded, or when a corrupt zoom reading would resolve to an implausible magnification, so it never aims on an assumption it cannot check. If the camera was asleep, waking it moves the gimbal and invalidates the frame you measured, so the call refuses instead of aiming on stale geometry — take a fresh snapshot and retry. Returns clamped:true if the target was outside the gimbal's range; the camera still moves, to the nearest reachable pose. Refuses (ok:false) instead of moving when the pixel lies past vertical from the current pose — reachable only by an "over the top" rotation that would swing the camera toward the opposite side of the room, not toward the target; tilt toward the pixel first, then re-aim.
| Name | Required | Description | Default |
|---|---|---|---|
| x | Yes | ||
| y | Yes | ||
| camera | No | ||
| source | No | device | |
| frameWidth | Yes | ||
| frameHeight | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of disclosing behavior. It does so thoroughly: it reads the camera's magnification from reported state, refuses instead of aiming on uncheckable assumptions, returns clamped:true and still moves when target is out of range, refuses with ok:false for over-the-top rotations, and invalidates stale geometry if the camera was asleep. This is exceptionally transparent about side effects and failure modes.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is long, but nearly every sentence earns its place by adding a caveat, refusal condition, or side-effect disclosure. The first sentence front-loads the primary purpose. The main weakness is that it is a dense wall of text with multiple embedded clauses, which could be easier to scan with better structuring.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity, the absence of an output schema, and no annotations, the description covers an impressive amount: prerequisites, refusal conditions, return flags (clamped, ok:false), and the virtual/ndi pass-through assumption. It does not specify the full shape of a successful response or explain the 'camera' parameter, so it is not fully complete, but it is close.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It explains x/y as pixel coordinates, emphasizes that frameWidth/frameHeight must come from the same snapshot, and describes the source enum with its default and pass-through caveat. However, the optional 'camera' parameter in the schema is never addressed, leaving one parameter under-specified.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description begins with a specific verb and resource: 'Point the camera at a specific pixel in a frame you just captured.' It clearly distinguishes this from sibling gimbal/zoom tools by targeting a pixel coordinate rather than angles, FOV, or zoom, and further differentiates itself by noting it takes no FOV/zoom argument.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives explicit when-to-use and when-not-to-use guidance: use coordinates and dimensions from the same obsbot_capture_snapshot result, declare the source, and only use virtual/ndi feeds that are unmodified pass-through. It also lists refusal conditions (AI tracking active, undecodable FOV, corrupt zoom, camera asleep, target past vertical) and tells the agent how to recover: take a fresh snapshot after wake, or tilt toward the pixel before re-aiming.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
obsbot_ai_trackA
Enable or disable AI tracking and choose the mode. When enabled the camera follows the subject; disabling stops tracking. mode is either a human framing (normal | upper-body | close-up | headless | lower-body) or a standalone scene mode (group | whiteboard | desk | hand); scene modes imply enabled:true. After writing, the tool polls the status block until the mode settles and returns { verified, matched } — the aiMode the device actually landed on (matched:false means no subject was being tracked, so the mode could not take effect yet).
| Name | Required | Description | Default |
|---|---|---|---|
| mode | No | normal | |
| camera | No | ||
| enabled | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description discloses polling behavior after writing, the return value with verified and matched fields, and clarifies that scene modes set enabled to true. No annotations are provided, so the description adequately covers behavioral traits.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is concise and well-structured, starting with the core action, then elaborating on mode semantics, and ending with polling behavior. It is not overly long but could be slightly more streamlined.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool has three parameters and no output schema, the description covers the main action, mode details, and return value. It lacks mention of prerequisites or error cases, but is largely complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The description adds meaning to mode by explaining human framing vs scene modes and the implication on enabled. However, the camera parameter is not explained. Schema coverage is 0%, so the description provides significant value but misses one parameter.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool enables or disables AI tracking and selects the mode. It mentions that the camera follows the subject and distinguishes from sibling tools like obsbot_ai_track_speed.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explains when to use the tool (to control AI tracking) and notes that scene modes imply enabled:true. It does not explicitly exclude alternatives, but provides enough context for proper use.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
obsbot_ai_track_speedA
Set the AI tracking-speed preset (OBSBOT Center's Standard/Sport). speed: standard (slower follow) | sport (snappier follow).
| Name | Required | Description | Default |
|---|---|---|---|
| speed | Yes | ||
| camera | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations exist, so description carries full burden. It states the action (set speed) but omits side effects, required system state (e.g., tracking must be active), or whether changes are persistent. The camera parameter is also undocumented in behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single sentence followed by a brief clarification of the enum values. No unnecessary words, front-loaded and efficient.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
While the core function is explained, the omission of the camera parameter and lack of output schema leave the tool's full behavior partially undefined. For a simple setter, it is adequate but not complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The description adds valuable meaning for the speed parameter ('slower follow' vs 'snappier follow') beyond the enum values. However, the camera parameter is not mentioned at all, leaving its purpose unclear.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool sets the AI tracking-speed preset, with explicit enumeration of the two options (standard/sport) and their meanings. It is specific and distinguishes from sibling tools like obsbot_ai_track.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No explicit guidance on when to use this tool versus alternatives, or when to choose standard vs. sport. The description implies use for adjusting tracking speed but lacks context or prerequisites.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
obsbot_capture_listA
List active recording/preview sessions (id, kind, source, output path, start time).
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations provided, so description carries full burden. It implies read-only behavior by stating 'list', but does not explicitly confirm no side effects or require any prerequisites.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Single concise sentence that includes all key elements: action, resource, and returned fields. No wasted words.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given no parameters or output schema, description adequately specifies what the tool returns. Could clarify what 'active' means or handle empty results, but sufficient for a simple list operation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Tool has zero parameters and schema coverage is 100%. Baseline for 0 parameters is 4; description does not add parameter info since none exist.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Description clearly states the verb 'List' and resource 'active recording/preview sessions', specifying the fields returned. Immediately distinguishes from sibling tools like obsbot_capture_stop or obsbot_capture_preview.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No explicit when-to-use or alternatives provided. However, the simplicity of the tool and absence of similar list tools among siblings makes the context clear enough.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
obsbot_capture_previewA
Open a live preview window of the camera (for the user to watch). NOTE: before calling, ensure the camera is focused (call obsbot_focus_auto for autofocus) unless otherwise directed. source: device|virtual|ndi. Returns a sessionId for obsbot_capture_stop.
| Name | Required | Description | Default |
|---|---|---|---|
| source | No | device |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Describes the preview window opening and return of sessionId, but does not detail side effects, blocking behavior, or resource impacts. No annotations to supplement.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two concise sentences plus a note, front-loaded with the core action, no extraneous content.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the simple parameter and no output schema, description covers return value, prerequisite, and source options; reasonably complete for this tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, description mentions the source parameter with its enum values but does not elaborate on each option's meaning.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Description clearly states the action (open a live preview window) and distinguishes from siblings like obsbot_capture_snapshot and obsbot_capture_stop.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides explicit prerequisite (call obsbot_focus_auto first) and lists source options, but does not explicitly state when not to use or alternatives.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
obsbot_capture_recordA
Start recording the camera to an MP4 (for the user). durationSec optional (open-ended recordings auto-stop after 60 min); audio defaults to on (the OBSBOT mic); outputPath optional (defaults to ~/Videos/OBSBOT on every platform, including macOS, where that is NOT the usual ~/Movies). NOTE: before calling, ensure the camera is focused (call obsbot_focus_auto for autofocus) unless otherwise directed. source: device|virtual|ndi. Returns a sessionId for obsbot_capture_stop.
| Name | Required | Description | Default |
|---|---|---|---|
| audio | No | ||
| source | No | device | |
| outputPath | No | ||
| durationSec | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries full disclosure burden and does it well: it reveals the 60-minute auto-stop for open-ended recordings, that audio is on by default from the OBSBOT mic, the cross-platform outputPath default including the macOS quirk, and the returned sessionId. These are exactly the non-obvious behaviors an agent needs before invoking.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense but front-loaded: the core action is in the first clause, and the parenthetical defaults and platform caveat come after. Some punctuation is run-on, but each clause adds operational value rather than filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with no annotations and no output schema, this covers the essentials: what it does, defaults, limits, prerequisite focusing step, and return value. An agent could invoke it correctly and know what to expect and how to stop it afterward.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, and the description compensates by explaining durationSec's open-ended behavior, audio's default and source, outputPath's default location, and the source enum values. Only 'source' semantics are thin beyond the enum list, but the schema already enumerates valid values.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and object: 'Start recording the camera to an MP4', which clearly distinguishes it from sibling capture tools like obsbot_capture_snapshot or obsbot_capture_preview. It also grounds the tool's role by noting it returns a sessionId used by obsbot_capture_stop, making its place in the workflow explicit.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives a concrete precondition: ensure the camera is focused before calling, and names obsbot_focus_auto as the way to satisfy it, with an exception ('unless otherwise directed'). It doesn't explicitly contrast with alternatives (e.g., when to choose snapshot vs record), but the recording target and MP4 context make the intended situation clear.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
obsbot_capture_snapshotA
Grab one still frame from the camera and return it as an image (for you to see and for framing/lighting/exposure checks). resolution is the longest edge in pixels, 256-1920, default 640 — larger images cost proportionally more tokens, so ask for more only when you need the detail. NOTE: before calling, ensure the camera is focused (call obsbot_focus_auto for autofocus) unless otherwise directed. source: device (default) | virtual | ndi. If the camera is in use by another app, returns a message instead of an image.
| Name | Required | Description | Default |
|---|---|---|---|
| camera | No | ||
| source | No | device | |
| quality | No | ||
| settleMs | No | ||
| resolution | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries full burden. It discloses token cost scaling with resolution, source options, and conflict behavior. Missing details on output format and blocking nature, but adequate for a simple read operation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Description is well-structured with purpose first, then resolution note, prerequisite, source list, and failure scenario. Each sentence adds value. Could integrate source line more smoothly, but overall concise and organized.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given 5 parameters, 0% schema coverage, no output schema, and no annotations, the description covers purpose, resolution, prerequisite, and failure mode but omits quality, settleMs, camera, and output format. Adequate for basic use but lacks completeness for nuanced agent decisions.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so description must compensate. Only resolution and source are explained; parameter quality, settleMs, and camera are not described. Quality and settleMs have no explanation of their effect, and camera accepts any string without guidance.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states 'Grab one still frame from the camera and return it as an image' with a specific verb and resource. It distinguishes from siblings like obsbot_capture_preview (likely video preview) by focusing on still frame capture.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides explicit prerequisite to ensure camera is focused via obsbot_focus_auto. Mentions source options (device/virtual/ndi) and a failure mode when camera is in use. Lacks explicit contrast with alternatives like obsbot_capture_preview, but the still frame vs preview distinction is clear enough.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
obsbot_capture_stopA
Stop a recording or preview session by its sessionId. Recordings are stopped gracefully so the MP4 finalizes correctly.
| Name | Required | Description | Default |
|---|---|---|---|
| sessionId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries full burden. It discloses that recordings are 'stopped gracefully so the MP4 finalizes correctly,' indicating safe termination. However, it does not describe what happens to preview sessions.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two concise sentences, no redundant information. Every word is necessary and front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's simplicity (one param, no output schema), the description covers the essential purpose and a behavioral note. It could mention that the session must be active or where to get the sessionId, but overall it's fairly complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so description must compensate. It only says 'by its sessionId' without explaining what a sessionId is or how to obtain it, adding minimal value beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action (stop) and resource (recording or preview session), and specifies the method (by sessionId). It distinguishes from sibling capture tools like capture_snapshot or capture_preview.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description does not provide any guidance on when to use this tool versus alternatives (e.g., when to stop vs. capture_snapshot), nor does it mention prerequisites like an active session.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
obsbot_devicesA
List attached OBSBOT-compatible cameras. Each entry is { serial?, locationId?, name, status, reason? }: serial is the value to pass as camera to any camera-addressing tool (present where obtainable — reading it requires briefly opening the camera); status is available (free to bind), bound (already opened by this process), or busy (could not be opened and identified, so it can't be targeted here — usually another process holds it, but it also covers a camera that opened yet would not answer). On a busy entry, reason carries the underlying error and distinguishes those cases.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description fully discloses behavioral details: reading serial requires opening the camera, status meanings and edge cases (busy with reason), and the distinction between availability conditions.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, front-loaded with the core purpose, then detailed structure explanation. No extraneous content.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Comprehensive coverage of return fields, including serial absence, status values, and error reasons. Complements lack of output schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema has zero parameters, so the description adds no parameter info; baseline 4 applies as per rules.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the verb ('List') and resource ('attached OBSBOT-compatible cameras'), and differentiates from sibling tools like 'obsbot_status' by focusing on device enumeration.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explains that the 'serial' field is needed for camera-addressing tools, implying a prerequisite relationship, but does not explicitly state when not to use this tool or provide alternatives.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
obsbot_focus_autoC
Enable continuous autofocus.
| Name | Required | Description | Default |
|---|---|---|---|
| camera | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure, but it only states the immediate effect. It does not mention whether this overrides manual focus, whether it persists across sessions, or what observable side effects occur on the camera.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is four words with zero filler, and the core action is immediately front-loaded. It is efficiently concise, though it verges on being too sparse to serve as a complete definition.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a one-parameter tool with no annotations and no output schema, the description leaves important gaps: the meaning of the camera parameter and the behavioral implications of enabling continuous autofocus are unaddressed. It identifies the intent but is insufficient for reliable autonomous invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The only parameter 'camera' is an undocumented string with 0% schema description coverage, and the description never mentions it. An agent has no way to know what values are acceptable, what the parameter identifies, or whether it is safe to omit.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a clear action ('Enable continuous autofocus') on a specific resource, making the tool's intent immediately understandable. It distinguishes itself from obsbot_focus_manual by the word 'continuous', though it does not explicitly name or contrast itself with siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is provided about when to use this tool versus alternatives like obsbot_focus_manual or obsbot_focus_face. There is no context for when continuous autofocus is appropriate or why an agent would select this over other focus-related tools.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
obsbot_focus_faceC
Enable or disable face-priority autofocus.
| Name | Required | Description | Default |
|---|---|---|---|
| camera | No | ||
| enabled | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations are absent, so description carries full burden. It fails to disclose behavioral traits like toggling behavior, effects on other settings, or state requirements.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Single sentence is concise but lacks structure or additional context. It is front-loaded but too sparse.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema and no description of return values or side effects. For a boolean toggle, missing information about toggling behavior and camera parameter.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and description does not explain the 'camera' parameter (optional string). The 'enabled' parameter is self-explanatory, but the undocumented parameter is a gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Description clearly states the verb (enable/disable) and resource (face-priority autofocus), distinguishing it from siblings like obsbot_focus_manual.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance on when to use this tool vs alternatives (e.g., manual focus). No context on prerequisites or typical use cases.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
obsbot_focus_manualB
Set the focus motor to position (0-100, near→far), mapped onto the device range.
| Name | Required | Description | Default |
|---|---|---|---|
| camera | No | ||
| position | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description must carry full behavioral disclosure. It mentions mapping onto the device range but does not state if the action is instant, reversible, or requires specific device state. No side effects or dependencies are noted.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Single sentence, no filler, directly conveys the core action and range. Efficiently structured with no wasted words.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Despite low complexity (2 params, no output schema), the description omits clarification for the 'camera' parameter and lacks behavior details. Not sufficiently complete for an agent to use correctly without additional context.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%. The description only explains the 'position' parameter (range and mapping), but does not clarify the 'camera' parameter, which is a string with no format guidance. The agent may not know how to specify the camera.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action (set), resource (focus motor), and range (0-100, near to far), distinguishing it from sibling tools like obsbot_focus_face which is presumably auto/manual focus.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance on when to use this tool vs alternatives (e.g., obsbot_focus_face), no prerequisites or exclusions mentioned. The name implies manual control, but explicit advice is missing.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
obsbot_gimbal_moveA
Move the gimbal to an absolute yaw/pitch angle (degrees); positive yaw pans to the camera's left, positive pitch tilts down. Yaw is clamped to [-130,130], pitch to [-90,90]. Absolute positioning (1:1 degrees), verified on hardware.
| Name | Required | Description | Default |
|---|---|---|---|
| yaw | Yes | ||
| roll | No | ||
| pitch | Yes | ||
| camera | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does substantial work: it declares sign conventions (positive yaw = camera's left, positive pitch = down), the clamp ranges, the 1:1 degree mapping, and that behavior is hardware-verified. It omits whether motion is blocking, error behavior at limits, or the meaning of the roll/camera inputs, so it is not fully complete.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single dense sentence ordered from action to semantics to limits; every clause adds information and nothing is repeated or padded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 4-parameter tool with no annotations and no output schema, the description covers the two required motion parameters well but leaves roll and camera unexplained, so an agent cannot confidently drive the full input surface.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It does explain yaw and pitch semantics, including sign convention and clamp bounds, but leaves roll (defaulted) and camera completely undefined in both schema and description, leaving half the parameters undocumented.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (move) and resource (gimbal) plus the exact coordinate mode (absolute yaw/pitch in degrees). This clearly separates it from siblings like obsbot_gimbal_recenter (relative reset) and obsbot_gimbal_position (a read), which an agent can distinguish without opening a schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies when to use it (absolute positioning) but never states when to prefer a sibling such as recenter, or that this is absolute versus relative motion. Usage is inferable from the scope statement but not explicitly guided.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
obsbot_gimbal_positionA
Read the gimbal's current absolute yaw/pitch in degrees (positive yaw = camera's left, positive pitch = down) via the standard UVC Pan/Tilt controls. This is a live hardware readout accurate to ±1°, reported rounded to 2 decimal places (finer digits would be noise, not precision): it is valid during a move as well as after one, and reflects motion the host did not command (speed moves, recenter, tracking).
| Name | Required | Description | Default |
|---|---|---|---|
| camera | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden, and it excels: it discloses accuracy (±1°), rounding behavior, validity during moves, and that it reflects externally-initiated motion. This gives the agent a nuanced model of what the tool does beyond a simple 'getter.'
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences and front-loads the core purpose before diving into precision details. The parenthetical about noise versus precision is slightly verbose but still informative; overall it is efficiently organized.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema, yet the description never specifies the return format (e.g., field names or a JSON structure), leaving an agent uncertain about how to interpret the result. The undocumented 'camera' parameter further reduces completeness, so important information for invoking the tool correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema has one parameter ('camera', a string) with 0% description coverage, and the description never mentions the camera parameter at all. Since the schema provides no meaning beyond its type, and the description does not compensate, an agent has no idea what to pass for this parameter.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Read') and resource ('the gimbal's current absolute yaw/pitch in degrees'), with units and sign conventions. It is clearly distinguished from sibling tools like obsbot_gimbal_move and obsbot_gimbal_recenter, which perform actions rather than read state.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It implies usage for querying the current gimbal position and even notes it reflects motion from external actions like speed moves, recenter, and tracking, which helps an agent know it is useful for live monitoring. However, it never explicitly says when not to use it or contrasts it with move/recenter, so it falls just short of a fully explicit routing guideline.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
obsbot_gimbal_recenterA
Recenter the gimbal — drives it back to yaw 0 / pitch 0 (level and facing forward). Returns as soon as the command is sent: the gimbal may still be moving, so poll obsbot_gimbal_position if you need to know it has arrived.
| Name | Required | Description | Default |
|---|---|---|---|
| camera | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description discloses the async nature (command returns before movement finishes) and recommends polling. It does not mention permissions or side effects, but the disclosed behavior is sufficient for safe usage.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences, front-loaded with the primary action, and every word adds value. No redundancy or filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description explains the action, target orientation, and async behavior with a polling suggestion. However, it omits documentation of the 'camera' parameter, which is a notable gap. For a simple tool with no output schema, it is otherwise adequate.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema has one optional parameter 'camera' (type string) with no description. The tool description does not explain this parameter at all, leaving its purpose unclear. With 0% schema coverage, the description should compensate but fails to do so.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool recenters the gimbal to yaw 0/pitch 0 (level and facing forward). It uses a specific verb 'recenter' and identifies the resource, distinguishing it from sibling tools like obsbot_gimbal_move or presets.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explains that the command returns immediately and the gimbal may still be moving, advising to poll obsbot_gimbal_position for arrival. It gives clear context for when to use and suggests an alternative, though it does not explicitly state when not to use.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
obsbot_image_adjustA
Adjust a standard image control: control is brightness | contrast | hue | saturation | sharpness | gain | backlight-compensation; level 0-100 is mapped onto the device's supported range for that control. Standard UVC (IAMVideoProcAmp), no auto. NOTE: gain and backlight-compensation are NOT implemented on the Tiny 2 — it reports them as zero-length controls — so they are refused with an error rather than silently doing nothing. The other five work.
| Name | Required | Description | Default |
|---|---|---|---|
| level | Yes | ||
| camera | No | ||
| control | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Discloses standard UVC protocol, no auto, and device-specific failures for some controls. But does not mention persistence, atomicity, or prerequisites (e.g., camera must be active). With no annotations, this is adequate but not thorough.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four concise sentences. First sentence defines purpose, second adds context, third and fourth provide critical caveats. No redundancy, optimally front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Covers tool purpose, controls, range, mapping, and device-specific limitations. Lacks information about return values or persistence, but for a simple adjustment tool this is reasonably complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Adds significant meaning beyond schema: explains control enum values implicitly (listed), level range is mapped to device range, and notes that gain and backlight-compensation may fail. Camera parameter is not explained, but overall compensates for 0% schema coverage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Clearly states the tool adjusts a standard image control (brightness, contrast, etc.) with a level 0-100, listing the controls and specifying the mapping to the device's range. This distinguishes it from sibling tools like exposure controls.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Implicitly indicates manual adjustment (no auto) and provides device-specific guidance about unimplemented controls (gain, backlight-compensation on Tiny 2), warning they will error. However, lacks explicit when-to-use/alternatives guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
obsbot_image_exposure_autoA
Enable auto-exposure. Optional priority 'global' | 'face' selects the metering region (face-priority meters for a detected face). Uses the proprietary V3 frame protocol (CAM_SET_EXPOSURE_TINY2, which carries mode and value in one command) because the standard UVC/IAMCameraControl V4L2 path is a stub on the Tiny 2.
| Name | Required | Description | Default |
|---|---|---|---|
| camera | No | ||
| priority | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Describes the proprietary protocol and why standard paths are avoided, providing useful behavioral context. Without annotations, it carries the transparency burden well, though missing details on side effects or auth.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, front-loaded with the main function, no unnecessary words. Efficiently uses space for essential information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Provides good context for a simple tool, including protocol details. Lacks explanation of the 'camera' parameter, but overall adequate given no output schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Adds meaning to the 'priority' parameter by explaining the metering region options and face detection. The 'camera' parameter is not explained, but schema coverage is 0% and description partially compensates.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Clearly states the tool enables auto-exposure and describes the optional priority parameter for metering region, distinguishing itself from the sibling obsbot_image_exposure_manual.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Implies usage for auto-exposure scenarios via the sibling manual tool, but lacks explicit guidance on when to use this tool versus alternatives or prerequisites.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
obsbot_image_exposure_manualA
Set exposure level 0-100 (0 darkest → 100 brightest), mapped onto the device's exposure range. Also returns raw: the device-native exposure value the level mapped to, for diagnostics — level is the number to reason with. Uses the proprietary V3 frame protocol (CAM_SET_EXPOSURE_TINY2, which carries mode and value in one command) because the standard UVC/IAMCameraControl V4L2 path is a stub on the Tiny 2.
| Name | Required | Description | Default |
|---|---|---|---|
| level | No | ||
| camera | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Discloses level mapping and raw return value, and explains the proprietary protocol. However, missing side effects, persistence, error behavior, and prerequisites (e.g., camera must be connected).
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two concise sentences, front-loaded with the main action and mapping. No fluff.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Explains level and return, but misses explanation of camera parameter, error handling, and relationship to auto mode. Adequate but not comprehensive.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The 'level' parameter is well explained (0-100 mapping). The 'camera' parameter is not described at all, leaving ambiguity about its value and purpose.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool sets exposure manually with a normalized level (0-100) mapped to the device's range, and distinguishes it from automatic exposure and other image adjustments.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No explicit guidance on when to use this tool versus alternatives like auto exposure or general image adjust. The protocol note is technical but doesn't clarify usage context or prerequisites.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
obsbot_image_fovB
Set the field of view. fov: wide (86°) | medium (78°) | narrow (65°).
| Name | Required | Description | Default |
|---|---|---|---|
| fov | Yes | ||
| camera | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description must carry the burden. It discloses the angle values for each FOV option, which is useful behavioral info, but does not mention side effects, permissions, or response behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two short sentences with no unnecessary information. It is front-loaded with the purpose and fits within a single line.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the lack of annotations and output schema, and the presence of an undocumented 'camera' parameter, the description is insufficient for full understanding. It omits behavioral details and parameter explanation for one parameter.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%. The description adds meaning for the 'fov' parameter by listing angle values, but provides no info about the optional 'camera' parameter, leaving it undocumented.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool sets the field of view, with a specific verb and resource. It does not explicitly differentiate from siblings, but the function is unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description lists the available FOV options with degrees, giving implied usage context, but lacks explicit guidance on when to use this tool versus alternatives or any prerequisites.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
obsbot_image_hdrC
Toggle HDR/WDR imaging on or off.
| Name | Required | Description | Default |
|---|---|---|---|
| camera | No | ||
| enabled | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description must carry full behavioral transparency. It merely says 'toggle on or off' without disclosing side effects, required permissions, or what changes occur during the operation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is very short (5 words), which makes it concise but under-specified for a tool with parameters. It could be more informative without losing brevity.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's simplicity and the existence of sibling tools, the description lacks explanation of HDR/WDR concepts and the impact of toggling. It is insufficient for an agent to fully understand the operation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, and the description does not explain the 'camera' parameter or the 'enabled' boolean. The tool has 2 parameters, so it fails to add meaning beyond the raw schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses the specific verb 'Toggle' and clearly identifies the resource as 'HDR/WDR imaging'. It effectively distinguishes this tool from siblings by focusing on a unique imaging mode.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use HDR/WDR toggle versus other settings, nor does it mention prerequisites or alternatives. It is a bare statement without usage context.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
obsbot_image_wb_autoC
Enable auto white balance.
| Name | Required | Description | Default |
|---|---|---|---|
| camera | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure, and it only states the literal action without any context. It does not disclose that this overrides manual white balance settings, whether the setting persists, whether the camera must be awake/connected, or what happens if the request fails. The description adds no information beyond the tool name's own meaning.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The four-word single sentence is maximally concise and front-loaded with no wasted words. It is efficient, though it borders on being so minimal that it adds little beyond the machine-readable tool name itself.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool is low-complexity (one optional string parameter, no output schema, no nested objects), so the bar is lower. However, the description still omits any explanation of the 'camera' parameter and any preconditions, and with 32 siblings including a direct manual-counterpart tool, a one-line pointer to the manual alternative would meaningfully improve completeness.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and the single 'camera' parameter is documented only as a string with no meaning. The description does not mention the parameter at all, so the model cannot know whether 'camera' is an index, ID, or name, or how to obtain valid values (presumably from obsbot_devices). The description fails to compensate for the schema's lack of parameter documentation.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Enable') and resource ('auto white balance'), making the intended action immediately clear. It distinguishes implicitly from the sibling obsbot_image_wb_manual through the word 'auto', though it does not explicitly name the alternative or describe what the camera actually does when enabled.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is given about when to use this tool versus obsbot_image_wb_manual, obsbot_image_adjust, or obsbot_image_exposure_auto. The only usage signal is the action itself ('Enable auto white balance'), which essentially restates the purpose and offers no exclusions, prerequisites, or alternative-selection logic.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
obsbot_image_wb_manualB
Set white balance to a colour temperature in Kelvin (clamped to the device's supported range).
| Name | Required | Description | Default |
|---|---|---|---|
| camera | No | ||
| temperature | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description discloses a useful behavioral trait: the temperature is clamped to the device's supported range, which is beyond what the schema shows. However, with no annotations provided, the description still does not address prerequisites, failure behavior, or whether this affects the current white-balance mode.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single efficient sentence with no filler. It front-loads the core operation and includes the most important behavioral caveat about clamping.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple two-parameter setter, the description conveys the main operation and the temperature constraint well. But without annotations or an output schema, and with the 'camera' parameter undocumented, the description is only minimally complete for correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The description adds meaning to the 'temperature' parameter by specifying Kelvin units and clamping, improving on the bare schema. However, schema coverage is 0% and the 'camera' parameter is not explained at all, leaving a gap for agents that need to know what camera identifiers are accepted.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action: setting white balance to a Kelvin color temperature, which is specific and tied to the manual white-balance functionality. It does not explicitly differentiate from obsbot_image_wb_auto, but the Kelvin mention makes the manual mode intent fairly clear.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to use this tool versus the sibling obsbot_image_wb_auto or other image adjustment tools. The description does not mention alternatives or provide context such as 'use this when a fixed color temperature is needed.'
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
obsbot_preset_deleteA
Delete preset slot 1|2|3. The slot must be occupied. Verifies by re-reading the slot list after writing.
| Name | Required | Description | Default |
|---|---|---|---|
| slot | Yes | ||
| camera | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations, so description carries burden. Discloses verification step 're-reading after writing', but lacks details on side effects, permissions, or error handling.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, no redundancy, front-loaded with action and constraints.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Short description missing return value, error conditions (e.g., slot not occupied), and camera parameter usage. For a deletion tool, more context is needed.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema has 2 params (slot, camera) with 0% coverage. Description only mentions slot implicitly via 'slot 1|2|3', but fails to describe the camera parameter entirely.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Clearly states verb 'Delete', resource 'preset slot', and specific slots 1,2,3. Distinguishes from siblings like preset_save, preset_recall.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Implies usage by stating 'The slot must be occupied' as a precondition, but no explicit when-to-use or alternatives compared to sibling tools.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
obsbot_preset_listB
Read the three gimbal preset slots (occupied/empty, name, pose in degrees). Reads flat XU selectors 12 (list) and 13 (entry cursor), NOT the vendor V3 framed-reply path (which is non-functional for preset data on this device).
| Name | Required | Description | Default |
|---|---|---|---|
| camera | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries full burden. It discloses that it reads from specific XU selectors (12 and 13) and explicitly states that the vendor V3 path is non-functional. This is helpful behavioral context beyond just 'read presets'.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences and front-loads the core purpose. The second sentence adds valuable technical context without being overly verbose. Slight reduction in conciseness due to technical jargon, but still efficient.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description gives partial output details (occupied/empty, name, pose) but no information on data structure, defaults, error states, or behavior when no presets exist. Missing output schema increases the need for such details, which are absent.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, and the description provides no explanation of the 'camera' parameter. The agent has no guidance on what value to provide or if it's required, leaving the parameter's semantics entirely opaque.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action 'Read' and the resource 'three gimbal preset slots', specifying exactly what data is returned (occupied/empty, name, pose in degrees). This distinguishes it from sibling preset tools that modify or recall presets.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No explicit guidance on when to use this tool versus alternatives like obsbot_preset_recall or obsbot_preset_save. The description only explains internal implementation details, not usage context.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
obsbot_preset_recallA
Recall preset slot 1|2|3, driving the gimbal to that slot's saved pose. The slot must be occupied. The gimbal may still be moving when this returns — verification only confirms the slot is still occupied, not that the pose has arrived.
| Name | Required | Description | Default |
|---|---|---|---|
| slot | Yes | ||
| camera | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations provided; description fully covers the asynchronous nature, verification limitation, and precondition. No contradictions.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, front-loaded with purpose, every sentence adds value. No fluff.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Explains core function and key behavior but lacks explanation of the 'camera' parameter and does not describe error handling (e.g., slot not occupied). Adequate but with gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema has 0% description coverage. Description explains the 'slot' parameter (values 1-3) but completely omits the optional 'camera' parameter, leaving its purpose unclear.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Clearly states the tool recalls a preset slot (1,2,3) to drive the gimbal. Distinguishes from siblings like save/delete by its action.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly mentions precondition (slot must be occupied) and asynchronous behavior (gimbal may still move). Does not explicitly list when to use versus alternatives, but context is sufficient.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
obsbot_preset_renameA
Rename preset slot 1|2|3. The slot must already be occupied. Names longer than 40 bytes are truncated to fit the wire frame. Verifies by re-reading the slot list after writing.
| Name | Required | Description | Default |
|---|---|---|---|
| name | Yes | ||
| slot | Yes | ||
| camera | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Discloses truncation of names longer than 40 bytes and a verification step by re-reading. With no annotations, this is good behavioral context beyond the schema.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three concise sentences that front-load the purpose, then provide important details. No wasted words.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Covers key behavioral aspects but lacks explanation of the camera parameter and return value. No output schema, so description should be more comprehensive.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Explains slot and name truncation but does not describe the camera parameter, which is present in the schema. Schema coverage is 0%, so description should cover all parameters.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Clearly states the action 'rename' and the resource 'preset slot 1|2|3', distinguishing it from sibling tools like delete or save. Includes a precondition that the slot must be occupied.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides a precondition (slot must be occupied) but does not explicitly guide when to use rename versus update or save siblings, nor when not to use it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
obsbot_preset_saveA
Save the gimbal's current live pose (yaw/pitch, via the standard UVC Pan/Tilt controls) into preset slot 1|2|3. Slots are create-once on this device — there is no overwrite, so an occupied slot is rejected (delete it first). Verifies by re-reading the slot list after writing.
| Name | Required | Description | Default |
|---|---|---|---|
| slot | Yes | ||
| camera | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description bears full burden. It discloses key behaviors: no overwrite, rejection of occupied slots, and verification by re-reading the slot list. This is thorough but could mention permissions or side effects.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two concise sentences with no redundancy. The first sentence states the action, the second adds constraints and verification. Information is front-loaded and efficient.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Description covers the core functionality but lacks explanation of the 'camera' parameter and does not describe the output/return value (no output schema). While verification step is mentioned, overall completeness is adequate but not thorough.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema has 0% description coverage. Description explains the 'slot' parameter (enum values 1-3) but does not explain the 'camera' parameter at all, leaving its purpose ambiguous. This is insufficient given the lack of schema descriptions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Description clearly states it saves the gimbal's current live pose to a preset slot (1|2|3). It differentiates from sibling preset tools (recall, delete, rename, update, list) by focusing on saving a new pose.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Describes the create-once behavior and the need to delete an occupied slot before saving, providing implicit guidance to use the delete tool. However, it does not explicitly mention when to use the update tool or contrast with alternatives.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
obsbot_preset_updateA
Overwrite preset slot 1|2|3 with the gimbal's current live pose (yaw/pitch via the standard UVC Pan/Tilt controls). The slot must already be occupied (save first to create it). Verifies by re-reading the slot list after writing.
| Name | Required | Description | Default |
|---|---|---|---|
| slot | Yes | ||
| camera | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Discloses that the tool overwrites (destructive action), requires pre-existing slot, and performs verification by re-reading after writing. No annotations present so the description carries full burden; good coverage of behavioral traits.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three concise sentences with no redundancy. First sentence states purpose, second adds prerequisite, third adds verification. Front-loaded and efficient.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given 2 parameters, no output schema, and no annotations, the description covers key aspects: overwrite behavior, prerequisite, and verification. Missing details on error handling or return value, but sufficient for a simple update tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 0% schema description coverage, the description adds value by explaining that 'slot' accepts values 1, 2, or 3 and corresponds to preset slots. However, the 'camera' parameter is not elaborated, leaving a gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action ('overwrite'), the resource ('preset slot 1|2|3'), and the target state ('current live pose'). It distinguishes from sibling tools like obsbot_preset_save (which creates) and obsbot_preset_recall (which reads).
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly states prerequisite: 'the slot must already be occupied (save first to create it)'. Also mentions verification step. No direct comparison to alternatives, but context from sibling names provides differentiation.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
obsbot_sleepA
Sleep the camera/gimbal (sends "sleep"). This MOVES the camera: sleeping STOWS the gimbal, tilting it face-down to roughly pitch 84°, so obsbot_gimbal_position will read ~84 rather than the pose you left it in. obsbot_wake un-stows it.
| Name | Required | Description | Default |
|---|---|---|---|
| camera | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and discloses the critical non-obvious side effect: sleeping moves the camera by stowing the gimbal to pitch ~84, which will change obsbot_gimbal_position. It also explains that wake reverses the action, so the agent knows the state impact.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and front-loaded: it states the action first, then the movement consequence, then the inverse. Every sentence adds relevant information with no repetition or filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple one-optional-parameter command with no output schema, the description covers purpose, side effects, and state consequences well. It falls just short of complete because the parameter's accepted values are undocumented and there is no explicit statement about the expected return/status, but these are minor for this tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0% and the description never mentions the single 'camera' parameter, so the agent must infer that it identifies which camera/gimbal to sleep. It does not state whether this expects a device ID, name, or index, though the parameter's name and optionality give a weak clue.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb and resource ('Sleep the camera/gimbal') and immediately clarifies what sleeping means by noting the gimbal is stowed face-down at roughly pitch 84. It also names the inverse sibling (obsbot_wake), so the tool is easily distinguished from wake and other gimbal commands.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It is clear that sleep is the command to stow the gimbal, and the mention that obsbot_wake un-stows it gives the agent the counterpart. It does not explicitly state exclusions or conditions (e.g., when to prefer gimbal_move or recenter), but the usage context is implied strongly enough.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
obsbot_statusA
Read the camera's live status block. Returns { awake, hdr, faceAe, aiMode, trackSpeed, fovMode, zoomPercent, focusMode, focusPosition }: faceAe is whether auto-exposure is metering for a detected face; aiMode is the current AI framing (no-tracking|normal|upper-body|close-up|headless|lower-body|desk|whiteboard|hand|group|unknown); trackSpeed is standard|sport|unknown; fovMode is the field-of-view mode (wide|medium|narrow|custom|unknown), where custom means a continuous zoom overrode the discrete modes; zoomPercent is the zoom position, 0-100; focusMode is auto|manual|unknown. focusPosition is present ONLY in manual mode, on the same 0-100 scale obsbot_focus_manual takes: under autofocus this camera does not expose the motor, it echoes the last written value, so reporting it would look like a live focus distance while being stale. Focus is a standard UVC control rather than a field of the status block, so it costs an extra read; a device that cannot answer it reports focusMode unknown rather than failing the whole read. Under --debug the result also carries raw: the full 60-byte status block as hex (for reverse-engineering undecoded offsets).
| Name | Required | Description | Default |
|---|---|---|---|
| camera | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Despite no annotations, the description fully discloses all behavioral traits: fields present only under certain conditions (focusPosition only in manual mode), edge cases (unknown values if device can't answer), and extra debug output. No contradictions.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with the main purpose but then becomes verbose with detailed field-by-field definitions. While informative, it could be shortened without losing critical information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description fully specifies the return value structure, including conditional fields, edge cases, and debug output. No important context is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The sole parameter 'camera' is not described in the text; the description focuses entirely on output. With 0% schema coverage, the description should explain what the parameter means, but it does not. This is a significant gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states it reads the camera's live status block and lists returned fields. It is a specific verb-resource pair and is distinct from sibling tools like obsbot_focus_manual (write) or obsbot_capture_snapshot (action).
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implicitly makes clear this is for reading current status, but does not explicitly contrast with alternatives or state when to avoid using it. The context is clear enough for an agent to infer usage.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
obsbot_tail2_af_trackB
Read or set the autofocus track mode: global, face or front (Center's radios; face measured on the wire, the other two are the presumed strings). Readback-verified write.
| Name | Required | Description | Default |
|---|---|---|---|
| mode | No | ||
| camera | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It does disclose that writes are readback-verified and that only 'face' is measured on the wire while the others are presumed, but it omits permissions, side effects, persistence, and error behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single compact sentence, front-loaded with the core action and mode list. The parenthetical about 'Center's radios' is cryptic but does not add excessive length.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given no annotations, no output schema, and a two-parameter tool, the description provides some useful behavior (readback-verified write) and mode details. However, it leaves the camera parameter, return values, and selection among sibling tools unspecified, so it is not fully complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It names the three mode values and adds reliability context for them, but the 'camera' parameter is entirely unexplained, leaving a significant gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: read or set the autofocus track mode, and enumerates the three mode values. It does not differentiate this tool from siblings like obsbot_ai_track or obsbot_focus_face, so it falls short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Offers no when-to-use guidance, prerequisites, or alternatives among the many sibling focus and tracking tools. The 'read or set' phrasing implies basic usage but leaves selection entirely to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
obsbot_tail2_ai_trackA
Enable or disable a Tail 2's AI tracking. enabled:false stops tracking (mode none). enabled:true starts the chosen mode — default humanTrackingSingleMode (single-person framing, the camera's power-on default); also group human framing, animal tracking, and object tracking (the mode strings are the camera's own enum). NOTE: while tracking is active it owns the gimbal — manual moves fight it, and a recenter can drop tracking to none. Mode strings verified against firmware 7.2.13.1.
| Name | Required | Description | Default |
|---|---|---|---|
| mode | No | humanTrackingSingleMode | |
| camera | No | ||
| enabled | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations at all, the description carries the full burden and delivers: it warns that active tracking owns the gimbal, that manual moves fight it, that a recenter can drop tracking to none, and that mode strings were verified against firmware 7.2.13.1. This is exactly the kind of stateful side-effect disclosure an agent needs for a mutation tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the primary action and the disabled-case first, then the enabled/mode detail, then a NOTE block for the operational caveat. Dense but every clause adds information; the firmware-verification sentence is useful provenance rather than filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema and no annotations, so the description must stand alone; it covers on/off behavior, mode selection, side effects, and firmware provenance well. The only material gap is the unexplained 'camera' parameter, which matters if multiple devices are targeted.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It explains 'enabled' and the 'mode' default and categorizes the enum families (single/group human, animal, object), but it never explains the 'camera' parameter at all, leaving one of three params undocumented in both schema and description.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Starts with a concrete verb+resource pair ('Enable or disable a Tail 2's AI tracking') and immediately contrasts the enabled:true/false semantics. The note about the gimbal being owned by tracking implicitly separates this tool from obsbot_tail2_recenter and manual gimbal moves.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Clearly states what enabled:false does ('stops tracking, mode none') and what enabled:true does, including the default mode and the available mode families. It lacks an explicit 'use X instead when...' routing to siblings like obsbot_tail2_track_target, but the recenter interaction gives practical context.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
obsbot_tail2_antiflickerB
Read or set the anti-flicker mode: off, 50hz or 60hz (60hz measured on this unit). Readback-verified write.
| Name | Required | Description | Default |
|---|---|---|---|
| mode | No | ||
| camera | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It does add a genuinely useful behavioral trait – 'Readback-verified write' tells the agent the write is confirmed by a read-back – but it omits permission requirements, failure behavior, and what a read returns.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single dense sentence with the action front-loaded and no filler; the enum caveat is parenthetical and does not derail the reading. Slightly terse for the load it must carry, but structurally sound.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 2-param tool with no output schema and no annotations, the description covers the core action and the read-back verification but leaves gaps: it does not explain what a read returns, what the 'camera' parameter designates, or how failures surface.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% for 2 parameters, so the description must compensate. It reinforces the enum values and adds the caveat that 60hz was measured on this unit, but the 'camera' parameter is completely undocumented and the read/write semantics of omitting 'mode' are left to inference.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: read or set the anti-flicker mode, and enumerates the valid values (off, 50hz, 60hz). It is clear what the tool touches; the only reason it is not a 5 is that it gives no sibling differentiation, though no sibling covers anti-flicker.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'Read or set' implies the usage contract (omit mode to read, supply mode to write), and the required-parameter count of 0 supports that reading, but the description never states the when-to-use condition or any prerequisites explicitly.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
obsbot_tail2_audioA
Read (bare call) or set the Tail 2's audio input: volume 0-100 and mute. Changes apply to the encoded/recorded audio stream.
| Name | Required | Description | Default |
|---|---|---|---|
| mute | No | ||
| camera | No | ||
| volume | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It usefully discloses the effect scope ('changes apply to the encoded/recorded audio stream'), but says nothing about permissions, whether changes persist across sessions, or what a read returns.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences, front-loaded with the read/write distinction, no filler or restatement of the name.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The core contract (bare read vs. set, stream scope, volume range) is present, but with no annotations, no output schema, and 0% schema coverage, the unexplained 'camera' selector and the shape of read-mode results are real gaps an agent would hit.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It covers volume (and the 0-100 range, which the schema's minimum/maximum already encodes) and mute, but the 'camera' parameter is never mentioned or explained, leaving a full third of the surface undocumented.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a concrete verb pair (read/set) and resource (Tail 2 audio input), plus the two controllable fields (volume, mute). No audio sibling exists in the tool list, so disambiguation is implicit rather than stated, which keeps it just short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly explains the dual invocation mode: a bare call reads state, supplying values sets it. That is genuine when-to-use guidance. It does not name an alternative tool, but none competes for this capability.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
obsbot_tail2_auto_zoom_speedA
Set the AUTO zoom speed 1-10 — how fast AI-tracking auto-zoom framing moves (Center's Console slider; distinct from the MANUAL zoom speed on obsbot_tail2_zoom's speed param). Rides UDP 9999 key 0x17. No REST-readable state; no ack — the frame is synthesized and works with Center closed.
| Name | Required | Description | Default |
|---|---|---|---|
| speed | Yes | ||
| camera | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full disclosure burden and does so substantially: it reveals the transport (UDP 9999, key 0x17), that there is no REST-readable state, no ack, and that the frame is synthesized and works even with Center closed. Only persistence/volatility of the setting across restarts is left unstated.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The value, its range, and its purpose are front-loaded in the first clause, with the sibling distinction and transport details following. It is dense with useful specifics, though the parenthetical asides make it slightly packed.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a raw UDP frame-control tool with no output schema, the description supplies the behavioral context an agent needs (transport, no-ack, no readable state). The only notable gap is the undocumented camera parameter.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must compensate. It explains what speed controls and its range (though the 1-10 bounds already appear in the schema), but the camera parameter is left completely undocumented in both schema and description. Partial compensation justifies the baseline 3.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Set) and resource (AUTO zoom speed), gives the range 1-10, and clarifies what the value controls (how fast AI-tracking auto-zoom framing moves). It explicitly distinguishes itself from the MANUAL speed on obsbot_tail2_zoom, so an agent can tell the two apart without opening a schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description clearly differentiates this tool from its closest sibling (obsbot_tail2_zoom's speed param) and notes it works with Center closed, which frames usage context well. It does not, however, spell out an explicit when-to-use scenario or any prerequisites, so it stops short of a full 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
obsbot_tail2_capture_photoB
Trigger a still-photo capture on the Tail 2 (POST capture/trigger). Needs storage (SD card — check obsbot_tail2_status's sdcard field). The ack means the camera accepted the trigger; the photo lands on the card (see obsbot_tail2_status / album).
| Name | Required | Description | Default |
|---|---|---|---|
| camera | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden and delivers real value: it explains that the ack means the camera only *accepted* the trigger (not that the photo exists) and that the file lands on the card asynchronously. It omits permissions, rate limits, or timeout behavior, but the ack semantics are a meaningful disclosure beyond structured data.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight sentences, front-loaded with the primary action, then precondition, then result semantics. No filler, though the parenthetical endpoint note is slightly dispensable.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Covers purpose, storage precondition, and async result semantics, and there is no output schema to compensate for. But it leaves the 'camera' parameter undefined and does not differentiate from snapshot siblings, so it is adequate rather than complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
There is one parameter ('camera') with 0% schema description coverage, and the description never mentions it, so it adds no semantic meaning. Since a lone, unnamed 'camera' string may still be inferable, this is above the floor but well short of compensating for the coverage gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Trigger a still-photo capture on the Tail 2') and even names the underlying endpoint (POST capture/trigger), so the action is unambiguous. However, it never distinguishes this from siblings such as obsbot_tail2_snapshot or obsbot_capture_snapshot, leaving the agent to guess which photo tool applies.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides a genuine precondition ('Needs storage (SD card — check obsbot_tail2_status's sdcard field)') that steers the agent to verify storage first. But it offers no when-to-use vs. alternatives guidance against the numerous snapshot siblings, so the routing decision remains implied.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
obsbot_tail2_devicesA
List registered OBSBOT Tail 2 cameras (network cameras; MAC-keyed). Each entry: mac (the identity — pass it as camera to any obsbot_tail2_* tool), name, and the host(s) it has answered on. Registration is not liveness: an offline camera still appears here; use obsbot_tail2_status to check one is reachable.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the behavioral burden and does so well: it discloses the critical caveat that offline cameras still appear, and that the MAC is the identity to pass as `camera` to other obsbot_tail2_* tools. It does not discuss ordering, caching, or error behavior, keeping it just short of 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight sentences, front-loaded with the identity/scope claim, then the return shape, then the routing caveat. No filler or restated name.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, so the description rightly names the returned fields (mac, name, hosts); combined with the liveness caveat and cross-tool `camera` usage, an agent has enough to call and interpret it. Minor gaps (ordering, empty-list behavior) remain unstated.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
There are zero input parameters, so the baseline is 4; nothing needs documenting. The description instead adds meaning to the output fields (mac as identity, name, host answers), which is useful but not strictly parameter semantics.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('List registered OBSBOT Tail 2 cameras') and immediately scopes it as network cameras keyed by MAC. It is clearly distinguishable from siblings like obsbot_tail2_status (liveness) and obsbot_devices (presumably other device families).
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly names the alternative and its selection condition: 'Registration is not liveness... use obsbot_tail2_status to check one is reachable.' It tells the agent both when this tool suffices and when to switch tools.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
obsbot_tail2_export_logA
Download the camera's full diagnostic bundle (GET /camera/test/log — the same archive OBSBOT Center's Export Log button produces) and save it to path. ~12 MB tar.gz containing the settings tree (ust.json), the runtime status snapshot, factory records, kernel logs. Takes ~10-20 s (server-side archive generation).
| Name | Required | Description | Default |
|---|---|---|---|
| path | Yes | ||
| camera | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and largely meets it: it discloses output size (~12 MB tar.gz), exact contents (ust.json settings tree, runtime status snapshot, factory records, kernel logs), and server-side latency (~10-20 s). It stops short of stating auth/permission needs or failure behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single dense sentence plus one follow-up, front-loaded with the action, endpoint, and destructive-free save semantics. Every clause (size, contents, duration) adds actionable information; no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema and no annotations, so the description does its own heavy lifting on return artifact, contents, and timing. It omits what `path` points to (file vs directory), error/failure behavior, and the `camera` param, leaving small but real gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must compensate. It clarifies `path` as the save destination for the downloaded archive, but says nothing about the optional `camera` parameter (presumably a device selector), leaving one of two params semantically dark.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (download/save) and resource (full diagnostic bundle / log archive), names the underlying endpoint GET /camera/test/log, and equates it to OBSBOT Center's Export Log button. No sibling tool produces a log bundle, so the agent can place it unambiguously.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied by the diagnostic-bundle framing (use this when troubleshooting), but there is no explicit when-to-use/when-not statement and no alternative tool is named or excluded. Adequate but leaves the agent to infer the trigger condition.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
obsbot_tail2_exposureA
Read (bare call) or set the Tail 2's exposure. mode manual|auto; in AUTO: face (face-priority AE, global|face) and evbias (-3.0..3.0, ~0.3 steps); in MANUAL: iso (100-6400) and shutter ("1/N", 1/6400..1/30). Manual values only exist in manual mode and face/evbias only in auto — a bare call reports whichever set the current mode exposes. Note: the camera's manual ISO granularity is finer than the doc claims (a live-AE value like 894 reads back).
| Name | Required | Description | Default |
|---|---|---|---|
| iso | No | ||
| face | No | ||
| mode | No | ||
| camera | No | ||
| evbias | No | ||
| shutter | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden, and it does disclose meaningful behavior: bare call reads whichever set the current mode exposes, manual values exist only in manual mode, face/evbias only in auto, plus an edge-case quirk about finer-than-documented ISO granularity. It still omits error/conflict behavior (e.g., passing iso while in auto) and auth requirements.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the core read/set behavior, then organizes the mode-specific detail efficiently. Dense but every clause adds usable information; the trailing ISO-granularity note is a worthwhile edge-case, not filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no annotations, no output schema, and 6 parameters, the description provides enough to call the tool correctly across both modes. The unnamed 'camera' parameter and unstated conflict handling are the only notable omissions.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate, and it documents 5 of 6 parameters with real semantics: mode (manual|auto), face (face-priority AE global|face), evbias range and step, iso range, and shutter range/format. Only the 'camera' parameter is left undocumented, which is a modest gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource: read (bare call) or set the Tail 2's exposure. Clear what the tool does, though it does not explicitly distinguish itself from the sibling obsbot_image_exposure_auto / obsbot_image_exposure_manual tools, so sibling differentiation is only implicit via the device prefix.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Clearly explains the two invocation modes (bare call to read, params to set) and the mode-dependent conditions under which parameters apply. It does not, however, name alternatives to route to or state when-not-to-use conditions, leaving sibling selection to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
obsbot_tail2_focusA
Read (bare call) or set the Tail 2's focus: mode afc|afs|mf, and position (0-100, the focus motor) which is only valid in mf — the camera errors on it in afc/afs and this tool refuses rather than trigger that. Bare call reports the mode, and the position when in mf.
| Name | Required | Description | Default |
|---|---|---|---|
| mode | No | ||
| camera | No | ||
| position | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does meaningful work: it discloses the dual read/write contract, that the camera errors on position outside mf, and that the tool proactively refuses that invalid combination. It does not mention permission requirements, device-state prerequisites, or what happens on other failure paths.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two dense sentences, no filler, with the read/write contract front-loaded and the constraint on position following immediately. Every clause carries operational information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with no annotations, no output schema, and three parameters, the description covers the return value of a bare call ('reports the mode, and the position when in mf') and the key validity constraint, which is enough to invoke it correctly. The undocumented 'camera' parameter and unspecified non-guard error behavior are minor omissions.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate, and it does for two of three parameters: it restates the mode enum and, more importantly, supplies the cross-parameter constraint that position is only meaningful in mf and rejected otherwise. The 'camera' parameter is never mentioned in either schema or description, which is the remaining gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description gives a specific verb and resource ('Read (bare call) or set the Tail 2's focus') and immediately enumerates the controllable fields (mode, position), so the agent knows exactly what the tool touches. It does not, however, distinguish itself from near siblings such as obsbot_tail2_focus_point, obsbot_focus_auto, or obsbot_focus_manual, leaving that routing to inference.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It clearly states the read path (bare call) versus the write path and gives a hard conditional: position is only valid in mf, and the tool refuses rather than letting the camera error. That is concrete when-to-use guidance. It stops short of naming an alternative tool for cases outside its scope.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
obsbot_tail2_focus_pointA
Tap-to-focus: move the focus window to a normalized frame coordinate (x/y 0.01-0.99, top-left ~0,0) AND start a point focus there. Effective in afc/afs focus modes only — in mf the motor position is the tool (obsbot_tail2_focus). Bare call reads the current window. Divide a snapshot pixel by frameWidth/frameHeight and pass it straight in.
| Name | Required | Description | Default |
|---|---|---|---|
| x | No | ||
| y | No | ||
| camera | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the behavioral burden, and it delivers real traits beyond the schema: the focus-mode dependency (only effective in afc/afs), the fallback tool for mf, and the read-write duality of a bare vs. populated call. It does not cover permissions, error behavior, or the return payload, so it stops short of a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the purpose and packs mode constraints and coordinate semantics into a few dense sentences with no filler. It is information-dense rather than padded, though the coordinate-conversion tail makes it slightly heavy for a single read.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no annotations and no output schema, the description covers purpose, mode gating, alternatives, and coordinate semantics well, plus the bare-call read behavior. The only meaningful gap is the camera parameter's meaning, but overall it is sufficient to call the tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must compensate, and it does for x/y: normalized 0.01-0.99, top-left at ~0,0, and a concrete conversion recipe (divide snapshot pixel by frameWidth/frameHeight). The camera parameter is left unexplained, so it is not fully complete.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a concrete verb and resource: 'move the focus window to a normalized frame coordinate AND start a point focus there.' It goes further by distinguishing itself from the sibling obsbot_tail2_focus, which is used in mf mode, so an agent can route correctly without opening either schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicit about when it applies (afc/afs focus modes only), when to use the alternative (in mf the motor position is the tool obsbot_tail2_focus), and the no-arg behavior (bare call reads the current window). All three selection conditions are spelled out.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
obsbot_tail2_gestureA
Read (bare) or set the gesture-control switches and the gesture zoom factor: lockedTarget, recording, zoom (booleans) and zoomFactor (the multiplier one zoom gesture applies, e.g. 2.0). All measured on the wire 2026-10-01. Writes are readback-verified; each field is optional.
| Name | Required | Description | Default |
|---|---|---|---|
| zoom | No | ||
| camera | No | ||
| recording | No | ||
| zoomFactor | No | ||
| lockedTarget | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full load; it does disclose the read/write duality, that writes are readback-verified, and the provenance of the measurements. It omits conflict/precedence rules between zoom and zoomFactor, required permissions, and what happens on a failed write.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight sentences: the read/set mode and affected fields come first, then a concrete example value, then the optionality and verification note. No filler, and the provenance tag is compact.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 5-parameter, zero-coverage, annotation-free tool with no output schema, the description covers most of what an agent needs but leaves the `camera` parameter undocumented and gives no defaults or failure behavior. Adequate but with clear gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must compensate, and it defines four of five parameters including the meaning of zoomFactor ("multiplier one zoom gesture applies, e.g. 2.0"), which the schema only bounds numerically. However the `camera` parameter is left entirely unexplained in both schema and description.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb pair (read/set) and a specific resource (gesture-control switches and gesture zoom factor), with the individual fields named. It is clearly distinguishable from generic zoom siblings like obsbot_tail2_zoom or obsbot_tail2_hybrid_zoom, though it never explicitly contrasts itself with them.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
"Read (bare) or set..." implies that calling with no arguments reads and passing fields writes, and "each field is optional" reinforces that. There is no explicit statement of when to prefer this over sibling gesture/zoom tools, nor any prerequisites such as device wake state.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
obsbot_tail2_gimbal_invertA
Read or set the Tail 2's control-direction inversion (gimbalinvert). Call with no arguments to read; pass enable to write (verified by readback).
| Name | Required | Description | Default |
|---|---|---|---|
| camera | No | ||
| enable | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations exist, so the description carries the full burden. It usefully discloses that writes are verified by readback, but says nothing about persistence of the setting, permission/auth needs, or how the value is returned, leaving meaningful behavioral gaps for a mutation-capable tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single sentence that front-loads the operation and then the read/write branch; no filler whatsoever.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a two-parameter, no-output-schema, no-annotation tool, the description covers the read/write mode logic and readback confirmation well; the only notable omission is the camera selector's meaning.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0% and there are 2 parameters, so the description must compensate. It explains the semantics of enable (pass it to write) but says nothing about the camera parameter, leaving half the inputs undocumented in both schema and prose.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb pair (read/set) and the exact resource (Tail 2 control-direction inversion, gimbalinvert), which is clearly distinct from siblings like gimbal_roll_bias or gimbal_speed.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly tells the agent how to select each mode: no arguments to read, pass enable to write. That covers the when-to-use decision for a dual-mode tool, though it names no alternative tool or prerequisites.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
obsbot_tail2_gimbal_moveA
Move the Tail 2's gimbal to an absolute yaw/pitch in degrees — the Tail 2's pose-record convention (obsbot_tail2_gimbal_position reports the same axes). There is no move-to-angle endpoint in the API, so this is a closed loop over the speed primitive: jog, read the pose (each read ~2s), correct, up to 5 rounds, tolerance ±1.5°. Slow by design — expect ~5-15s. REFUSES while AI tracking is active (tracking drives the gimbal itself and would fight the loop); obsbot_tail2_ai_track enabled:false first.
| Name | Required | Description | Default |
|---|---|---|---|
| yaw | No | ||
| pitch | No | ||
| camera | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations present, the description carries the full burden and does so excellently: it discloses the closed-loop implementation (jog, read pose ~2s, correct, up to 5 rounds, tolerance ±1.5°), expected duration (~5-15s), and refusal condition with the exact prerequisite (disable AI tracking first). This gives an agent rich behavioral context beyond the schema.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with the core action, then systematically adds implementation detail, performance expectations, and prerequisites. Every sentence earns its place; there is no redundant or filler content despite the length.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a complex motion tool with no annotations, no output schema, and 0% schema coverage, the description provides strong contextual completeness: it explains behavior, timing, tolerance, and the critical AI-tracking prerequisite. The notable gap is the unmentioned 'camera' parameter and optionality, which leaves one of three parameters entirely unexplained.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% for all 3 parameters. The description adds meaning for yaw and pitch (absolute degrees, pose-record convention, matching gimbal_position axes) beyond the raw numeric ranges in the schema. However, it completely omits the 'camera' parameter and does not address requiredness (0 required parameters), so it only partially compensates for the lack of schema documentation.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: 'Move the Tail 2's gimbal to an absolute yaw/pitch in degrees.' It distinguishes from siblings by clarifying the pose-record convention and that there is no move-to-angle endpoint, implying a closed-loop approach instead of the speed primitive. An agent can tell what this does and how it differs from gimbal_speed.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides clear context: use for absolute yaw/pitch moves, and explicitly says it REFUSES while AI tracking is active, requiring obsbot_tail2_ai_track enabled:false first. However, it does not name an alternative sibling tool (e.g., gimbal_speed for relative moves) to use instead, so it stops short of full when-to-use-vs-alternative guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
obsbot_tail2_gimbal_positionA
Read the Tail 2's live gimbal pose in degrees (yaw/pitch/roll) plus the zoom ratio. The API has no direct pose read — this works by saving the current pose into an empty preset slot, reading it back, and deleting the slot (takes ~2s). Requires at least one empty preset slot: overwriting a preset is irreversible (no pose-by-value write exists to restore it). Leftover 'pose-probe' slots from crashed reads are deleted, not trusted.
| Name | Required | Description | Default |
|---|---|---|---|
| camera | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does so thoroughly: it discloses the non-obvious implementation (save → read → delete preset), the ~2s latency, the irreversible-overwrite hazard for an occupied slot, and that stale 'pose-probe' slots are cleaned rather than trusted. These are exactly the side effects and risks an agent needs before invoking.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three dense sentences, front-loaded with the returned values, then the mechanism, then the constraint. No sentence is filler — the implementation detail is what justifies the constraint that follows.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, and the description compensates by naming the return fields and units (degrees; zoom ratio). It also covers the side-effect profile and latency, so an agent can call it correctly and know what to expect.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
There is one parameter ('camera') with 0% schema description coverage, and the description never mentions it — not its format, whether it names a device, or whether it is optional. The baseline-4 exemption applies to zero-parameter tools, and this exceeds that; the description fails to compensate for the coverage gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('read the Tail 2's live gimbal pose in degrees (yaw/pitch/roll) plus the zoom ratio'), which is distinct from the sibling gimbal-move/preset tools. An agent knows exactly what it gets back without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explains the context that forces this tool's use ('the API has no direct pose read') and states the precondition (at least one empty preset slot). It does not explicitly name a sibling alternative or a when-not condition, but the usage context is clear.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
obsbot_tail2_gimbal_speedA
Jog the Tail 2's gimbal: per-axis SPEED (sign = direction, magnitude = speed, ±150 of the firmware's ±178 range) with an automatic stop after durationMs (default 500, max 5000). This is the joystick primitive (POST ptz/gimbalcontrol) — there is no absolute move endpoint, so the camera moves continuously until stopped; this tool always stops it. Measured: command 20 ≈ 7.9°/s, and a POSITIVE yaw command DECREASES the recorded yaw (obsbot_tail2_gimbal_position's convention). Disable AI tracking first or it will fight the move.
| Name | Required | Description | Default |
|---|---|---|---|
| yaw | No | ||
| roll | No | ||
| pitch | No | ||
| camera | No | ||
| durationMs | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden and largely meets it: it discloses that motion is continuous until stopped, that this tool always stops it, the auto-stop duration (default 500ms, max 5000ms), the ±150 vs ±178 firmware range, the empirical calibration (~20 command ≈ 7.9°/s), and the counter-intuitive sign convention where positive yaw decreases recorded yaw. That is unusually rich behavioral context.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
It is dense but front-loaded: the core behavior and defaults come first, then the API primitive, then the calibration and the tracking caveat. Every sentence carries information; the measured °/s figure and the yaw convention are the only slightly extraneous extras, but both are actionable.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no annotations, no output schema, and no schema descriptions, the description supplies nearly everything an agent needs to call this safely and correctly. The one gap is the undocumented 'camera' selector, which an agent cannot disambiguate from the description alone.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate and does so well for four of five parameters: sign-as-direction and magnitude-as-speed for yaw/pitch/roll, the ±150 cap, and the durationMs default (500) that does not even appear in the schema. It says nothing about the 'camera' parameter, leaving one required meaning undocumented.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('jog the Tail 2's gimbal') and immediately qualifies the semantics as velocity-based rather than absolute, naming the underlying primitive (POST ptz/gimbalcontrol). It distinguishes itself from siblings by explicitly noting there is no absolute move endpoint, so an agent can tell it apart from obsbot_tail2_gimbal_move/recenter without opening either schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives a clear prerequisite — 'Disable AI tracking first or it will fight the move' — and explains the continuous-motion model with the automatic stop. However, it never names the alternative sibling tools (obsbot_tail2_gimbal_move, obsbot_tail2_recenter) or states when a discrete move is preferable to a jog.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
obsbot_tail2_hdrB
Read (bare call) or set (enabled) the Tail 2's HDR.
| Name | Required | Description | Default |
|---|---|---|---|
| camera | No | ||
| enabled | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries full behavioral burden, and it usefully discloses the dual read/write mode via the optional parameter. It does not state persistence, whether the change is reversible, permission requirements, or what a read returns, leaving real gaps for a mutation-capable tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single efficient sentence with the read/write distinction front-loaded and no wasted words. Appropriately sized for a simple toggle.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a two-parameter toggle with no output schema and no annotations, the description covers the set case adequately but leaves the camera parameter and any return information unaddressed, so it is not fully self-sufficient.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It clarifies the meaning of "enabled" (set the HDR) but says nothing about the "camera" parameter, which is entirely undocumented in both schema and description.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a clear verb and resource: read or set the Tail 2's HDR. The resource is easily distinguished from siblings like wb, exposure, and image_adjust. It stops short of explicitly differentiating from the sibling obsbot_image_hdr, but the purpose itself is unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
"Read (bare call) or set (enabled)" tells the agent which invocation form to use (omit enabled to read, supply it to write), which is genuine usage guidance. However, it offers no guidance on when to choose this tool over related siblings such as obsbot_image_hdr or exposure settings.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
obsbot_tail2_hybrid_zoomA
Unlock (or re-lock) the Tail 2's hybrid digital zoom — the control OBSBOT Center owns. Out of the box the camera's advertised 12x range is walled at its 5x OPTICAL ceiling across the entire REST API: PUT ptz/zoom above 5.0 is acknowledged but silently pins at 5.0. Center unlocks the 5-12x digital region through its private UDP 9999 protocol; this tool speaks that protocol directly (both checksums decoded — the synthesizer reproduces Center's captured frames byte-for-byte; TAIL2-PROTOCOL.md §10). Verified by readback from the live status snapshot (zoom_infos.digital_enable); returns settled:false if the readback hadn't caught up — behavioral double-check: obsbot_tail2_zoom ratio 6.0 settles above 5.0 only when enabled. Works with OBSBOT Center closed. Hardware-verified both directions 2026-09-30, including full synthesis with a fresh sequence number.
| Name | Required | Description | Default |
|---|---|---|---|
| camera | No | ||
| enabled | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden and delivers exceptional detail: the tool speaks a private UDP 9999 protocol, returns settled:false if readback hasn't caught up, performs a behavioral double-check via obsbot_tail2_zoom ratio 6.0, works with OBSBOT Center closed, and has been hardware-verified in both directions.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the purpose well, but the description is dense and includes implementation provenance (checksums decoded, byte-for-byte reproduction, hardware verification date) that is extraneous for an agent selecting and invoking the tool. The core information could be conveyed in half the length.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with no annotations, no output schema, and 0% schema description coverage, the description must compensate for parameter and return-value gaps. It richly covers behavior but says nothing about what the 'enabled' and 'camera' parameters do, leaving an agent unable to invoke it correctly without guessing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and the description never mentions the 'enabled' boolean or the 'camera' string parameter. The phrase 'Unlock (or re-lock)' hints at a boolean toggle, but there is no explicit mapping to parameter values, and the 'camera' parameter is completely unexplained.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Unlock (or re-lock)') and resource ('hybrid digital zoom') and explains its unique purpose relative to the camera's normal REST API ceiling. Does not explicitly name sibling tools, but the contrast with 'PUT ptz/zoom above 5.0 is acknowledged but silently pins at 5.0' implicitly differentiates it from ordinary zoom tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides clear context for when to use: to access the 5–12x digital region that OBSBOT Center normally unlocks via its private protocol, and notes it works with Center closed. No explicit when-not or named alternatives, but the context of the 5x ceiling is sufficient to infer applicability.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
obsbot_tail2_image_adjustA
Read (bare call) or set the Tail 2's image style: brightness/contrast/hue/saturation/sharpness, each an absolute 0-100, plus styleMode standard|outdoor|pastel|manual. Individual control writes are MODE-GATED (measured): they apply only in styleMode manual — HTTP 500 "style mode is not manual" otherwise — so this tool applies a given styleMode FIRST and refuses a control write in a preset mode. A styleMode switch preserves the current values. Bare call returns the full style bundle.
| Name | Required | Description | Default |
|---|---|---|---|
| value | No | ||
| camera | No | ||
| control | No | ||
| styleMode | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations exist, so the description carries the full burden and does so richly: it discloses mode-gating, the exact HTTP 500 failure condition, that a styleMode switch preserves current values, and that a bare call returns the full style bundle. This is exactly the behavioral context an agent needs for a mutation tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One dense paragraph with the read/write behavior and the mode-gating constraint front-loaded before the failure detail. Every sentence earns its place, though the parenthetical HTTP 500 message is slightly verbose.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no annotations and no output schema, the description covers return behavior for the bare call, mutation constraints, and error semantics. Only the camera parameter's role is left implicit, a minor gap for an otherwise complete definition.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It documents the value range (absolute 0-100), the control enum, and the styleMode enum values, adding meaning beyond the bare schema. The camera parameter is left unexplained, keeping it from a 5.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific dual-purpose verb+resource: read via bare call or set the Tail 2's image style, enumerating exactly which controls (brightness/contrast/hue/saturation/sharpness) and modes. An agent can distinguish this from generic siblings like obsbot_image_adjust or obsbot_tail2_wb.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Clearly explains when writes succeed (styleMode manual) and when they are refused (preset modes, with the HTTP 500 message), and what a bare call does. It does not explicitly name alternative sibling tools, but the conditions for correct invocation are spelled out.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
obsbot_tail2_infoA
Read a Tail 2's identity and static configuration in one call: device_info (name, MAC, wired/wireless IPs), range (zoom 1.0–12.0, focus 1–100, exposure, white balance 2000–10000K, image adjust ranges), and networkconfig (NDI/stream encoder settings). Read-only.
| Name | Required | Description | Default |
|---|---|---|---|
| camera | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden, and it discharges the most important part by declaring 'Read-only', which tells the agent this is a safe non-mutating call. It also enumerates the returned data blocks in detail, though it says nothing about authentication, rate limits, or behavior when the camera is unreachable.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One dense sentence that front-loads the verb and the payoff ('in one call') before enumerating contents; the trailing 'Read-only' is well placed. It is information-packed without padding, though the mid-sentence enumeration is a heavy run-on.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, so the description properly takes on the job of describing return values, and it does so concretely for all three sub-blocks. The only real gap is the undocumented 'camera' argument, which leaves the agent unable to know how to target a device.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The single 'camera' parameter has 0% schema description coverage, so the description is the only place meaning could be added, and it never mentions the parameter at all. With one undocumented parameter, the description fails to compensate for the coverage gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Read) plus the resource (a Tail 2's identity and static configuration) and enumerates the three returned blocks (device_info, range, networkconfig), so the agent knows exactly what it gets. It does not, however, differentiate itself from nearby read siblings such as obsbot_tail2_status or obsbot_tail2_devices, which an agent must choose between.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'in one call' implies this is the consolidated read for static configuration, which hints at when to prefer it, but there is no explicit when/when-not statement and no named alternative. Usage is only implied, not stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
obsbot_tail2_iso_rangeA
Read or set the auto-ISO range (the double-ended slider: both bounds, 100-6400 in doubling stops). Pass min and/or max; both default to the current value when omitted. Readback-verified write.
| Name | Required | Description | Default |
|---|---|---|---|
| max | No | ||
| min | No | ||
| camera | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden and adds useful context: read-or-write semantics, current-value defaults, doubling-stop granularity, and readback-verified write behavior. It omits permissions, failure modes, or response shape, but the readback verification and defaults are meaningful disclosures for a camera control tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight sentences with no wasted words. The core purpose and slider constraint are front-loaded, followed by parameter defaults and the readback-verified write note.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 3-parameter camera control tool with no annotations and no output schema, the description covers the ISO-range parameters and write verification well. It still leaves the 'camera' parameter unexplained, which is a gap given the low schema coverage and the many sibling tools that likely share this parameter.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It does well for 'min' and 'max' by explaining the double-ended slider, 100-6400 range, doubling stops, and omission defaults, but it leaves the 'camera' parameter completely undocumented.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb pair ('read or set') and a precise resource ('auto-ISO range'), including the double-ended slider nature and bounds. It is clearly distinct from exposure, WB, and HDR tools in the sibling list, though it does not explicitly name or compare against those siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explains how to invoke the tool ('Pass min and/or max; both default to the current value when omitted') and that it can read or write. However, it gives no explicit when-to-use guidance versus alternatives such as obsbot_tail2_exposure or obsbot_image_exposure_auto/manual.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
obsbot_tail2_live_statusA
Read a FRESH runtime snapshot (GET /camera/test/status) distilled to what the REST tree hides: live gimbal pose (euler + joint angles — the WS push has no pose), zoom internals including digital_enable (the hybrid-zoom state, the readback obsbot_tail2_hybrid_zoom verifies against), runtime exposure truth (shutter/ISO/aperture as actually running), battery/lens temperatures, boot stage and uptime, and accessory state (360° base, BT remote, tally). Regenerated server-side per request.
| Name | Required | Description | Default |
|---|---|---|---|
| camera | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does well: it declares the operation is a read, notes the response is regenerated server-side per request (freshness semantics), and specifies the precise data categories returned. It omits auth/permission needs and any rate-limit or latency context, which keeps it short of a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single dense sentence front-loads the core action and then enumerates the payload; every clause adds discriminating detail (euler+joint angles, digital_enable, running shutter/ISO/aperture). The parentheticals approach over-packing but do not waste words.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only GET with no output schema, the description adequately compensates by enumerating the returned fields (pose, zoom internals, exposure truth, temperatures, boot stage, accessory state). The only real gap is the undocumented camera parameter, which an agent needs in order to call the tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
There is one parameter (camera) with 0% schema description coverage, and the description never mentions it — no format, default, or meaning. The baseline of 3+ for schema-rich tools does not apply here since the schema documents nothing, so the description should have compensated but does not.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Read) and resource (FRESH runtime snapshot from GET /camera/test/status) and enumerates what the snapshot contains. It further distinguishes itself from the WS push (no pose) and from the REST tree, so an agent can differentiate it from siblings like obsbot_tail2_status.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Implicitly conveys when to use it — when you need a fresh per-request snapshot of things the REST tree or WS push hide (gimbal pose, zoom internals, running exposure). However, it never names sibling alternatives (obsbot_status, obsbot_tail2_status, obsbot_tail2_info) or states exclusions, so routing is left to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
obsbot_tail2_only_meA
Read (bare call) or set (enabled) the Tail 2's OnlyMe human-tracking switch — track only the nearest/locked person rather than reframing for everyone.
| Name | Required | Description | Default |
|---|---|---|---|
| camera | No | ||
| enabled | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden. It usefully discloses the dual read/write invocation model, which is non-obvious and valuable. It does not describe side effects (e.g., whether enabling OnlyMe disables other tracking modes), permission/auth needs, or failure behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One well-formed sentence that front-loads the read/set modes and then explains the feature semantics. No filler, every clause earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple two-parameter device switch with no annotations or output schema, the description covers invocation modes and feature meaning adequately. The unnamed 'camera' parameter and undisclosed side effects are the only remaining gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It conveys that 'enabled' drives the set operation but says nothing about the 'camera' parameter (target device selection) and gives no type/format details. Partial compensation only.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Names a specific verb pair (read/set) and a specific resource (the OnlyMe human-tracking switch), and elaborates what the switch does functionally ('track only the nearest/locked person rather than reframing for everyone'). It does not explicitly distinguish itself from tracking siblings like obsbot_tail2_ai_track or obsbot_tail2_track_target, so it stops short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives concrete operational guidance: a bare call reads the current state, while passing 'enabled' sets it. This tells the agent how to invoke either mode. It does not, however, say when to prefer this over the broader AI-tracking tools, so the routing guidance is incomplete.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
obsbot_tail2_portraitA
Rotate a Tail 2's barrel 90° for portrait framing (enable) or back to landscape (disable) — a motorized physical rotation no Tiny 2 has. Takes ~1.5s; the tool polls the status push for the new orientation and returns settled:false if the motor hadn't finished within ~5s. MEASURED 2026-09-26: while the motor is in motion the camera can acknowledge a write and silently drop it — if settled stays false, retry the command. Hardware-verified both directions.
| Name | Required | Description | Default |
|---|---|---|---|
| camera | No | ||
| enable | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden and does so richly: ~1.5s motion time, polling of the status push, the settled:false return signal, the measured hazard that writes can be silently dropped while the motor moves, and a retry remedy. Hardware-verified framing gives the agent confidence in the reported behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the action and the enable semantics, then layers timing, failure mode, and retry guidance. Every sentence adds operative information with no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, and the description still explains the meaningful return value (settled:false) and what to do about it. Combined with the timing and failure-mode detail, an agent has everything needed to call and recover from this tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0% for 2 parameters, so the description must compensate. It fully defines enable (portrait vs landscape) but leaves the camera parameter unaddressed, relying on the shared device-selection convention across sibling tools.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Rotate a Tail 2's barrel 90°'), plus the two framing outcomes tied to the enable boolean. It also explicitly differentiates from siblings by noting this is a motorized rotation no Tiny 2 has, so the agent can rule out look-alike gimbal tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear conditional guidance: enable for portrait framing, disable to return to landscape, and adds a retry rule when settled stays false. No sibling competes for this action, so an explicit 'use X instead' clause isn't needed, though the description never states prerequisites such as device selection.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
obsbot_tail2_preset_deleteB
Delete a Tail 2 preset slot (1–3), freeing it. Verified by the slot disappearing from the list.
| Name | Required | Description | Default |
|---|---|---|---|
| slot | Yes | ||
| camera | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden. It usefully discloses the post-condition ('verified by the slot disappearing from the list') and that deletion frees the slot, but says nothing about irreversibility, error behavior on an empty/invalid slot, or permissions.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences, front-loaded with the action and scope; every clause carries information with no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a destructive, annotation-free tool with no output schema, the description covers the core action and a verification cue but omits the undocumented 'camera' parameter, failure modes, and whether deletion is permanent.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It restates the slot range (1–3, redundant with the schema's min/max) but completely ignores the second 'camera' parameter, leaving half the interface undocumented in both schema and prose.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Delete a Tail 2 preset slot'), scoped to slots 1–3, so the agent knows exactly what is affected. It does not, however, distinguish itself from the closely named generic sibling obsbot_preset_delete, leaving the model-specific vs generic choice implicit.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description never states when to use this over obsbot_preset_delete, obsbot_tail2_preset_save, or obsbot_tail2_preset_rename, nor any prerequisite (e.g., slot must already exist). Usage is only implied by the verb.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
obsbot_tail2_preset_listA
List a Tail 2's three preset slots: occupied/empty, decoded name, and pose in degrees + zoom ratio (pitch/yaw/roll/ratio). Slot numbers are 1–3 (the camera's own ids are 0-based; mapped for consistency with the Tiny 2 preset tools).
| Name | Required | Description | Default |
|---|---|---|---|
| camera | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the burden and discloses real behavior: exactly three slots, an id-mapping quirk (camera ids 0-based, exposed as 1–3 for consistency with Tiny 2 tools), and the decoded contents of each slot. It is clearly a non-destructive read. It stops short of permissions, connection state, or error behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, zero padding, and the return-value content is front-loaded before the 1–3 slot-numbering note, which is exactly the fact an agent needs before calling.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema, so the description must describe the return shape — and it does, listing occupancy, decoded name, and pose fields. Only the 'camera' selector and failure/disconnected behavior are unaddressed, which is minor for a read-only listing tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and the single optional 'camera' parameter is never mentioned, so the schema and description together leave it undocumented. The param is optional and its meaning is inferable from sibling tools, which keeps this from being a hard failure, but the description does not compensate for the coverage gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource (list a Tail 2's preset slots) and even enumerates the returned fields (occupied/empty, name, pose degrees + zoom ratio). It implicitly distinguishes itself from the sibling preset mutators (save/recall/delete/rename) by being the read path, and from the generic obsbot_preset_list by being Tail 2-scoped.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is only implied: an agent can infer you call this to inspect slots before recalling or renaming, but the description never says when to prefer it over obsbot_preset_list or the other Tail 2 preset tools, and gives no prerequisites (e.g., device must be connected).
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
obsbot_tail2_preset_recallA
Recall a Tail 2 preset slot (1–3): drives the gimbal and zoom to the saved pose. Refuses an empty slot. Arrival is verified through the only live observable — the zoom ratio from the status push — so gimbal axes are open-loop (the Tail 2 reports no live pose). Disable AI tracking first or it will fight the move. Hardware-verified 2026-09-26 (zoom 1.5 → saved 3.0 observed).
| Name | Required | Description | Default |
|---|---|---|---|
| slot | Yes | ||
| camera | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does so: it discloses the empty-slot refusal, that gimbal axes are open-loop because the Tail 2 reports no live pose, that arrival is verified only via the zoom ratio from the status push, and the AI-tracking conflict. This is exactly the kind of operational detail an agent cannot infer.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Five short sentences, front-loaded with the action and effect before the caveats. Dense but each clause adds operational value; the hardware-verification note is the only borderline sentence.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a two-parameter mutation tool with no annotations and no output schema, the description covers preconditions, failure modes, and how success is observed. The only material gap is the undocumented 'camera' parameter.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It documents the slot semantics and its 1–3 range, but the 'camera' parameter is never mentioned, leaving half the parameters unexplained in both schema and prose.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Recall a Tail 2 preset slot') with scope (1–3) and the effect ('drives the gimbal and zoom to the saved pose'). Clearly distinguishes itself from sibling preset_save/list/delete/rename without opening their schemas.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives a concrete precondition ('Disable AI tracking first or it will fight the move') and failure condition ('Refuses an empty slot'), which tell the agent when the call will succeed. It stops short of explicitly naming alternatives such as preset_save or preset_list for the save/browse cases.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
obsbot_tail2_preset_renameB
Rename a Tail 2 preset slot (1–3). Names travel base64 on the wire; the tool encodes/decodes transparently (up to 40 chars).
| Name | Required | Description | Default |
|---|---|---|---|
| name | Yes | ||
| slot | Yes | ||
| camera | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It usefully discloses a non-obvious wire behavior (names travel base64, encoded/decoded transparently), which an agent could not infer elsewhere. However, for a mutation it omits whether an existing name is overwritten, what happens on an empty slot, and whether errors are returned.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short, front-loaded sentences with no filler; the core action comes first and the encoding caveat second. Slightly dense on wire-format detail but nothing is wasted.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
A mutation tool with no annotations, no output schema, and an entirely undocumented parameter set leaves notable gaps: the undocumented camera selector and the lack of any error/overwrite behavior. The base64 note is a bright spot but does not fill the overall picture.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%. The description restates the slot range (1–3) and the 40-char limit already encoded in the schema, and only incidentally clarifies the name parameter's base64 wire format. The third parameter (camera) is never mentioned in either the schema or the description.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (rename), resource (Tail 2 preset), and scope (slot 1–3). This clearly distinguishes it from the sibling preset_save, preset_recall, preset_delete, and preset_list tools without needing to open their schemas.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The verb 'rename' combined with the slot range implies when to use it within the preset family, but the description names no alternatives and states no prerequisites (e.g., whether the slot must already be saved). Usage is implied rather than specified.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
obsbot_tail2_preset_saveA
Save a Tail 2's CURRENT live pose into a slot (1–3): aim the camera first (via tracking, a recall, or physically), then save. name defaults to P1/P2/P3. NOTE: unlike the Tiny 2, save OVERWRITES an occupied slot (no create-once), and there is no explicit-pose write in this API at all — a slot always captures the camera's CURRENT pose (measured 2026-09-26; TAIL2-PROTOCOL.md §3) — deleting first is unnecessary. Hardware-verified 2026-09-26.
| Name | Required | Description | Default |
|---|---|---|---|
| name | No | ||
| slot | Yes | ||
| camera | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With zero annotations the description carries the full burden and does so: it discloses that save OVERWRITES an occupied slot, that there is no create-once behavior, that no explicit-pose write exists in this API, and that a slot always captures the current pose. These are exactly the mutation semantics an agent needs before calling.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The core instruction and the overwrite warning are front-loaded and useful, and the sentence structure is tight. The trailing provenance noise ('measured 2026-09-26; TAIL2-PROTOCOL.md §3', 'Hardware-verified 2026-09-26') earns little for an invoking agent.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 3-param mutation tool with no annotations and no output schema, the behavior story is essentially complete — overwrite semantics, current-pose capture, and preconditions are all covered. The one gap is the unexplained 'camera' parameter, which is left to inference.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must compensate. It adds real value for two params — slot range 1–3 and the name default of P1/P2/P3, which the schema does not state — but never explains the 'camera' parameter at all, leaving one of three parameters undocumented in both places.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Save a Tail 2's CURRENT live pose into a slot') and pins the scope to slots 1–3, immediately separating it from siblings like obsbot_tail2_preset_recall, preset_rename, and preset_delete. An agent can tell exactly which preset operation this is without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives an explicit prerequisite sequence ('aim the camera first (via tracking, a recall, or physically), then save') and notes that deleting first is unnecessary, which prevents a plausible wrong workflow. It stops short of naming a sibling tool as the alternative path, so it is clear context rather than full when/when-not routing.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
obsbot_tail2_preset_speedB
Read or set the preset switching speed (1-5, Center's Console page). How fast recalls drive the gimbal. Readback-verified write.
| Name | Required | Description | Default |
|---|---|---|---|
| speed | No | ||
| camera | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It does add genuine behavioral context by stating writes are 'Readback-verified', which tells the agent writes are confirmed. However, it omits how the dual read/write mode is selected (omit 'speed'?), permission requirements, and whether changing speed affects in-flight gimbal motion.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short, front-loaded sentences with no filler; the core action and its effect lead. Slightly terse at the expense of completeness, but nothing is wasted.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No annotations, no output schema, 0% schema coverage, and an undocumented 'camera' parameter. For a dual-mode read/write tool the description does not explain how to invoke read versus write, making it under-specified for correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% for two parameters. The description compensates for 'speed' by giving its range (1-5) and meaning, but 'camera' is entirely undocumented in both the schema and the description, leaving a meaningful gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb pair and resource: read or set the preset switching speed, with the numeric range (1-5) and location ('Center's Console page'). The elaboration 'How fast recalls drive the gimbal' clarifies the effect, distinguishing it from preset_recall/preset_save siblings, though it stops short of naming those siblings directly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'Read or set' implies two modes, but there is no guidance on when to use this versus obsbot_tail2_gimbal_speed or the preset recall tools, and no prerequisite or context is given. An agent must infer usage entirely.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
obsbot_tail2_recenterA
Recenter a Tail 2's gimbal (yaw 0 / pitch 0 / roll 0). Returns as soon as the command is acknowledged — the Tail 2 exposes no live pose to verify arrival against (open-loop, like the Tiny 2 on Linux), and note that commanding a reset can drop AI tracking to none (observed once 2026-09-26; re-enable with obsbot_tail2_ai_track if that matters). Hardware-verified 2026-09-26.
| Name | Required | Description | Default |
|---|---|---|---|
| camera | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries full behavioral burden and does so well: it discloses that the call returns on acknowledgement only, that arrival cannot be verified (open-loop, no live pose), and that a reset can silently drop AI tracking to none with the recovery tool named. This is exactly the operational context an agent needs before invoking a fire-and-forget hardware command.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the action and target state, then the caveats, which is the right ordering. Slightly cluttered by journaling asides ('observed once 2026-09-26', 'Hardware-verified 2026-09-26') that add provenance but little actionable value.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema or annotations exist, and the description adequately covers behavior, caveats, and the recovery path for a one-command tool. The remaining gap is the unexplained 'camera' parameter, which an agent would have to guess at.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0% and the single 'camera' parameter has no description in either the schema or the description. The description never says what 'camera' expects (device id, name, index) or that it is optional, so it fails to compensate for the coverage gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Recenter a Tail 2's gimbal') and pins the target state (yaw 0 / pitch 0 / roll 0), so the agent knows exactly what will result. The 'Tail 2' qualifier differentiates it from the sibling obsbot_gimbal_recenter and from gimbal_move.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives an implicit usage cue by noting the AI-tracking side effect and pointing to obsbot_tail2_ai_track as the remedy, which helps an agent decide sequencing. However, it never states when to prefer this over obsbot_tail2_gimbal_move or obsbot_tail2_gimbal_position, so alternative selection is left to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
obsbot_tail2_recordA
Read (bare call) or control (enable) the Tail 2's recording switch over REST. Recording needs storage — with no SD card the state still reads ("off") but starting will not succeed; check obsbot_tail2_status's sdcard field first.
| Name | Required | Description | Default |
|---|---|---|---|
| camera | No | ||
| enable | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does meaningful work: it discloses a hidden failure mode (state reads 'off' even with no SD card, but starting will not succeed) and the storage dependency. It does not cover auth/permissions or response shape, so it stops short of full transparency.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, zero filler, and the operational mode is front-loaded before the storage caveat. Every clause adds actionable information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema and no annotations, so the description needs to be self-sufficient. It covers the recording precondition well but leaves the 'camera' parameter and the returned state shape undefined, which an agent would need in order to call this correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% for both parameters, so the description must compensate. It does explain 'enable' semantics well via the read/control duality, but the 'camera' parameter is never mentioned, leaving one of two parameters undocumented.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a concrete verb pair (read/control) and resource (the Tail 2's recording switch), and the dual-mode framing is clear: bare call reads, enabling writes. It does not explicitly differentiate from the similarly-named siblings obsbot_capture_record / obsbot_capture_stop, which would have pushed it to a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It tells the agent exactly when each mode applies (bare call for read, 'enable' for control) and adds a prerequisite workflow: check obsbot_tail2_status's sdcard field before attempting to start. No explicit naming of alternative tools, but the usage condition is concrete.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
obsbot_tail2_roll_biasA
Set a Tail 2's roll trim angle in degrees (horizon correction / deliberate tilt), clamped to ±90. Unlike the Tiny 2 (whose roll slot is inert), this moves the image. Verified write+readback on hardware 2026-09-26.
| Name | Required | Description | Default |
|---|---|---|---|
| angle | Yes | ||
| camera | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden: it discloses the ±90 clamp, that this is a write that can be read back, and that hardware behavior was verified. It stops short of saying whether the trim persists across sessions, requires a connected/powered camera, or how failure is reported.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two front-loaded sentences: purpose and semantics come first, then the hardware-verification note. The dated verification stamp is borderline noise but conveys trustworthiness and costs little.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple two-parameter setter with no annotations and no output schema, the description covers intent, units, range, and readback verification. It leaves the `camera` selector and any persistence/state caveats unexplained, which is a real gap at 0% schema coverage.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It does explain the required `angle` parameter with units and semantics (roll trim, ±90 clamp), but the optional `camera` parameter is left entirely undefined, and the clamp value merely duplicates the schema's minimum/maximum.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Set) plus resource (Tail 2's roll trim angle) and the unit (degrees), and immediately disambiguates the intent as horizon correction vs. deliberate tilt. The contrast with the Tiny 2's inert roll slot further separates it from superficially similar siblings like obsbot_tail2_image_adjust or gimbal_move.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It implies the use case (horizon correction / deliberate tilt) and flags that this actually moves the image, but never says when to prefer it over siblings such as obsbot_tail2_gimbal_move, recenter, or image_adjust. No exclusions or prerequisites are stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
obsbot_tail2_scanA
Discover OBSBOT Tail 2 cameras on the network and register them. Listens for the camera's own mDNS announcements (~5s — it multicasts its MAC, name and IPs every few seconds, the same channel OBSBOT Center discovers by) and falls back to an HTTP subnet sweep only when nothing was heard (multicast-filtered networks; ~20s on a full /24). Returns the cameras found; they stay registered afterwards. Run this once to onboard a Tail 2 whose IP you don't know.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does so richly: it discloses the mDNS listen window (~5s), the fallback HTTP subnet sweep and its cost (~20s on a full /24), and the crucial side effect that cameras 'stay registered afterwards'. This is the kind of persistence/timing/auth-relevant context an agent needs before invoking a network-scanning tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the action and resource, then layered with timing and fallback detail; every sentence adds operational value. The parenthetical about OBSBOT Center is slightly decorative but supports trust in the discovery mechanism.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no annotations and no output schema, the description still states the return value ('Returns the cameras found'), the registration side effect, and the discovery strategy. Nothing an agent needs in order to call this correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes no parameters, so per the rubric the baseline is 4. There is nothing for the description to disambiguate, and it correctly does not invent options.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Specific verb+resource: 'Discover OBSBOT Tail 2 cameras on the network and register them' tells the agent exactly what happens. Scope is further narrowed by 'onboard a Tail 2 whose IP you don't know', which implicitly separates it from the listing siblings (e.g. obsbot_tail2_devices), though no sibling is named explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Clear usage context: 'Run this once to onboard a Tail 2 whose IP you don't know' states when the tool is appropriate. It does not name a specific alternative for the already-known-IP or already-registered case, so no explicit when-not guidance is given.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
obsbot_tail2_snapshotA
Grab one still frame from a Tail 2 and return it as an image (for framing/lighting checks — and for measuring pixels). Source: the camera's SRT output via ffmpeg SRT caller (port 5000) — requires SRT listener mode in OBSBOT Center; exclusive with NDI, cleared by reboots, and Center only allows changing streaming settings while the output is off. Serves immediately once on. If SRT is off, returns a report of what it found with the available options and their costs (SRT toggle, NDI Webcam bridge caveats, NDI Tools install) so you can pick the cheapest path instead of dead-ending. resolution is the longest edge (256–1920, default 640); quality 1–100 (default 80). Needs ffmpeg. Frames are 16:9 1080p-sourced.
| Name | Required | Description | Default |
|---|---|---|---|
| camera | No | ||
| quality | No | ||
| resolution | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does so richly: it discloses the transport path (SRT/ffmpeg port 5000), the required OBSBOT Center listener mode, exclusivity with NDI, that reboots clear the setting, the streaming-settings-off constraint, the ffmpeg dependency, and — importantly — the graceful degradation when SRT is off. That is exactly the behavioral context an agent needs.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the action and its purpose, and nearly every clause carries distinct operational information. It is dense and its long fallback parenthetical runs on, but little of it is filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with no output schema and no annotations, this covers invocation prerequisites, failure behavior, dependencies, and parameter ranges well. The un-described 'camera' parameter and the absence of return-format detail (e.g., image encoding) are the remaining gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must compensate, and it explains resolution (longest edge, 256–1920, default 640) and quality (1–100, default 80) plus the 16:9 1080p source. Only the 'camera' selector is left undefined, which is a minor residual gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Grab one still frame from a Tail 2 and return it as an image') and even names the intent (framing/lighting checks, pixel measurement). It does not differentiate from close siblings like obsbot_tail2_capture_photo or obsbot_capture_snapshot, so an agent must infer which snapshot path applies.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The intent clause ('for framing/lighting checks — and for measuring pixels') implies when to reach for it, and it explains the SRT-off fallback. But there is no explicit comparison against the photo-capture siblings, so the alternative-selection guidance is only implied.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
obsbot_tail2_statusA
Read a Tail 2's full live status block (one WebSocket push): power, recording, portrait rotation, AI mode + tracking settings, zoom ratio, roll bias, focus modes, NDI/RTSP/SRT/RTMP enablement, SD card, preset list with poses (degrees + zoom ratio), and per-subsystem health (gimbal/AI/battery/lens/ToF). NOTE: no live yaw/pitch — the Tail 2 reports pose only via saved presets, so gimbal moves are open-loop (same limitation as the Tiny 2 on Linux).
| Name | Required | Description | Default |
|---|---|---|---|
| camera | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does well: it discloses the transport semantics ('one WebSocket push', i.e. a point-in-time snapshot) and a real capability limit (no live yaw/pitch, poses only via saved presets, so gimbal moves are open-loop). It stops short of stating permissions, error behavior, or refresh/caching semantics.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The long comma-delimited field list is front-loaded after the verb and every item earns its place because there is no output schema to document the return shape. The trailing NOTE is justified as a caveat, though the sentence is dense.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema and no annotations, the description steps in and documents the return contents field-by-field plus the key limitation. Only the camera parameter and any failure/timeout behavior remain uncovered.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The single optional 'camera' parameter has 0% schema description coverage and is never mentioned in the description, so an agent gets no guidance on what value to supply or how it selects among multiple devices. The neighboring obsbot_tail2_devices/obsbot_tail2_scan tools imply multi-device setups where this matters.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Read') and resource ('a Tail 2's full live status block') and then enumerates the exact fields returned, so an agent knows precisely what this yields versus the single-purpose siblings like obsbot_tail2_gimbal_position or obsbot_tail2_preset_list.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied by the completeness of the status block, but the description never says when to prefer this over the sibling obsbot_tail2_info or the generic obsbot_status, nor when to fall back to obsbot_tail2_gimbal_position for live pose. No explicit when/when-not guidance is given.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
obsbot_tail2_streamA
Read (bare call) or set the Tail 2's active network output: ndi|rtsp|srt|off — exactly ONE is active at a time (setting srt displaces ndi, and a camera reboot resets to off). This is the programmatic way to arm SRT for obsbot_tail2_snapshot without touching OBSBOT Center.
| Name | Required | Description | Default |
|---|---|---|---|
| camera | No | ||
| output | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does meaningful work: it discloses mutual exclusivity ('exactly ONE is active at a time'), the displacement side effect of setting srt, and the reboot-resets-to-off state loss. It omits any auth/permission or error behavior, which keeps it from a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One dense, front-loaded sentence pairing the read/set duality with the enum, followed by two short clauses on exclusivity and the SRT use case. No filler, though the em-dash chain makes it slightly run-on.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Covers read-vs-set behavior, the full value set, state-transition side effects, and a cross-reference to the snapshot sibling, so an agent needs nothing more for correct invocation. The undocumented 'camera' parameter and unspecified read-return shape are the only gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It explains the state semantics of the 'output' values (mutual exclusivity, reboot reset) — genuine meaning beyond the bare enum — but says nothing about the 'camera' parameter (presumably a target selector), leaving half the parameters undocumented.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a precise dual verb ('Read (bare call) or set') on a specific resource ('Tail 2's active network output') and enumerates the accepted values. An agent can distinguish it from every gimbal/image/preset sibling at a glance.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives a concrete motivating context ('programmatic way to arm SRT for obsbot_tail2_snapshot without touching OBSBOT Center') and implies the GUI alternative. It does not spell out when a user should prefer NDI/RTSP or when not to change output at all, so it stops short of full when/when-not guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
obsbot_tail2_stream_configA
Read (bare) or set the stream encoder configuration shared by the network outputs: encoder (h264/h265), resolution (camera format like 1920X1080P30), bitrate (Mbps). The bare read also includes the RTSP URLs for both network interfaces. GATED (MEASURED 2026-10-01): resolution writes 500 while an output is live — set obsbot_tail2_stream output:"off" first (encoder and bitrate writes are ungated). Readback-verified writes.
| Name | Required | Description | Default |
|---|---|---|---|
| camera | No | ||
| bitrate | No | ||
| encoder | No | ||
| resolution | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden and does well: it discloses the gated resolution-write failure (HTTP 500 while live), the prerequisite tool/action, which writes are ungated, and that writes are readback-verified. It does not cover permissions/auth or persistence/rollback behavior, but it is far above typical transparency.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single dense paragraph front-loads the read/write duality, then lists settings, then adds the critical gate. Every clause adds operational value, though the style is packed rather than cleanly sectioned.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a four-parameter hardware config tool with no annotations and no output schema, the description supplies critical operational constraints: gating, prerequisite tool call, readback verification, and the bare-read RTSP URL payload. Missing semantics for the camera parameter and write return details keep it from a 5.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and there are four parameters. The description explains encoder values, resolution format with an example, and bitrate units, but the 'camera' parameter is never mentioned. This is only partial compensation for the coverage gap, so a baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource: read or set the stream encoder configuration shared by the network outputs, and enumerates the settings (encoder, resolution, bitrate). It distinguishes itself from the sibling obsbot_tail2_stream by focusing on encoder config rather than output on/off, so an agent can identify it without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly explains bare read vs parameterized write, and gives the critical gating condition: resolution writes return 500 while an output is live, so obsbot_tail2_stream output:"off" must be called first. It also names which writes are ungated (encoder, bitrate), providing both when-to-use and when-not-to-use guidance plus an alternative tool.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
obsbot_tail2_track_customA
Configure CUSTOM tracking speed (Center's customized track speed): enable it, set the per-axis Pan/Tilt speeds 1-10, and toggle the per-axis Auto buttons. Bare call reads current state from the status push (tracking_settings, locks included read-only). Writes ride UDP 9999 keys 04/07/0a/06/09 (pan/tilt travel as float32 slider/10 on the wire; keys 06/09 are the Auto buttons — hardware-verified against horizontal_auto/vertical_auto). The AXIS LOCKS have no known write path (not UDP, REST guesses 404) — shown read-only. Verified by tracking_settings readback.
| Name | Required | Description | Default |
|---|---|---|---|
| pan | No | ||
| tilt | No | ||
| camera | No | ||
| enabled | No | ||
| panAuto | No | ||
| tiltAuto | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full behavioral burden and does so richly: it discloses the write transport (UDP 9999, keys 04/07/0a/06/09), the wire encoding (float32 slider/10, keys 06/09 as Auto buttons verified against horizontal_auto/vertical_auto), that readback comes from tracking_settings, and that the axis locks have no write path and are read-only. This is far beyond what structured fields supply.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the action (enable, set speeds, toggle autos), then the bare-call read behavior, then the wire-level details. Dense but each clause carries information; the protocol keys/encoding are arguably more than some callers need, but they are the tool's differentiating content.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 6-param mutation tool with no annotations and no output schema, the description covers writes, reads, encoding, verification, and the read-only lock caveat. The missing `camera` parameter semantics and the unstated default behavior when only some axes are passed are the remaining gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate, and it largely does: it explains the pan/tilt 1-10 speed semantics, the enabled toggle, and maps the Auto toggles to specific wire keys. The only gap is the `camera` parameter, which is never mentioned anywhere in the description or schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource: configure the CUSTOM ('customized') tracking speed, with per-axis Pan/Tilt values 1-10 and per-axis Auto toggles. It even names the internal concept (Center's `customized` track speed), which helps distinguish it conceptually. However it never names the sibling `obsbot_tail2_track_speed` or `obsbot_ai_track_speed`, so the agent must infer the split between preset speed vs custom speed.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides one useful usage signal: a bare call reads current state from the status push rather than writing, which tells the agent it can use this as a readback. But there is no explicit when-to-use / when-not, no mention of when to prefer `obsbot_tail2_track_speed`, and no stated prerequisites for the write path.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
obsbot_tail2_track_speedA
Set a Tail 2's tracking follow speed: superLazy | lazy | slow | fast | crazy | customized — the camera's own six-speed enum (NOT the Tiny 2's standard/sport pair). Verified by readback on hardware 2026-09-26.
| Name | Required | Description | Default |
|---|---|---|---|
| speed | Yes | ||
| camera | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden. It adds useful provenance ('verified by readback on hardware') and a cross-model caveat, but says nothing about whether tracking must already be enabled, whether the setting persists across sessions, or what happens with 'customized' — meaningful gaps for a mutation tool with zero annotation coverage.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Single sentence, front-loaded with the verb and the enum, with the differentiator placed where it is read first. The 'verified by readback ... 2026-09-26' clause is provenance rather than instruction and is slightly expendable, but overall there is little waste.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a two-parameter mutation tool with no annotations and no output schema, the description covers the core enum well but omits the 'camera' parameter, preconditions, and the effect of the setting. It is adequate but leaves real gaps an agent would have to probe elsewhere.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It enumerates and gives semantic flavor to the 'speed' values (superLazy through crazy) beyond the bare schema enum, but the second parameter, 'camera', is never mentioned anywhere, and 'customized' is left unexplained.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Set a Tail 2's tracking follow speed') and pins the exact enum it accepts. It also distinguishes itself from a near-miss sibling family by noting this is Tail 2's six-value enum, not the Tiny 2's standard/sport pair, so an agent can separate it from obsbot_tail2_gimbal_speed and obsbot_ai_track_speed.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied by the device/model scoping (Tail 2 only, and explicitly not Tiny 2), but there is no explicit when-to-use statement, no prerequisite (e.g. AI tracking must be active), and no routing to an alternative tool. It gives enough to infer context but leaves selection partly to the agent.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
obsbot_tail2_track_targetA
Tap-to-track: give a normalized frame coordinate (x/y 0.01-0.99, top-left ~0,0) and the camera engages AI tracking on the subject there — measured to arm humanTrackingSingleMode on its own from mode none. Divide a snapshot pixel by frameWidth/frameHeight and pass it straight in. The mode that engages depends on what is at the point, so read it back (obsbot_tail2_status ai_mode) when it matters; objectTracking cannot be entered this way (it demands a bounding box the API grammar for which is not yet decoded).
| Name | Required | Description | Default |
|---|---|---|---|
| x | Yes | ||
| y | Yes | ||
| camera | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does so well: it discloses the resulting mode is data-dependent ('depends on what is at the point'), that humanTrackingSingleMode is armed autonomously from mode none, and that objectTracking is unreachable because its bounding-box grammar is undecoded. These are non-obvious state-transition facts an agent could not infer from the schema.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the 'Tap-to-track' concept and the coordinate contract, then layers mode behavior and the readback caveat. It is dense and dash-heavy but each clause adds distinct information; slightly tighter phrasing would help.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 3-param mutation tool with no annotations and no output schema, the description covers the coordinate contract, mode transitions, verification path, and one hard exclusion. Only the camera parameter's meaning is left unexplained, so it is nearly, but not fully, complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must compensate, and it does for x/y by defining the normalized 0.01-0.99 range, the top-left ~0,0 origin, and the exact derivation (snapshot pixel / frameWidth|frameHeight). The third parameter, camera, receives no explanation at all, leaving a gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('the camera engages AI tracking on the subject there') anchored to an input coordinate, and it distinguishes the behavior from adjacent paths by naming the modes it can and cannot reach (humanTrackingSingleMode vs objectTracking). An agent can tell this apart from gimbal-aiming or generic track tools from the description alone.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear context: feed a normalized coordinate and tracking engages from mode none, plus an explicit exclusion (objectTracking cannot be entered this way) and a readback instruction via obsbot_tail2_status. It does not, however, route the agent between this and sibling tracking entry points such as obsbot_tail2_ai_track or obsbot_aim_at_pixel, so the alternative-selection guidance is incomplete.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
obsbot_tail2_usb_modeB
Read or set the USB-C function mode: mtp or uvc. Control rides the network either way, so switching is safe remotely. mtp is the measured wire value; uvc is the presumed counterpart (GET-measured 2026-10-01). Readback-verified write.
| Name | Required | Description | Default |
|---|---|---|---|
| mode | No | ||
| camera | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden, and it does contribute real behavioral context: remote safety of switching, network-based control path, and a readback-verified write. However it omits permission/auth requirements, what happens on failure, and any latency or disconnection risk, which matters for a mode switch that could drop the UVC stream.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight sentences, front-loaded with the action and valid values. The measurement-date parenthetical is slightly opaque but short enough to be tolerable.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, so the description should clarify what a read returns, and with no annotations and 0% schema coverage, the `camera` parameter and the omitted-mode read behavior are gaps. It is adequate but leaves an agent guessing on read semantics and targeting.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It usefully explains the mtp/uvc value semantics ('measured wire value' vs 'presumed counterpart') and the measurement provenance, but the second parameter `camera` is never mentioned, leaving half the inputs undocumented.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb pair (read or set) and resource (USB-C function mode), and enumerates the two valid values. It is distinguishable from siblings, but doesn't explicitly contrast with any of them or state the read-vs-write default when `mode` is omitted.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no when-to-use guidance or alternative routing. The note 'Control rides the network either way, so switching is safe remotely' is a safety reassurance, not a usage condition, and nothing tells the agent how this differs from obsbot_tail2_stream or other connection-related siblings.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
obsbot_tail2_wbA
Read (bare call) or set the Tail 2's white balance: mode auto|daylight|fluorescent|tungsten|cloudy|manual, plus temperature in Kelvin (2000-10000) which applies in manual mode.
| Name | Required | Description | Default |
|---|---|---|---|
| mode | No | ||
| camera | No | ||
| temperature | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It discloses the dual read/write behavior and the key constraint that temperature only takes effect in manual mode, which is genuinely useful. It does not state whether setting requires the camera to be active, what the read returns, or any rate/permission constraints.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence that packs read/set duality, the enumeration, and the manual-mode condition without wasted words.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no annotations, no output schema, and 0% schema coverage, the description covers the core read/write semantics and parameter meaning well but omits the camera parameter entirely and says nothing about return format or error/edge behavior.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must compensate. It restates the enum values and the 2000-10000 Kelvin range (duplicating schema) but adds real semantics: temperature applies only in manual mode. The third parameter, camera, is undocumented in both schema and description, leaving a genuine gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific action pair (read via bare call, or set) on a specific resource (Tail 2's white balance), and enumerates the mode values inline. An agent immediately understands this covers both querying and configuring white balance, distinguishable from exposure, focus, and image-adjust siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies the two usage modes: call bare to read, pass mode/temperature to set. However it never names or contrasts the closest siblings (obsbot_image_wb_auto, obsbot_image_wb_manual), so an agent must infer which white-balance tool to pick.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
obsbot_tail2_zone_trackingB
Toggle zone tracking (Center's Console page switch) on or off. Rides the camera's private UDP 9999 channel (generic key-value command, key 03 — decoded from a labeled capture 2026-10-01, TAIL2-PROTOCOL.md §10b). No ack and no REST-readable state — the effect shows in the tracking behavior. Synthesized frame, works with Center closed.
| Name | Required | Description | Default |
|---|---|---|---|
| camera | No | ||
| enabled | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does so well: it discloses the transport (private UDP 9999, key 03), the critical fact that there is no ack and no REST-readable state, and that the effect must be observed indirectly in tracking behavior. That is exactly the kind of side-effect/no-confirmation detail an agent needs. It stops short of 5 because the verification cue ('effect shows in the tracking behavior') is vague about latency or how to confirm success.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the action and its UI equivalence, then the operational caveats. The capture date and TAIL2-PROTOCOL.md §10b citation are implementation provenance that does not help an agent select or invoke the tool, but they are brief.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a no-annotation, no-output-schema, two-parameter mutation tool, the description covers the behavioral risks (no ack, no readable state) but leaves the `camera` parameter and the boolean's exact semantics undocumented. Adequate but with a clear gap on the input contract.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% for both parameters. The description's 'on or off' loosely explains `enabled`, but `camera` — arguably the more consequential parameter for a multi-device setup — is never mentioned or explained in either the schema or the description.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Toggle zone tracking') and anchors it to a known UI concept ('Center's Console page switch'), so the agent knows exactly what capability is being flipped. It does not, however, differentiate itself from plausible siblings like obsbot_tail2_ai_track or obsbot_tail2_track_custom, which also concern tracking behavior.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied rather than stated: 'works with Center closed' tells the agent this path is usable when the app is not running, which is a real contextual cue. But there is no explicit when-to-use/when-not guidance and no naming of an alternative when the effect matters.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
obsbot_tail2_zoomA
Set absolute zoom ratio on a Tail 2 (1.0–12.0 — twelve x, six times the Tiny 2's range; scale is the camera's own ratio, not a magnification multiple). speed 1–10 is REQUIRED by the firmware and defaults to 5. Zoom ramps mechanically; the tool polls the readback and returns settled:false if it hadn't arrived within ~5s — the command was still sent. Hardware-verified 2026-09-26.
| Name | Required | Description | Default |
|---|---|---|---|
| ratio | Yes | ||
| speed | No | ||
| camera | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and delivers: firmware requires speed (defaults to 5), zoom ramps mechanically, the tool polls the readback, and it returns settled:false if the value hadn't arrived within ~5s while still having sent the command. This is exactly the kind of timing, side-effect, and return-behavior detail an agent needs.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with purpose and range, then the speed requirement, then the behavioral caveat. Dense but every sentence carries distinct, useful information; the parenthetical qualifiers clarify rather than pad.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no annotations or output schema, the description supplies the missing behavioral and return semantics (settled flag, ~5s polling window, mechanical ramp). It stops short of explaining the 'camera' selector or what the settled payload contains, but is otherwise complete for a single-axis zoom command.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must compensate, and it does well for two of three params: it explains the ratio range (1.0–12.0) and its scale semantics and the speed range (1–10) plus the firmware requirement/default. The 'camera' parameter is left completely undocumented, a small but real gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description gives a specific verb+resource ('Set absolute zoom ratio on a Tail 2') and clarifies the scale is the camera's own ratio, not a magnification multiple. However, it never distinguishes itself from the many sibling zoom tools (obsbot_zoom_uvc, obsbot_zoom_vendor, obsbot_zoom_to_fit), leaving the agent to infer which zoom primitive to pick.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no explicit when-to-use guidance and no comparison to alternative zoom tools in the sibling list. The range and speed notes hint at usage, but the agent gets no rule for choosing this tool over the UVC/vendor/to-fit zoom variants.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
obsbot_tail2_zoom_typeA
Read or set the auto-zoom framing pattern (Center's Console slider under Single/Group tracking: off/3/5/7/9/…/24). Off=normal, 7/9/16/24=P7/P9/P16/P24 (subject-size framing percentages), 3/5≈halfBody/fullBody (3 disabled in group mode). Arming AI tracking sets this to shot — which is why the status block's zoom_type reads shot while tracking. Writes are GATED (MEASURED 2026-10-01): the camera 500s unless human tracking is armed — obsbot_tail2_ai_track enabled:true first. Readback-verified write.
| Name | Required | Description | Default |
|---|---|---|---|
| type | No | ||
| camera | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the entire behavioral burden and does so well: it discloses that writes are gated and fail with a 500 unless tracking is armed, that arming AI tracking silently forces the value to `shot`, that 3 is disabled in group mode, and that writes are readback-verified. This is exactly the mutability and side-effect context an agent needs.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core action before the enum explanation, gating caveat, and status-read explanation. It is dense and parenthetical but every clause contributes operationally relevant information; slightly more structure would help readability but nothing is wasted.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read/write tool with no annotations and no output schema, the description covers the critical gaps: write gating, the silent `shot` side effect, mode-dependent value availability, and value semantics. Only the `camera` selector and any return-shape expectations remain unaddressed.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must compensate, and it largely does for `type`: it maps off/3/5/7/9/16/24 to normal, halfBody, fullBody, and P7/P9/P16/P24 framing percentages, adding real meaning beyond the bare enum list. The `camera` parameter is left undocumented in both schema and description, which is the only gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb pair and resource: 'Read or set the auto-zoom framing pattern', and enumerates what the framing values mean (subject-size framing percentages), which implicitly separates it from optical-zoom siblings like obsbot_tail2_zoom and obsbot_tail2_hybrid_zoom. It never names an alternative sibling explicitly, so it stops short of full differentiation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear operational context: the values map to Console slider positions under Single/Group tracking, and writes will 500 unless human tracking is armed first. It states the prerequisite ('obsbot_tail2_ai_track enabled:true first') but never routes the agent away from sibling tools for adjacent needs.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
obsbot_wakeB
Wake the camera/gimbal (sends "run"). This MOVES the camera: waking un-stows the gimbal and brings it back to level (pitch ~0). Most control commands also wake the camera implicitly as a side effect.
| Name | Required | Description | Default |
|---|---|---|---|
| camera | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description reveals that waking moves the camera (un-stows, brings to level) and sends a 'run' command. However, without annotations, it does not fully disclose behavioral traits like idempotency, duration, or safety of calling while already awake.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is concise with three sentences, each adding value. No unnecessary words. Could be slightly improved by front-loading the most critical information, but overall well-structured.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's simplicity (one optional param, no output schema), the description covers the core action and its physical effect. However, it omits details about the camera parameter and any return behavior, leaving some gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The description does not mention the single parameter 'camera' at all. With 0% schema coverage, the description fails to add any meaning beyond the raw schema, leaving the agent unaware of the parameter's purpose or usage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool wakes the camera/gimbal, specifies the physical action (un-stows, brings to level), and distinguishes from sibling tools like gimbal_move or gimbal_recenter. It also notes that most control commands wake implicitly, helping differentiate use cases.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies that explicit waking may be unnecessary if a control command follows, but it does not explicitly state when to prefer this tool over alternatives or when not to use it. Lacks explicit guidance on prerequisites or exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
obsbot_zoom_to_fitA
Frame a region of a frame you just captured: centre the gimbal on it and zoom so the region fills the frame. Give x/y/width/height of the region plus the frameWidth/frameHeight from THE SAME obsbot_capture_snapshot result — mixing a region from one frame with dimensions from another frames the wrong place and cannot be detected. Must come from a snapshot, and takes the same source declaration as obsbot_aim_at_pixel. margin (default 0.1) backs the zoom off by that fraction so the region isn't framed edge-to-edge; the tighter of the region's two axes decides the zoom, so the WHOLE region stays visible rather than being cropped on one side. Moves the gimbal BEFORE zooming, since zooming first can push the region's centre out of frame. Refuses on the same conditions as obsbot_aim_at_pixel: AI tracking active, the camera was asleep (waking it moves the gimbal and invalidates the frame), the FOV mode can't be decoded, a corrupt zoom reading, or the region's centre lying past vertical from the current pose, or a frame that isn't 16:9 (obsbot_capture_snapshot always returns 16:9; a non-16:9 pair looks transposed). Also refuses a region that isn't within the frame (edges included), or has non-positive width/height. The requested zoom is clamped to the camera's [1x, 4x] magnification range and reported via clamped; a partial fit still moves and zooms to the limit. Zoom ramps rather than jumping, so the tool polls the status block for up to 3s waiting for it to arrive and returns settled:false (not an error) if it didn't — a frame captured mid-ramp is at an unknown magnification, so check settled before trusting a follow-up snapshot.
| Name | Required | Description | Default |
|---|---|---|---|
| x | Yes | ||
| y | Yes | ||
| width | Yes | ||
| camera | No | ||
| height | Yes | ||
| margin | No | ||
| source | No | device | |
| frameWidth | Yes | ||
| frameHeight | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With zero annotations, the description carries the full burden and meets it exceptionally: it discloses execution ordering ('moves the gimbal BEFORE zooming'), silent failure modes ('mixing a region from one frame with dimensions from another... cannot be detected'), nine specific refusal conditions, zoom clamping to [1x, 4x] with a `clamped` report, the 3-second polling window, and that `settled:false` is returned as a non-error. This goes far beyond what any annotation set would provide.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is long (roughly 250 words) but appropriately so for a tool with 9 parameters and numerous edge cases; no sentence is filler. It is front-loaded with the core purpose and flows logically from input constraint (same snapshot) to margin behavior, execution ordering, refusal conditions, clamping, and settling. The dense refusal-condition list is slightly hard to parse, but every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool this complex — no annotations, no output schema, 0% schema descriptions — the description covers an extraordinary amount: prerequisites, ordering, silent-failure warning, refusal conditions, clamping, and the meaningful return fields `clamped` and `settled`. The only real gaps are the coordinate system origin/units for the region parameters and the full success-return shape beyond `clamped`/`settled`, which keeps it from a perfect score.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate, and it substantially does: it explains the critical same-snapshot dependency between x/y/width/height and frameWidth/frameHeight, the margin parameter's backing-off behavior with default 0.1, the source enum's shared semantics with obsbot_aim_at_pixel, and the 16:9 transposition tell for swapped dimensions. Minor gaps remain — pixel units and coordinate origin for x/y/width/height are never stated — but the essential semantic relationships are well covered.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The opening sentence 'Frame a region of a frame you just captured: centre the gimbal on it and zoom so the region fills the frame' names a specific composite operation (aim + zoom) on a specific resource (a region of a captured frame). It clearly differentiates from siblings like obsbot_aim_at_pixel (single-pixel aim), obsbot_gimbal_move (raw movement), and obsbot_zoom_uvc/zoom_vendor (raw zoom), and references these relations explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description establishes a clear workflow context: use this after obsbot_capture_snapshot, and it says it 'takes the same source declaration as obsbot_aim_at_pixel' and 'refuses on the same conditions as obsbot_aim_at_pixel', anchoring its behavior to a sibling. It does not explicitly enumerate when-not-to-use versus raw zoom/move tools, but the snapshot-dependent precondition and framing purpose make the selection criteria largely unambiguous.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
obsbot_zoom_uvcA
Standard UVC zoom: set an absolute zoom ratio, clamped to [1.0, 2.0]. Snaps to the requested target exactly (unlike obsbot_zoom_vendor, whose ratio scale differs and may not land exactly where asked). Waits for the zoom to actually arrive and returns { settled }: the ramp is not instant (a full 1.0->2.0 sweep takes about 2.4s), and obsbot_aim_at_pixel and obsbot_zoom_to_fit both refuse while it is in flight, so this returning early would just move the failure downstream. settled:false means the zoom had not arrived within the timeout — the command was still sent.
| Name | Required | Description | Default |
|---|---|---|---|
| ratio | Yes | ||
| camera | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Discloses clamping, exact snapping, wait behavior, timeout (settled:false), duration (~2.4s sweep), and downstream impacts. With no annotations, this is comprehensive and honest.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with core purpose; behavioral details follow logically. Slightly long but every sentence adds value. Could be more compact while retaining clarity.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Covers zoom behavior and return value well, but omits camera parameter meaning. No output schema, so return format is clarified, but missing parameter info reduces completeness.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The 'ratio' parameter is described as absolute zoom ratio clamped to [1.0,2.0], but the 'camera' parameter is entirely undocumented. Schema coverage is 0%, so description should cover both parameters.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Clearly states 'set an absolute zoom ratio' with specific resource and verb. Distinguishes from sibling obsbot_zoom_vendor by noting exact snapping vs. vendor's differing scale.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly contrasts with obsbot_zoom_vendor and warns that obsbot_aim_at_pixel and obsbot_zoom_to_fit refuse while zoom is in flight, providing clear context for when to use this tool.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
obsbot_zoom_vendorA
Vendor zoom path with adjustable speed: zoom to a ratio at a chosen speed. This tool's ratio scale differs from obsbot_zoom_uvc's and may not land exactly on the requested target. ratio is clamped to [1.0,2.0]; speed 0=device default, 1-10 slow→fast, 255=maximum.
| Name | Required | Description | Default |
|---|---|---|---|
| ratio | Yes | ||
| speed | No | ||
| camera | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It reveals that the ratio is clamped, speed has special values, and landing may not be exact. This offers good insight into tool behavior, though it does not disclose whether the operation is safe or requires authorization.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is concise, consisting of two sentences that front-load the main purpose and then add differentiation and parameter details. It avoids unnecessary words, though a more structured format could improve readability.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with 3 parameters, no output schema, and no annotations, the description covers the main purpose, parameter ranges, and differentiation from siblings. However, it omits the camera parameter entirely and does not mention return values or error conditions, leaving some context incomplete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description should fully compensate. It explains ratio and speed semantics (clamping, range, special values) but provides no explanation for the 'camera' parameter, leaving its purpose unclear.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose: it zooms to a ratio at a chosen speed on a vendor zoom path. It distinguishes itself from the sibling tool obsbot_zoom_uvc by noting a different ratio scale and potential inexact landing.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides guidance on when to use this tool versus obsbot_zoom_uvc, noting the scale difference and inexact landing. It also gives specific ranges for ratio and speed parameters. However, it does not specify prerequisites or when not to use the tool.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
14 tool updates
v0.9.2- Added
obsbot_tail2_af_track - Added
obsbot_tail2_antiflicker - Added
obsbot_tail2_auto_zoom_speed - Added
obsbot_tail2_export_log - Added
obsbot_tail2_gesture - Added
obsbot_tail2_hybrid_zoom - Added
obsbot_tail2_iso_range - Added
obsbot_tail2_live_status - Added
obsbot_tail2_preset_speed - Added
obsbot_tail2_stream_config - Added
obsbot_tail2_track_custom - Added
obsbot_tail2_usb_mode - Added
obsbot_tail2_zone_tracking - Added
obsbot_tail2_zoom_type
32 tool updates
v0.9.0- Added
obsbot_tail2_ai_track - Added
obsbot_tail2_audio - Added
obsbot_tail2_capture_photo - Added
obsbot_tail2_devices - Added
obsbot_tail2_exposure - Added
obsbot_tail2_focus - Added
obsbot_tail2_focus_point - Added
obsbot_tail2_gimbal_invert - Added
obsbot_tail2_gimbal_move - Added
obsbot_tail2_gimbal_position - Added
obsbot_tail2_gimbal_speed - Added
obsbot_tail2_hdr - Added
obsbot_tail2_image_adjust - Added
obsbot_tail2_info - Added
obsbot_tail2_only_me - Added
obsbot_tail2_portrait - Added
obsbot_tail2_preset_delete - Added
obsbot_tail2_preset_list - Added
obsbot_tail2_preset_recall - Added
obsbot_tail2_preset_rename - Added
obsbot_tail2_preset_save - Added
obsbot_tail2_recenter - Added
obsbot_tail2_record - Added
obsbot_tail2_roll_bias - Added
obsbot_tail2_scan - Added
obsbot_tail2_snapshot - Added
obsbot_tail2_status - Added
obsbot_tail2_stream - Added
obsbot_tail2_track_speed - Added
obsbot_tail2_track_target - Added
obsbot_tail2_wb - Added
obsbot_tail2_zoom
8 tool updates
v0.7.0- Added
obsbot_aim_at_pixel - Added
obsbot_capture_record - Added
obsbot_focus_auto - Added
obsbot_gimbal_position - Added
obsbot_image_wb_auto - Added
obsbot_image_wb_manual - Added
obsbot_sleep - Added
obsbot_zoom_to_fit
13 tool updates
v0.6.2- Added
obsbot_capture_list - Added
obsbot_capture_preview - Removed
obsbot_capture_record - Added
obsbot_capture_snapshot - Removed
obsbot_focus_auto - Removed
obsbot_gimbal_position - Removed
obsbot_image_wb_auto - Added
obsbot_preset_delete - Added
obsbot_preset_recall - Added
obsbot_preset_save - Removed
obsbot_zoom_to_fit - Added
obsbot_zoom_uvc - Added
obsbot_zoom_vendor
23 tool updates
v0.6.2- First observed
obsbot_ai_track - First observed
obsbot_ai_track_speed - First observed
obsbot_capture_record - First observed
obsbot_capture_stop - First observed
obsbot_devices - First observed
obsbot_focus_auto - First observed
obsbot_focus_face - First observed
obsbot_focus_manual - First observed
obsbot_gimbal_move - First observed
obsbot_gimbal_position - First observed
obsbot_gimbal_recenter - First observed
obsbot_image_adjust - First observed
obsbot_image_exposure_auto - First observed
obsbot_image_exposure_manual - First observed
obsbot_image_fov - First observed
obsbot_image_hdr - First observed
obsbot_image_wb_auto - First observed
obsbot_preset_list - First observed
obsbot_preset_rename - First observed
obsbot_preset_update - First observed
obsbot_status - First observed
obsbot_wake - First observed
obsbot_zoom_to_fit
TDQS
Scored across 80 tools
The set includes two overlapping families (generic obsbot_* seemingly for Tiny 2 and obsbot_tail2_*) without a device qualifier in the generic names, so an agent must infer target hardware from descriptions. Within each family, numerous tracking, preset, and zoom tools share conceptual territory, increasing misselection risk despite detailed descriptions.
All names use snake_case with the obsbot_/obsbot_tail2_ prefix, but the convention is mixed between verb_noun (e.g., obsbot_gimbal_move, obsbot_tail2_preset_list) and bare noun phrases (e.g., obsbot_tail2_audio, obsbot_tail2_wb, obsbot_tail2_info). The prefixing is consistent, but the lack of a predictable read/write action pattern makes names only moderately predictable.
80 tools is far beyond the recommended 3-15 range and exceeds even the 25+ threshold for 'too many'. Although a large camera-control surface could justify many endpoints, this volume imposes extreme context load and selection difficulty.
The surface covers CRUD/lifecycle for presets, gimbal, zoom, focus, exposure, white balance, image adjust, AI tracking, capture, streaming, status, and device discovery for two camera families. Missing operations are minor (e.g., no Tiny 2 network scan, no firmware update), so agents can work around gaps.
Maintenance
Related MCP Connectors
Turns a phone into a camera+Bluetooth remote so AI assistants can see and control any PC.
Drive real devices from your AI Coding tool. Embed a client SDK (Unity, Godot, Flutter, iOS/macOS, Android, React Native, Web) in your app, then capture screenshots, traverse the UI tree, inject taps and key events, and run automated test tasks on the physical device over a secure relay.
Operate Linux, macOS and Windows from your LLM. Every action runs through an auditable allowlist.
Control Unreal Engine to browse assets, import content, and manage levels and sequences. Automate…
Related MCP Servers
- AlicenseBqualityDmaintenanceEnables PTZ camera control with gimbal positioning, snapshots, and AI visual analysis for OBSBOT and UVC cameras. Supports autonomous scanning patterns and integrates with vision-language models for real-time camera analysis.71MIT
- AlicenseNot gradedqualityDmaintenanceEnables LLM-based AI agents to control SO-ARM100 and SO-101 robots through natural language commands and camera feedback. It supports various transport protocols and provides tools for both autonomous robotic movement and manual keyboard operation.85Apache 2.0
- FlicenseAqualityDmaintenanceControls Reachy Mini robot, enabling movement of head, body, and antennas, playing animations, capturing images, text-to-speech, and WebRTC streaming via natural language.19-
- FlicenseNot gradedqualityDmaintenanceProvides visual recognition and PTZ camera control for DIY MOSS (ESP32) smart assistants, bridging cloud AI with local camera capabilities via WebSocket.-