Fast Video Cataloger
OfficialFast Video Cataloger MCP server
Search, view and tag your local video library from an AI assistant.
Fast Video Cataloger keeps a searchable catalog of your own video files on Windows: scene thumbnails, keywords, transcripts and faces. This MCP server lets an assistant such as Claude work with that catalog. You ask for footage in your own words; the assistant searches, looks at the actual shots and does the filing for you.
Find videos and moments by what was said, with timecodes
Find scenes by describing the picture ("a drone shot over a city at night"), nothing tagged first
Fetch scene thumbnails so the assistant can see and judge a shot
With an Editor key: add or remove keywords, build bins, index new videos
It runs against the Fast Video Cataloger server on your own computer. No video is uploaded; the assistant receives the titles, transcript lines and scene images it asks for.
More: MCP server for your video library
Requirements
Windows 10 or 11 (x64)
Fast Video Cataloger 10.4 or later. The free 30-day trial works.
fvc-mcp.exeis also installed with the program, together with the .NET 10 runtime it needs.Your catalog shared through the Fast Video Cataloger server: in the program, open the start page's Server section and click Setup. The assistant connects to the server, not to the desktop program, so it works with the program closed.
An API key: Manage Users in the same Server section.
Viewerto search and read,Editorto also tag and build bins.
Related MCP server: mcp-video-vision
Install
Claude Desktop: one click
Download fvc-mcp-<version>.mcpb from Releases and
open it. Claude Desktop asks for the server address (leave http://localhost:8754 when the server
runs on this computer) and your API key.
Claude Desktop: config file
Use the copy installed with Fast Video Cataloger. In claude_desktop_config.json:
{
"mcpServers": {
"fast-video-cataloger": {
"command": "C:\\Program Files\\FastVideoCataloger\\fvc-mcp.exe",
"env": {
"FVC_SERVER_URL": "http://localhost:8754",
"FVC_API_KEY": "fvc_your_key_here"
}
}
}
}Claude Code
claude mcp add fast-video-cataloger --env FVC_SERVER_URL=http://localhost:8754 --env FVC_API_KEY=fvc_your_key_here "--" "C:\Program Files\FastVideoCataloger\fvc-mcp.exe"Keep the quotes around "--" in PowerShell.
Gemini CLI
gemini mcp add fast-video-cataloger "C:\Program Files\FastVideoCataloger\fvc-mcp.exe" -s user -e FVC_SERVER_URL=http://localhost:8754 -e FVC_API_KEY=fvc_your_key_hereFull setup, things to try and troubleshooting: Connect an AI Assistant (MCP).
Configuration
Variable | Meaning |
| Address of your catalog server. Default |
| The API key. |
|
|
Tools (23)
Tool | What it does |
| Find videos by free text or by keyword. |
| Find individual scenes, with their timecodes. |
| Find scenes by describing the picture, or scenes that look like a given one. Needs the visual search index and, for descriptions, the Visual search model on the server. |
| Find spoken words, with their video and timecode. |
| Find people by name or tag. |
| Every keyword in the catalog, so the assistant knows what it can search for. |
| One catalog entry in full. |
| A video's scene thumbnails in timecode order. |
| The transcript as timed lines. |
| Keywords and cast. |
| The thumbnail image itself, so the assistant can see the shot. |
| Catalog size, the build serving it, and existing bins. |
| Add keywords to a video or a single scene. Editor key. |
| Remove named keywords again. Editor key. |
| Build a collection. Editor key. |
| Index catalog entries that have no thumbnails yet. Editor key. |
It cannot delete videos or bins, take a video out of a bin, play video or drive the player.
Source
src/FvcMcpServer is the source of the fvc-mcp.exe shipped with Fast Video Cataloger, copied
from each release. It is a thin client over the Fast Video Cataloger server's REST API: it holds
no state and every tool is one or two REST calls. Build with the .NET 10 SDK:
dotnet publish src/FvcMcpServer -c Release -r win-x64 --self-contained false -p:PublishSingleFile=trueReleased builds are code-signed by VideoStorm Sweden AB. Issues and pull requests are welcome; the code here follows the version in Fast Video Cataloger, so accepted changes ship with the next release.
License
MIT for the source in this repository, see LICENSE. Bundled packages: THIRD-PARTY-NOTICES.md. Fast Video Cataloger itself is commercial software with a free 30-day trial.
Available Tools
23 toolsadd_video_to_binAdd a video to a binA
Put a video into a bin. Needs an API key with the Editor role.
| Name | Required | Description | Default |
|---|---|---|---|
| binId | Yes | The bin's id. | |
| videoId | Yes | The video's catalog id. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only declare destructiveHint=false, so the description carries most of the burden. It adds genuinely useful context — the required Editor-role API key — but says nothing about idempotency, what happens if the video is already in the bin, or failure modes.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences, with the action front-loaded ahead of the permission note; both sentences earn their place. It is very terse, bordering on under-specified, but there is no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple two-parameter mutation with full schema coverage and no output schema, the description supplies the key missing piece (authorization requirement). Little more is needed for an agent to call it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, with both binId and videoId documented inline, so the baseline is 3. The description adds no format, range, or sourcing detail beyond what the schema already provides.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
'Put a video into a bin' states a specific verb and resource, and the parameter names (binId, videoId) confirm the exact relationship. It is clear enough to distinguish from read-oriented siblings like get_bin_videos, though it does not explicitly name an alternative.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives a prerequisite (an API key with the Editor role) but never states when to choose this tool versus alternatives such as create_bin or tag_video, nor any exclusions. Usage is only implied by the name.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
catalog_statsCatalog sizeARead-only
How many videos, actors and images the Fast Video Cataloger catalog holds. This is the user's local video library, not cloud storage or the file system. Useful as a first call to confirm the connection works and to judge how broad a search should be.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The readOnlyHint annotation already covers the safety profile, and the description adds real context beyond it: the counts come from the user's local video library and explicitly not from cloud storage or the file system, and the call doubles as a connectivity check. It does not describe the response shape or any cost of a large catalog, which keeps it out of the top band.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, each earning its place: what is counted, what scope it covers, and when to call it. The payload (counts) is front-loaded and nothing is repeated from the schema or annotations.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter read-only stats tool with no output schema, the description covers what is counted and the data scope, which is enough for correct invocation. It leaves the exact return field names and any error/empty-catalog behavior unstated, a minor gap given the tool's simplicity.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so there is no parameter semantics for the description to add; the baseline for a no-arg tool is 4. The description correctly implies the call is global with no filtering options.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names the exact resource (the Fast Video Cataloger catalog) and the exact quantities returned (counts of videos, actors and images), which is a specific verb+resource statement rather than a restatement of the title. It is immediately distinguishable from siblings like search_videos or get_video, which return entities rather than aggregate counts.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives explicit usage context: "Useful as a first call to confirm the connection works and to judge how broad a search should be." That tells the agent when to call it (startup, sizing a search) but names no alternative tool or exclusion, so it stops short of a full when/when-not comparison.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
create_binCreate a binA
Create a new bin to collect videos in. Returns the new bin including its id. Needs an API key with the Editor role.
| Name | Required | Description | Default |
|---|---|---|---|
| label | Yes | Name for the bin, e.g. 'trailer candidates'. | |
| parentId | No | Id of a parent bin, to nest this one inside it. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only declare destructiveHint=false, so the description adds genuinely new behavioral context: the authorization requirement (Editor role) and the fact that it returns the new bin including its id. It does not cover idempotency, label uniqueness, or nesting side effects for parentId, but the added auth and return details exceed the annotation baseline.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, front-loaded with the action and result, then the permission constraint. Every sentence carries information an agent needs; no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a two-parameter mutation with no output schema, the description correctly compensates by stating the return payload (new bin with id) and the auth requirement. The only omitted items are failure modes (duplicate label, invalid parentId), which would make it fully self-sufficient.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so both 'label' and 'parentId' are already documented with examples and purpose. The description adds no extra parameter semantics beyond the schema, which is the expected baseline rather than a gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Specific verb+resource: 'Create a new bin to collect videos in' immediately tells the agent what it produces and what a bin is for. It is distinguishable from siblings like list_bins and add_video_to_bin, though it never names a sibling explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description supplies a real prerequisite ('Needs an API key with the Editor role'), which tells the agent when the call can succeed. However, it gives no guidance on when to create a bin versus reusing one, nor any reference to alternative tools such as list_bins for checking existing bins.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_bin_videosGet a bin's videosARead-only
List the videos collected in one bin. A bin can hold the whole catalog, so this returns the first 50 by default - check totalCount and page with offset.
| Name | Required | Description | Default |
|---|---|---|---|
| binId | Yes | The bin's id. | |
| limit | No | Maximum number of videos. Defaults to 50. | |
| offset | No | Skip this many videos, for paging through a large bin. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With readOnlyHint=true already covering the safety profile, the description still adds real behavioral context: a bin can hold the whole catalog, so results are truncated to 50 by default and paging via offset is required. It does not mention ordering or what happens with an invalid binId.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, zero waste, with the scoping statement front-loaded and the pagination caveat immediately after. Every clause earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a three-parameter read tool with no output schema, the description covers the key trap (silent truncation at 50) and names the totalCount field the agent must inspect. It could be slightly more complete by noting there is no filtering here versus search_videos.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so binId, limit and offset are already documented; the description mostly restates the default-50 and paging behavior. It adds the rationale (bins can be catalog-sized) and points at totalCount, which is mildly useful but not new parameter semantics.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ("List the videos collected in one bin"), which cleanly separates it from get_video (single video), list_bins, and search_videos. An agent can tell what this returns without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explains pagination mechanics (default 50, check totalCount, page with offset) but gives no explicit when-to-use versus alternatives such as search_videos or list_videos_needing_indexing. Usage is implied by the bin-scoped framing rather than stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_scene_imageLook at a sceneARead-only
Fetch the actual image of a scene thumbnail so you can see what is in the shot. Get thumbnail ids from search_scenes, search_scenes_semantic or get_video_scenes.
| Name | Required | Description | Default |
|---|---|---|---|
| thumbnailId | Yes | The thumbnail's id. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, so the safety profile is covered. The description still adds value by clarifying that the call returns viewable image content (not a URL or metadata blob), which tells the agent what kind of payload to expect. It does not discuss image size, format, or failure behavior for stale ids.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences, no redundancy, and the core purpose is front-loaded ahead of the input-sourcing hint. Every clause earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a one-parameter read tool with no output schema, the description covers purpose, the nature of the return value, and how to supply the input. Nothing an agent needs to invoke it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% for the single parameter, so the baseline is 3. The description goes beyond the schema's bare 'The thumbnail's id.' by telling the agent where to obtain a valid id, which is the key semantic gap for this argument.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Fetch the actual image of a scene thumbnail') and adds the agent-facing rationale ('so you can see what is in the shot'). This clearly separates it from siblings like search_scenes or get_video_scenes, which return ids/metadata rather than the image itself.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly tells the agent the prerequisite step and names three concrete sources for thumbnailId (search_scenes, search_scenes_semantic, get_video_scenes). It lacks any when-not-to-use or cost/latency caveat, but the usage context is unambiguous.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_videoGet a videoBRead-only
Get one catalog entry by its id, including title, description, path, length and rating.
| Name | Required | Description | Default |
|---|---|---|---|
| videoId | Yes | The video's catalog id. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
readOnlyHint=true already tells the agent this is a safe read, so the description isn't carrying the safety burden. It adds the field list, which is genuinely useful, but says nothing about behavior when the id doesn't exist (error vs. empty) or how missing fields are represented.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One sentence, front-loaded with the verb and resource, with no filler. The trailing field enumeration is slightly listy but earns its place as a stand-in for the absent output schema.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a single-parameter read tool with no output schema, the description supplies the essentials: what it returns and how it is keyed. Only the not-found/error behavior is unstated, which is a minor gap given the readOnlyHint.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% with a single integer parameter already documented as 'The video's catalog id.' The description's 'by its id' adds no format, range, or validity detail beyond the schema, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (get) and resource (one catalog entry) plus the lookup key (id), and enumerates the returned fields (title, description, path, length, rating). This distinguishes it from the deeper siblings like get_video_scenes or get_video_transcript, though it doesn't name any alternative explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'by its id' weakly implies you need a known id from a prior lookup, but there is no statement of when to use this versus search_videos for discovery or the get_video_* siblings for related data. No exclusions or prerequisites are given.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_video_actorsGet a video's castCRead-only
List the people recorded as appearing in a video.
| Name | Required | Description | Default |
|---|---|---|---|
| videoId | Yes | The video's catalog id. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The readOnlyHint annotation already establishes this as a safe read operation. The description adds no behavioral context beyond that — no mention of ordering, pagination, whether an empty list is possible, or what happens for videos with no recorded cast.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single efficient sentence with no wasted words, front-loaded with the verb and resource. It is appropriately sized, though it is arguably under-specified rather than optimally concise.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple one-parameter read tool with annotations covering safety and no output schema, the description is minimally adequate. It does not address return shape, ordering, or the empty-result case, leaving small but real gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% with a single parameter, so the schema already documents videoId fully. The description adds no format or semantic detail beyond what the schema provides, which matches the baseline 3 for high coverage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (List) and resource (people recorded as appearing in a video), which is clear and unambiguous. It does not explicitly differentiate itself from the sibling search_actors, but the scoping to a single video makes the distinction inferable.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to use this tool versus search_actors, get_video_scenes, or other siblings, nor any stated prerequisites. The only implied usage comes from the required videoId parameter.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_video_scenesGet a video's scenesARead-only
List the scene thumbnails of one video in timecode order. Use this to see how a video is structured before deciding which moment to look at. Returns totalCount so you can tell whether there are more than you asked for.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | Maximum number of scenes. Defaults to 100. | |
| offset | No | Skip this many scenes, for paging through a long video. | |
| videoId | Yes | The video's catalog id. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, so safety is covered; the description adds ordering behavior (timecode order) and the key return signal (totalCount reveals whether more scenes exist beyond the requested page). That pagination-relevant disclosure is genuinely useful since no output schema exists, though offset/limit mechanics are left to the schema.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight sentences with no filler, and the primary action is front-loaded ahead of the usage hint and return-value note. Every sentence contributes, though the return-value sentence could be trimmed slightly.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple three-parameter read tool with annotations covering the safety profile and a fully documented schema, the description supplies the missing pieces: ordering, purpose, and the totalCount pagination signal that no output schema provides. Nothing essential is absent.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so videoId, limit, and offset are already fully documented. The phrase "more than you asked for" hints at the limit parameter but adds no syntax, defaults, or edge-case guidance beyond the schema, so the baseline of 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ("List the scene thumbnails of one video") and adds scope ("in timecode order"), which distinguishes it from collection-wide siblings like search_scenes. It does not name those siblings explicitly, so differentiation is implied rather than stated.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
"Use this to see how a video is structured before deciding which moment to look at" gives a concrete when-to-use scenario, making the tool's role clear relative to moment-level tools like get_scene_image. No when-not-to-use or named alternative is provided, so it stops short of the top band.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_video_tagsGet a video's tagsBRead-only
List the keywords a video is tagged with.
| Name | Required | Description | Default |
|---|---|---|---|
| videoId | Yes | The video's catalog id. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, so the safety profile is covered without the description. The description adds only that the result is 'keywords', not pagination, ordering, or whether absent tags yield an empty list, so it does not go beyond the structured fields.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with no wasted words. It is efficient, though it nearly restates the title 'Get a video's tags' rather than adding new information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a one-parameter read-only lookup with full schema coverage and readOnlyHint annotations, the description is sufficient to call the tool correctly. Minor gaps remain (return shape/ordering), but no output schema exists and none of the essentials are missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
There is a single parameter and schema coverage is 100%, with the schema documenting 'videoId' as the catalog id. The description adds no syntax, format, or edge-case detail beyond what the schema already provides, so baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('List') and resource ('keywords a video is tagged with'), making it clear this is a read of tag metadata. However, it does not name or distinguish itself from close siblings such as list_keywords, get_video_actors, or tag_video/untag_video, leaving differentiation to inference.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No explicit when-to-use guidance, no prerequisites, and no alternatives named. Usage is only implied by the fact that the tool lists a video's tags rather than another video's data.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_video_transcriptGet a transcriptARead-only
Get the transcript of a video as timed lines. Empty if the video has not been transcribed. A feature-length transcript is very long, so this returns the first 200 lines by default - check totalCount and page with offset if you need the rest.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | Maximum number of transcript lines. Defaults to 200. | |
| offset | No | Skip this many lines, for reading further into a long transcript. | |
| videoId | Yes | The video's catalog id. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, so the safety profile is covered. The description adds real behavioral context beyond that: the empty-transcript condition, the default 200-line truncation, and the existence of a totalCount field to detect truncation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, front-loaded with the core purpose, then the edge case, then the pagination rule. Every sentence carries information and none is wasted.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description does the work of naming the returned unit (timed lines) and the totalCount field, and it covers the empty and truncated cases. Only the sibling-routing question is left open, which is a minor gap for a simple read tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3. The description goes slightly further by explaining why limit/offset exist (feature-length transcripts are very long) and how they interact with totalCount, which is not documented anywhere in the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Get the transcript of a video') plus the return shape ('as timed lines'), which is more than a restatement of the title. It does not, however, distinguish itself from the sibling search_transcripts, leaving the agent to infer which one to call.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives clear guidance on pagination ('check totalCount and page with offset if you need the rest') and notes the empty-transcript case, which is implied usage context. But it never states when to prefer this over search_transcripts or get_video, so alternative selection is left to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
index_videoIndex a videoA
Ask the server to index a video that is in the catalog but has no scene thumbnails yet. Extracts thumbnails and duration, plus whatever else the server is configured for such as transcription. This takes minutes and runs in the background - poll get_video_scenes to follow it. Needs the Editor role.
| Name | Required | Description | Default |
|---|---|---|---|
| videoId | Yes | The video's catalog id. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only declare destructiveHint=false, so the description carries the real behavioral burden and does so well: it discloses that the operation is asynchronous and long-running ('takes minutes and runs in the background'), what it extracts (thumbnails, duration, optionally transcription), and an authorization requirement (Editor role). These are exactly the traits an agent cannot infer from structured fields.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, each earning its place: purpose/precondition, what it produces, then runtime and follow-up. Front-loaded and free of filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, but the description compensates by explaining the async nature, the polling mechanism, and the role requirement, which is everything an agent needs to invoke and track this correctly. Nothing material is missing for a one-parameter tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
There is a single parameter with 100% schema description coverage, so the schema already documents videoId fully and the description adds nothing about it. Baseline 3 is appropriate when the schema does all the parameter work.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource (index a video) and scopes it precisely to catalog videos lacking scene thumbnails. It also distinguishes itself from the sibling used afterward (get_video_scenes) and implicitly from list_videos_needing_indexing, so an agent can route without opening a schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives an explicit precondition for use ('in the catalog but has no scene thumbnails yet') and names the correct alternative action for a different need (poll get_video_scenes to follow progress). The only minor omission is a direct pointer to list_videos_needing_indexing for discovery, but the when-to-use guidance is unambiguous.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_binsList binsARead-only
List the bins in the user's Fast Video Cataloger catalog. A bin is a user-made collection of videos, like a folder of shortlisted clips.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, so the safety profile is covered by structured data. The description adds domain context about what a bin is but says nothing about whether all bins are returned, ordering, or pagination behavior — and with no output schema, the shape of the result is left to inference.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences with no filler: the action comes first and the clarifying definition of 'bin' follows immediately. Nothing could be removed without losing meaning.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a trivial, zero-argument read-only list tool, the description covers what is listed and why the resource exists, which resolves the main ambiguity (what a 'bin' is). The one residual gap is that with no output schema, the fields of a returned bin (id, name, etc.) are never described.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so per the baseline there is nothing for the description to disambiguate. The sentence explaining what a bin is provides conceptual grounding for the returned entities, which is the only parameter-adjacent value available.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('List') and resource ('bins') and grounds the resource in its domain ('user-made collection of videos, like a folder of shortlisted clips'), which separates it from sibling concepts like keywords, tags, and actors. It does not explicitly distinguish itself from nearby siblings such as get_bin_videos or create_bin, so it falls short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is only implied by the verb 'List' — there is no statement of when to call this versus get_bin_videos (contents of one bin) or create_bin. For a zero-parameter, read-only enumeration tool this is survivable, but no explicit alternative or condition is named.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_keywordsList the catalog's keywordsARead-only
List every keyword that exists in the whole Fast Video Cataloger catalog, rather than the keywords on one video. Use it to find out what you can actually search for before guessing - a keyword that is not in the catalog simply returns nothing. This is the vocabulary only, with no counts, and it mixes keywords the user typed with ones written automatically by object detection, so a catalog can list things like 'zebra' or 'toaster' that appear nowhere in the footage. Both scopes can contain either kind. To find out whether a keyword is really used, search for it.
| Name | Required | Description | Default |
|---|---|---|---|
| scope | No | Which keywords to list: 'videos' for keywords on videos, 'scenes' for keywords on scene thumbnails. Defaults to videos. | videos |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only declare readOnlyHint=true; the description adds substantive behavioral context beyond that: the result is vocabulary only with no counts, it mixes user-typed and object-detection-generated keywords, and it can contain terms (e.g. 'zebra', 'toaster') that appear nowhere in the footage. This tells the agent exactly what the output can and cannot be trusted to mean.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four sentences, front-loaded with the catalog-wide scope before the caveats, and every sentence carries information. Slightly long: the closing 'search for it' advice restates the earlier 'not in the catalog returns nothing' warning, so there is minor redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description does the work of explaining what comes back (a vocabulary list, no counts, mixed provenance). For a single optional-parameter, read-only listing tool, nothing an agent needs to call it correctly or interpret the result is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the 'videos'/'scenes' scope values are already documented, giving a baseline of 3. The description adds a genuinely non-obvious semantic point — that both scopes can contain either user-typed or auto-detected keywords — which clarifies the parameter's meaning beyond the schema text.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('List every keyword that exists in the whole Fast Video Cataloger catalog') and immediately scopes it against the per-video alternative ('rather than the keywords on one video'). An agent can distinguish this from sibling keyword tools like get_video_tags without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit when-to-use context ('find out what you can actually search for before guessing') and a when-not-to-trust-this signal ('a keyword that is not in the catalog simply returns nothing'), plus a redirect ('To find out whether a keyword is really used, search for it'). The alternative is described by behavior rather than named as a specific sibling tool, so it falls just short of the top tier.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_videos_needing_indexingList unindexed videosARead-only
List catalog entries that have no scene thumbnails or no known length yet. Pass each id to index_video to work through them. A newly added folder can leave thousands pending, so this returns the first 50 by default - check totalCount and page with offset to walk the rest.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | Maximum number of entries. Defaults to 50. | |
| offset | No | Skip this many entries, for working through a long backlog. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare readOnlyHint=true, so safety is already covered. The description adds real behavioral context beyond that: the default cap of 50, the reason for it (a new folder can leave thousands pending), and the totalCount/offset walk pattern. It does not describe entry shape or ordering, which keeps it from a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, front-loaded with the definition of what gets listed, then the next action, then the pagination caveat. No sentence is redundant with the schema or the title.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a two-parameter, zero-required read tool with no output schema, the description covers identification, follow-up action, and pagination behavior. It even names totalCount, which is otherwise undiscoverable since no output schema exists.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3. The description still adds value by tying the parameters to a workflow: the 50 default is explained as a backlog-size protection, and offset is framed as the mechanism for walking a long queue, which is more than the schema's bare 'skip this many entries'.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (List) plus the exact filter criterion that defines membership: entries with no scene thumbnails or no known length. This is distinguishable from siblings like search_videos, get_video, or catalog_stats without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly routes the agent onward ('Pass each id to index_video to work through them'), which is the key when-to-use relationship in this toolset. It lacks an explicit when-not-to-use or a named contrasting sibling (e.g. search_videos for already-indexed content), so it stops short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
search_actorsSearch actorsARead-only
Search the people recorded in the user's Fast Video Cataloger catalog, by name or by the tags on them. These are the cast of their own footage, not contacts.
| Name | Required | Description | Default |
|---|---|---|---|
| tags | No | Comma-separated tags the actor must have. | |
| limit | No | Maximum number of people to return. Defaults to 50. | |
| query | No | Free text to match against actor names. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
readOnlyHint=true already tells the agent this is a safe read. The description adds worthwhile domain context that results come from the user's own catalog and are the cast of their footage rather than contacts. It says nothing about result size, ordering, or empty-match behavior beyond what annotations cover.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences with zero filler, and the core purpose is front-loaded. Every clause carries meaning, with the 'not contacts' note earning its place as a disambiguator.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple, annotation-covered read tool with three fully documented parameters and no output schema, the description covers what the tool returns and how it matches. Only the absence of guidance on result ordering/pagination keeps it short of full completeness.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so query, tags, and limit are already fully documented in the schema. The description's 'by name or by the tags on them' loosely maps to the query and tags parameters but adds no format or syntax detail beyond the schema, so baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Search) and resource (people/actors in the Fast Video Cataloger catalog) plus the matching basis (name or tags). The clarifying phrase 'not contacts' usefully disambiguates the domain, but it does not distinguish this from the sibling get_video_actors, which also deals with actors.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage by saying it searches by name or by tags, but never states when to choose this over get_video_actors or other search tools. No exclusions or prerequisites are given, leaving the routing decision to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
search_scenesSearch scenesARead-only
Search individual scenes inside the user's Fast Video Cataloger videos, rather than whole videos. This is how you find the moment inside a video: each result carries a thumbnail id, the video it belongs to, and its timecode. Pass the thumbnail id to get_scene_image to look at it.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | Maximum number of results. Defaults to 40. | |
| query | No | Free text to match against scene keywords. | |
| videoId | No | Restrict the search to one video by its id. | |
| keywords | No | Comma-separated keywords. A scene matches if it is tagged with ANY of them, not all of them - listing more keywords widens the search rather than narrowing it. Scene keywords come from object detection, so they are concrete things that appear in the frame ('person', 'laptop', 'car'), not themes or descriptions. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only declare readOnlyHint=true, so the description adds valuable context by disclosing the result fields (thumbnail id, video, timecode) and the chaining step to get_scene_image. It does not cover auth or rate limits, but the lower bar with annotations makes this sufficient.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tightly written sentences, front-loaded with the scope distinction, then result shape, then next action. No waste.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given no output schema, the description adequately explains what results contain and how to use them. It could mention the semantic-search sibling to fully disambiguate, but is otherwise complete for this search tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all four parameters are already clearly documented. The description does not add meaning beyond what the schema provides, making the baseline 3 appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (search) and resource (individual scenes) with scope (inside the user's videos) and explicitly contrasts with whole-video search. It does not name the semantic-search sibling, so it falls short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear context for when to use it (to find the moment inside a video) and an implicit exclusion (rather than whole videos), plus the next tool to call (get_scene_image). It does not mention alternatives like search_scenes_semantic.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
search_scenes_semanticFind scenes by descriptionARead-only
Find scenes by describing what is in the picture - 'a dog running on a beach', 'close-up of hands typing', 'city skyline at night'. Nothing has to be tagged: every indexed scene is ranked by how well it matches the words, so use this when the keywords search_scenes needs do not cover what you are after, and use similarTo to find footage that looks like a scene you already have (the source scene itself is left out of those results). Results come best first, each with the thumbnail id, the video, the time (seconds and hh:mm:ss) and a relevance score. Judge by the order, not the number: for a description a good match scores only about 0.15-0.30, while scenes that look like a given one score 0.8 and above, and scores are comparable only within one search. Pass a thumbnail id to get_scene_image to check what a scene really shows. If the answer says the model or the index is missing, tell the user what to do; do not retry.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | How many of the best matches to return. Defaults to 20, maximum 200. | |
| query | No | What the scene should show, in plain words. Ignored when similarTo is given. | |
| videoId | No | Restrict the search to one video by its id. | |
| similarTo | No | A thumbnail id: rank scenes by how much they look like this one instead of by a description. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only declare readOnlyHint=true, so the description carries most of the burden and does so richly: nothing needs tagging, every indexed scene is ranked, the similarTo source scene is excluded from results, and the error path is specified. It even discloses the score semantics (0.15-0.30 for descriptions vs 0.8+ for visual similarity) and warns that scores are comparable only within a single search – context an agent cannot get from the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
It is a dense single paragraph but front-loads the clearest signal (the concrete example queries) before the routing and scoring rules. Every sentence carries a distinct instruction – mode selection, exclusion rule, result format, score interpretation, verification, error handling – though the run-on structure makes it slightly heavy to parse at a glance.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description must describe returns and it does: results come best-first with thumbnail id, video, time in seconds and hh:mm:ss, and relevance score. Combined with the error-handling instruction and the cross-reference to get_scene_image, an agent has everything needed to call and interpret this tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, and the description adds real meaning on top: query is ignored when similarTo is supplied and similarTo takes a thumbnail id and ranks by visual likeness rather than description. It explains return-field semantics (thumbnail id, video, time, relevance score) rather than leaving them implicit.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb+resource ('Find scenes by describing what is in the picture') and immediately grounds it with three concrete query examples. It also explicitly separates the description mode from keyword search_scenes and from the similarTo visual-similarity mode, so an agent can distinguish it from siblings without opening a schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It names the condition that selects this tool ('use this when the keywords search_scenes needs do not cover what you are after') and the condition that selects the alternative mode ('use similarTo to find footage that looks like a scene you already have'). It even routes further downstream ('Pass a thumbnail id to get_scene_image to check what a scene really shows') and states a when-not ('do not retry' on missing model/index).
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
search_transcriptsSearch spoken wordsARead-only
Search what is said out loud in the user's Fast Video Cataloger videos. Returns the matching transcript lines themselves - each with the id of the video it belongs to (videoFileID) and its start and end time in seconds, so you can go straight to the moment. Use get_video_transcript to read the lines around a hit. Only videos that have been transcribed are searched.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | Maximum number of matching lines to return. Defaults to 25, maximum 200. | |
| query | Yes | The words to look for in the transcripts. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only declare readOnlyHint, so the description usefully adds the return shape (matching lines with videoFileID and start/end seconds) and the index limitation that only transcribed videos are searched. No pagination or ranking behavior is disclosed, so it stops short of a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight sentences, front-loaded with the core purpose, then return value, then the alternative tool and the indexing caveat. Every clause carries information; nothing is wasted.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description carries the burden of describing results and does so precisely (line content, videoFileID, start/end seconds). Combined with the transcription precondition and the follow-up tool, an agent has everything needed to call and use it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% (limit's default/max and query's meaning are both documented in the schema). The description adds no syntax, matching, or format detail beyond the schema, so baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource with scope: 'Search what is said out loud in the user's ... videos'. The phrase 'said out loud' immediately differentiates it from search_videos, search_scenes, and search_actors without the agent needing to consult any other schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Names a follow-up tool explicitly ('Use get_video_transcript to read the lines around a hit') and states a real precondition ('Only videos that have been transcribed are searched'). It does not give explicit when-not guidance against the other search_* siblings, but the context is clear.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
search_videosSearch videosARead-only
Search the user's Fast Video Cataloger catalog - their local library of video files, not cloud storage or the file system. Use this for any question about their videos, footage or clips. Use 'query' for a free text search over titles, descriptions and keywords, or 'keywords' for a comma-separated keyword match. Returns catalog entries including each video's id, title, path and length.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | Maximum number of results. Defaults to 25. | |
| query | No | Free text to search for. Leave empty to browse with the other filters. | |
| actors | No | Comma-separated actor ids, to find the videos a person appears in. Get ids from search_actors. This is the reverse of get_video_actors. | |
| offset | No | Skip this many results, for paging through a large set. | |
| keywords | No | Comma-separated keywords, e.g. 'launch,apollo'. By default a video matches if it carries ANY of them; set matchAllKeywords to require every one. | |
| minRating | No | Only videos rated at least this (1-5). | |
| matchAllKeywords | No | Require EVERY keyword rather than any one of them. Use this to narrow a search, e.g. keywords 'beach,sunset' with matchAllKeywords true finds only videos that are both. A keyword that does not exist in the catalog is ignored rather than making the result empty, so check list_keywords if a narrowed search looks too broad. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only declare readOnlyHint=true, so the safety profile is covered but little else. The description adds real behavioral content the annotations lack: it enumerates the returned fields (id, title, path, length), which matters because there is no output schema. It does not discuss paging behavior beyond what the schema states.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, each carrying distinct load: scope, routing, then return shape. The domain constraint is front-loaded and nothing is repeated from the schema or annotations.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only, fully schema-documented search tool with no output schema, the description covers scope, usage, mode selection and return fields. An agent has everything needed to call it correctly without further exploration.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3, but the description adds field-level meaning for 'query' (free text over titles, descriptions and keywords) that the schema's generic 'Free text to search for' does not convey, plus a plain-language gloss of the query-vs-keywords choice. The remaining five parameters get no added meaning from the description.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Names a specific verb and resource ('search the user's Fast Video Cataloger catalog') and immediately bounds the domain: local video library, not cloud storage or the file system. That scoping also distinguishes it from siblings like search_transcripts, search_scenes and search_actors without the agent opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit context ('use this for any question about their videos, footage or clips') and routes between the two main search modes: 'query' for free text, 'keywords' for keyword matching. It does not name sibling tools or say when a different search tool is the better choice, so it stops short of full when/when-not guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
tag_sceneTag a sceneA
Add keywords to a single scene thumbnail rather than the whole video. Needs an API key with the Editor role.
| Name | Required | Description | Default |
|---|---|---|---|
| tags | Yes | Comma-separated keywords to add. | |
| thumbnailId | Yes | The thumbnail's id. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only set destructiveHint=false; the description adds that an API key with the Editor role is required, clarifying this is a privileged write operation. It does not cover idempotency or duplicate handling, but the auth context is valuable beyond annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, front-loaded with purpose and scope, then the auth prerequisite. No wasted words.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple two-parameter tagging tool with full schema coverage and a destructiveHint annotation, the description covers purpose, scope, and auth. Return values aren't needed without an output schema, so it is complete enough.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, with both parameters documented. The description adds no syntax or format details beyond the schema's 'comma-separated keywords' and 'thumbnail's id', so baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb 'Add keywords' and resource 'single scene thumbnail', explicitly distinguishing from tagging the whole video. An agent can tell it apart from sibling tag_video without opening schemas.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides clear context for use—tagging an individual scene thumbnail rather than the whole video—but does not name the alternative tool (tag_video) or state exclusions. The implied usage is clear, yet explicit routing is absent.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
tag_videoTag a videoA
Add keywords to a video. Existing tags are kept. Needs an API key with the Editor role.
| Name | Required | Description | Default |
|---|---|---|---|
| tags | Yes | Comma-separated keywords to add, e.g. 'trailer,apollo'. | |
| videoId | Yes | The video's catalog id. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only declare destructiveHint=false, so the description adds real value: it clarifies tags are appended, not overwritten, and that an API key with the Editor role is required. It stops short of covering duplicate-tag handling or any rate limits.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, front-loaded with the action, then the additive constraint, then the permission requirement. Every sentence carries distinct information with no padding.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema, so return values need not be explained, and the key behavioral facts (additive tagging, Editor role) are present for this simple 2-parameter tool. Only edge cases like duplicate tags or invalid videoId are unaddressed.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, with both videoId and tags fully documented including the comma-separated format. The description adds no parameter detail beyond the schema, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ("Add keywords to a video") and, via "Existing tags are kept," distinguishes its additive semantics from the sibling untag_video and from any replace-style tagging behavior. An agent can pick this over untag_video without opening either schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The additive-only behavior gives implicit guidance on when to use this versus untag_video, but no sibling is named and there are no explicit when-not conditions. Usage is inferable rather than stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
untag_sceneRemove a scene's keywordsADestructive
Remove keywords from a single scene thumbnail - the counterpart to tag_scene. Only the named keywords are removed from that one scene; other scenes and the video's own tags are untouched. A keyword that no scene uses any more also disappears from the catalog's scene keyword list. Returns what was removed and what did not match. Needs an API key with the Editor role.
| Name | Required | Description | Default |
|---|---|---|---|
| tags | Yes | Comma-separated keywords to remove, named exactly as they appear on the scene. | |
| thumbnailId | Yes | The thumbnail's id. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Goes well beyond the destructiveHint annotation by disclosing cascade behavior (a keyword unused by any scene disappears from the catalog's scene keyword list), the response shape (what was removed vs. what did not match), and an authorization prerequisite (API key with the Editor role). These are exactly the operational facts the annotations cannot convey.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four compact sentences, front-loaded with the action and the tag_scene relationship before secondary details. Every sentence carries a distinct fact (scope, cascade, return value, auth), though it is slightly denser than strictly necessary.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, so the description correctly covers return values, plus cascade side effects and the auth role required. For a 2-parameter destruction tool whose safety hint is only destructiveHint=true, nothing an agent needs to invoke it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and both parameters are documented in the schema, so the baseline is 3. The description adds matching semantics (only the named keywords, other scenes untouched) and hints at partial-match handling in the return, but no format or syntax detail for the parameters themselves.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Remove keywords from a single scene thumbnail') and immediately scopes it as 'the counterpart to tag_scene'. The scope is also distinguished from untag_video by noting 'other scenes and the video's own tags are untouched', so an agent can separate it from siblings without opening a schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Clear context: it operates on one scene's keywords and is the counterpart to tag_scene, and it distinguishes itself from video-level tagging. It stops short of an explicit 'use X when Y' routing sentence, but the boundary conditions are stated with enough precision to select the right tool.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
untag_videoRemove a video's keywordsADestructive
Remove keywords from a video - the counterpart to tag_video, for undoing a tag that was added in error. Name the keywords exactly as get_video_tags reports them; only those are removed and every other tag on the video is left alone. Other videos keep the keyword, but a keyword that no video uses any more also disappears from the catalog's keyword list. Returns what was removed and what did not match. Needs an API key with the Editor role.
| Name | Required | Description | Default |
|---|---|---|---|
| tags | Yes | Comma-separated keywords to remove, named exactly as they appear on the video. | |
| videoId | Yes | The video's catalog id. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only supply destructiveHint=true; the description goes well beyond that by stating precisely what is destroyed (only the named keywords, all other tags untouched), the catalog-level side effect (a keyword unused by any video is removed from the keyword list), that other videos keep the keyword, and the required Editor-role API key.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four tight sentences, front-loaded with purpose and the sibling relationship, then scoping, then side effects and return/auth details. No filler sentences; every clause carries operational information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description still covers the return value ('what was removed and what did not match'), the auth prerequisite, the destructive scope, and the catalog side effect, leaving nothing an agent needs in order to call it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so both parameters are already documented, including the comma-separated format and exact-name requirement. The description's 'name the keywords exactly as get_video_tags reports them' reinforces the source of truth but adds little beyond the schema, so it sits at the high-coverage baseline.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Remove keywords from a video') and explicitly positions itself as 'the counterpart to tag_video', which disambiguates it from the sibling that shares a near-identical name.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives a clear triggering scenario ('for undoing a tag that was added in error') and names the counterpart tool tag_video. It does not mention untag_scene, the scene-level sibling, so an agent working on scene tags gets no routing hint, but the primary when-to-use is unambiguous.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
23 tool updates
v10.4.1- First observed
add_video_to_bin - First observed
catalog_stats - First observed
create_bin - First observed
get_bin_videos - First observed
get_scene_image - First observed
get_video - First observed
get_video_actors - First observed
get_video_scenes - First observed
get_video_tags - First observed
get_video_transcript - First observed
index_video - First observed
list_bins - First observed
list_keywords - First observed
list_videos_needing_indexing - First observed
search_actors - First observed
search_scenes - First observed
search_scenes_semantic - First observed
search_transcripts - First observed
search_videos - First observed
tag_scene - First observed
tag_video - First observed
untag_scene - First observed
untag_video
TDQS
Scored across 23 tools
Each tool targets a clearly distinct resource or action, with little overlap. The various search tools (search_videos, search_scenes, search_transcripts, search_scenes_semantic) are well differentiated by the entity they search and the method used. Tagging and untagging tools are split by scope (video vs scene), and bin tools are distinct from video and scene tools.
Almost all tools follow a consistent verb_noun snake_case pattern (e.g., get_video, search_scenes, tag_video, list_bins). The only notable deviation is catalog_stats, which uses a noun_noun form without a verb. This is a minor inconsistency in an otherwise predictable naming scheme.
With 23 tools, the set falls into the borderline-heavy range for an MCP server. While the domain (video catalog with bins, scenes, transcripts, actors, tags, and indexing) is broad, some tools could potentially be consolidated (e.g., add/remove tag pairs), and the count feels high relative to the missing lifecycle operations.
The toolset covers search, tagging, bins, indexing, and transcript/scene access well, but lacks basic update and delete operations for core entities: there is no update_video, delete_video, delete_bin, or remove_video_from_bin. These gaps mean the surface is not a complete CRUD/lifecycle set for the catalog.
Maintenance
Related MCP Connectors
- VivuOAuthio.github.vivuai
Search your video library with natural language and retrieve relevant moments through MCP.
Your YouTube library in Claude, ChatGPT and Cursor: transcripts, breakdowns, summaries, search.
Your YouTube library in Claude, ChatGPT and Cursor: transcripts, breakdowns, summaries, search.
Search your Flashback video library with natural language to instantly find relevant moments. Get…
Related MCP Servers
- AlicenseNot gradedqualityDmaintenanceEnables video processing operations (speed adjustment, keyframe optimization, concatenation, and file management) using FFmpeg via Claude Desktop.89 npm3MIT
- FlicenseBqualityDmaintenanceMCP server that enables video analysis capabilities to Claude, including frame extraction, scene detection, and video metadata retrieval.8-
- FlicenseNot gradedqualityDmaintenanceBridges Claude and video content by extracting keyframes and transcribing audio, enabling Claude to analyze video files.-
- AlicenseAqualityCmaintenanceEnables natural language search of local photo archives using AI-powered semantic understanding, with integration into Claude Desktop via the Model Context Protocol.42MIT