Skip to main content
Glama
frelo

Fast Video Cataloger

Official

Find scenes by description

search_scenes_semantic
Read-only

Find scenes by describing what they show, like 'a dog running on a beach', without needing tags. Alternatively, rank scenes by visual similarity to a thumbnail id.

Instructions

Find scenes by describing what is in the picture - 'a dog running on a beach', 'close-up of hands typing', 'city skyline at night'. Nothing has to be tagged: every indexed scene is ranked by how well it matches the words, so use this when the keywords search_scenes needs do not cover what you are after, and use similarTo to find footage that looks like a scene you already have (the source scene itself is left out of those results). Results come best first, each with the thumbnail id, the video, the time (seconds and hh:mm:ss) and a relevance score. Judge by the order, not the number: for a description a good match scores only about 0.15-0.30, while scenes that look like a given one score 0.8 and above, and scores are comparable only within one search. Pass a thumbnail id to get_scene_image to check what a scene really shows. If the answer says the model or the index is missing, tell the user what to do; do not retry.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
limitNoHow many of the best matches to return. Defaults to 20, maximum 200.
queryNoWhat the scene should show, in plain words. Ignored when similarTo is given.
videoIdNoRestrict the search to one video by its id.
similarToNoA thumbnail id: rank scenes by how much they look like this one instead of by a description.

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observedv10.4.1

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only declare readOnlyHint=true, so the description carries most of the burden and does so richly: nothing needs tagging, every indexed scene is ranked, the similarTo source scene is excluded from results, and the error path is specified. It even discloses the score semantics (0.15-0.30 for descriptions vs 0.8+ for visual similarity) and warns that scores are comparable only within a single search – context an agent cannot get from the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

It is a dense single paragraph but front-loads the clearest signal (the concrete example queries) before the routing and scoring rules. Every sentence carries a distinct instruction – mode selection, exclusion rule, result format, score interpretation, verification, error handling – though the run-on structure makes it slightly heavy to parse at a glance.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description must describe returns and it does: results come best-first with thumbnail id, video, time in seconds and hh:mm:ss, and relevance score. Combined with the error-handling instruction and the cross-reference to get_scene_image, an agent has everything needed to call and interpret this tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3, and the description adds real meaning on top: query is ignored when similarTo is supplied and similarTo takes a thumbnail id and ranks by visual likeness rather than description. It explains return-field semantics (thumbnail id, video, time, relevance score) rather than leaving them implicit.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb+resource ('Find scenes by describing what is in the picture') and immediately grounds it with three concrete query examples. It also explicitly separates the description mode from keyword search_scenes and from the similarTo visual-similarity mode, so an agent can distinguish it from siblings without opening a schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It names the condition that selects this tool ('use this when the keywords search_scenes needs do not cover what you are after') and the condition that selects the alternative mode ('use similarTo to find footage that looks like a scene you already have'). It even routes further downstream ('Pass a thumbnail id to get_scene_image to check what a scene really shows') and states a when-not ('do not retry' on missing model/index).

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.