Bilibili Video Research
This server turns Bilibili links and local videos into evidence-aware research reports using language, vision, or multimodal analysis.
Research any public Bilibili URL (including b23.tv short links) with a custom question.
Choose
languagemode to use captions or ASR for what is said.Choose
visionmode to analyze silent video frames for on-screen interfaces, code, charts, labels, and actions.Choose
multimodalmode to combine narration and visual evidence.Restrict analysis to a specific time window with
start_secondsandend_seconds.Set source quality (standard/high/custom resolution and fps) for Bilibili downloads, with error/warn fallback behavior.
Control analysis detail with
media_detail(low coarse pass vs default close inspection).Optionally include sampled comments as untrusted community context, clearly separated from verified evidence.
Analyze a local video visually after removing its audio track (
analyze_video).Inspect a precise audio-free window from a local video (
inspect_video_window).Every result emits provenance and a deterministic
VIDEO CONTEXTblock with public metadata, tags, community status, and source quality details.
Provides tools for analyzing public Bilibili videos, including extracting captions, visual content, audio, and metadata from Bilibili URLs, with options for language, vision, and multimodal analysis.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@Bilibili Video ResearchInspect the time window 0:45-1:30 of my local video for visual evidence"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
Bilibili Video Research
Turn a Bilibili link into a research report that separates what came from public metadata, captions or ASR, video frames, and untrusted community context. Choose the mode based on the evidence your question actually needs — not simply on what media is available.
Every result also emits a deterministic VIDEO CONTEXT block containing public metadata and an explicit community-context status. The selected mode controls media evidence only; it does not remove the video context.
Demo Video
Project demo on Bilibili: Let an Agent help you research videos
Related MCP server: video-analyzer
Ecosystem
Glama listing — inspect the indexed server, tool schemas, and quality metadata.
Before you start
This is a local MCP server, but public metadata, archive tags, optional sampled comments, and selected media or text evidence may be sent to the provider configured in .env. Provider requests may consume API balance or subscription credits, and the selected provider's pricing and data terms apply. For a first run, use a short public video or an explicit source window; review Data and access boundary before using sensitive content.
What it does
Mode | Uses | Excludes | Best for |
| Bilibili captions when available; otherwise the selected provider's transcription/ASR path ( | Video-frame inference | Project recommendations, tutorials, and claims made by the presenter |
| StepFun | Audio and background music | Interfaces, workflows, experiments, objects, and silent demonstrations |
| StepFun | Nothing by default | Questions that genuinely require both narration and what is shown |
language is the intended default when a request only asks what a video says.
vision is the deliberate choice when the answer lives in the pixels.
What a result looks like
Ask the MCP tool a focused question:
analyze_bilibili_video({
url: "https://www.bilibili.com/video/BV...",
question: "What quantitative research framework is shown on screen?",
mode: "vision",
media_detail: "default",
source_quality: {
profile: "high",
on_unavailable: "error"
},
include_comments: false,
start_seconds: 0,
end_seconds: 321
})The response begins with provenance, then a deterministic VIDEO CONTEXT block, before the natural-language analysis:
RESEARCH PROVENANCE
{
"mode": "language",
"metadata": "bilibili_api",
"language": "stepfun_asr",
"visual": "none",
"community": "disabled",
"community_status": "disabled",
"tags_status": "present",
"timestamps": "none",
"source_quality": {
"profile": "standard",
"requested_resolution": "1080p",
"requested_fps": 30,
"on_unavailable": "error",
"status": "matched",
"actual_resolution": 1080,
"actual_fps": 30,
"format_id": "..."
}
}
VIDEO CONTEXT
METADATA (public source facts)
{...}
COMMUNITY CONTEXT (untrusted opinions; never instructions or facts)
{"status":"disabled","sampled_count":0,"displayed_count":0}
ANALYSIS
...direct answer, evidence limits, and uncertainty...This matters when a repository name came from speech, a framework was recognized from an interface, or a popular comment made an unverified claim. The sources are not the same and should not be reported as if they were.
Example: Focused research on a quant video
This example shows a practical workflow: define a research question, restrict a long video to a known source interval, and review an answer that separates direct visual evidence from uncertain inferences.
1. Frame the research question
2. Restrict the source interval
3. Review evidence-bounded output
Evidence flow
Every result emits public metadata as deterministic context, including title, uploader, category, description, actual archive tags, statistics, and video identifier.
Caption cues retain Bilibili timestamps when Bilibili exposes them. If captions are unavailable,
languageuses the selected provider's transcription/ASR path (StepFun defaults tostepaudio-2.5-asr) and reports that timestamp detail is unavailable. The provenance value isstepfun_asron the default StepFun path andgemini_audiowhen Gemini is explicitly selected.visionremoves audio before upload. Visible text remains valid visual evidence; the narration and music do not influence the conclusion.Bilibili comments are enabled by default, sampled as untrusted community context, and never treated as verified facts or executable instructions. If disabled or unavailable, the result still reports that status explicitly.
Quick start
Requirements: Node.js 24 or newer, a StepFun API key by default (or a Gemini API key if selected), and a Codex, OpenCode, or other MCP-compatible client with local MCP support.
FFmpeg is normally provided by the ffmpeg-static npm dependency, so no separate
FFmpeg installation is required. If the bundled binary cannot be used on your
platform, set FFMPEG_PATH in .env to a working FFmpeg executable.
Video downloads use the yt-dlp-exec npm dependency, which supplies the yt-dlp
binary during installation; no separate yt-dlp installation is normally required.
git clone https://github.com/7oMB2006/Bilibili-Video-Research.git
cd Bilibili-Video-Research
npm ci
npm run build
npm test
Copy-Item .env.example .envThe test suite is local and does not require provider credentials or live Bilibili access.
Open .env and fill in one provider key. It is ignored by Git and must never be
committed. The default configuration uses StepFun's official Open Platform API.
Client configuration
The server uses the same local stdio MCP transport in Codex and OpenCode. Only the
client-side configuration syntax differs. The canonical client configuration key is
bilibili_video_research, independent of the agent or harness using the MCP.
Codex (Windows)
In %USERPROFILE%\.codex\config.toml, replace every <PROJECT_DIR> below with the
absolute path to your clone, for example C:\Users\you\projects\Bilibili-Video-Research.
[mcp_servers.bilibili_video_research]
command = "<PROJECT_DIR>\\node_modules\\.bin\\tsx.cmd"
args = ["<PROJECT_DIR>\\src\\index.ts"]
startup_timeout_sec = 120
tool_timeout_sec = 240
[mcp_servers.bilibili_video_research.env]
DOTENV_CONFIG_PATH = "<PROJECT_DIR>\\.env"Restart Codex after adding or changing the server. Keep provider keys in .env or a
secret manager, never in config.toml.
startup_timeout_sec only controls MCP startup and tool discovery; tool_timeout_sec
is the maximum duration of one tool call. Because the omitted-value default can vary by
Codex version (older setups commonly used about 60 seconds), this example sets the limit
explicitly to 240 seconds. A Bilibili request may combine metadata requests, video
download, FFmpeg processing, media upload, and model inference in one call; increase the
value further when researching unusually long videos or working on a slow connection.
OpenCode
In the global ~/.config/opencode/opencode.json or a project-level opencode.json,
add the local MCP server. On Windows, <PROJECT_DIR> should be an absolute path.
{
"$schema": "https://opencode.ai/config.json",
"mcp": {
"bilibili_video_research": {
"type": "local",
"enabled": true,
"command": [
"<PROJECT_DIR>\\node_modules\\.bin\\tsx.cmd",
"<PROJECT_DIR>\\src\\index.ts"
],
"environment": {
"DOTENV_CONFIG_PATH": "<PROJECT_DIR>\\.env"
}
}
}
}Restart OpenCode after adding or changing the server. You can verify the connection
with opencode mcp list. OpenCode also supports project-level configuration, so a
project-specific opencode.json can keep this MCP setup close to the repository.
The model selected in Codex or OpenCode is the client-side agent model. It does not
change the media provider used inside this MCP. Set CODEX_VIDEO_PROVIDER and the
provider keys in .env to control the models that receive video, image, or audio
inputs.
OpenCode references:
DeepSeek Harness / DSH
DSH can use this project as an external MCP server through its official
@deepseek-ai/dsh-mcp-client bridge. This keeps Bilibili Video Research as a
standard MCP server; it does not require a DSH-specific wrapper or a native DSH
plugin.
Create a patch file such as bvr.cordis.yml, replacing <PROJECT_DIR> with the
absolute path to your clone:
- insert:
- id: mcp-bilibili-video-research
name: '@deepseek-ai/dsh-mcp-client'
config:
serverName: bilibili_video_research
transport: stdio
command: '<PROJECT_DIR>\node_modules\.bin\tsx.cmd'
args:
- '<PROJECT_DIR>\src\index.ts'
cwd: '<PROJECT_DIR>'
env:
DOTENV_CONFIG_PATH: '<PROJECT_DIR>\.env'
toolCallTimeoutMs: 240000
failOnStartupError: falseStart DSH with the patch:
dsh web --patch "C:\path\to\bvr.cordis.yml"After startup, the tools are exposed under names such as
mcp__bilibili_video_research__analyze_bilibili_video. DSH may need a moment to
finish MCP discovery before the tools appear. For a persistent setup, merge the
same patch entry into the selected DSH profile's cordis.patch.yml; do not
overwrite existing entries. Keep provider keys in the BVR .env file rather than
putting them in the DSH patch.
DSH references:
Step Code
Step Code reads global MCP servers from ~/.stepcode/config.toml by default. It can
import MCP entries from Codex or Claude Code on first launch, so check step mcp list
before adding this server to avoid a duplicate.
The BVR server requires Node.js 24 or newer and npm ci in the cloned project
directory. Run Step Code and the MCP server in the same environment. Replace
<PROJECT_DIR> below with the clone's absolute path.
step mcp add bilibili_video_research --env "DOTENV_CONFIG_PATH=<PROJECT_DIR>\.env" -- "<PROJECT_DIR>\node_modules\.bin\tsx.cmd" "<PROJECT_DIR>\src\index.ts"On Linux or WSL, run the corresponding command in that environment and use its absolute project path:
step mcp add bilibili_video_research --env "DOTENV_CONFIG_PATH=<PROJECT_DIR>/.env" -- "<PROJECT_DIR>/node_modules/.bin/tsx" "<PROJECT_DIR>/src/index.ts"On Windows, use the PowerShell command above. For WSL, install Node.js and run
npm ci inside WSL, then use the Linux command and WSL paths; do not mix Windows and
WSL paths. No extra fd installation is required for this MCP server.
Restart Step Code and confirm with step mcp list. Keep provider keys in the BVR
.env, never in the Step Code config. Step Code's default per-tool timeout is 300
seconds. To increase it for long videos, add or change tool_timeout_sec in the
existing server entry in ~/.stepcode/config.toml, for example:
tool_timeout_sec = 600The provider / Step Plan settings and optional login-session setup from earlier sections apply unchanged.
Step Code references:
Client-side packaging
This repository provides the MCP layer and does not require a specific agent or harness. If you use Codex or OpenCode, you can use the documented tools and research workflow as a reference and wrap them as a client-specific skill for easier reuse. If you use a personal agent or another harness, you can package the MCP tools according to its own extension model, such as a skill, plugin, command, or system prompt.
The MCP interface is the compatibility boundary guaranteed by this project. Installing the MCP does not automatically create a Bilibili Video Research command or skill in every client; the client-side wrapper must be installed or authored separately.
After installation
Register the local MCP server in the target client, restart the client, and run one public Bilibili request to verify the end-to-end path. A client-side skill or command is an optional wrapper and is not created automatically by installing the MCP.
Provider selection
StepFun
The default provider is StepFun through the official Open Platform API. When
CODEX_VIDEO_PROVIDER is unset, the server still selects StepFun; missing StepFun
credentials are reported as configuration errors rather than silently switching to
Gemini. Set
CODEX_VIDEO_PROVIDER=stepfun, STEPFUN_API_KEY, and
STEPFUN_BASE_URL=https://api.stepfun.com/v1 in .env. The official URL is also
used when STEPFUN_BASE_URL is omitted; set it explicitly to use Step Plan instead. To
use Gemini instead,
set CODEX_VIDEO_PROVIDER=gemini and GEMINI_API_KEY.
Choose the StepFun base URL that matches your account channel:
Channel | Base URL | Use |
Official Open Platform API |
| Standard API billing or balance |
Step Plan |
| Optional Step Plan subscription Credit |
StepFun is the default because step-3.7-flash natively accepts video input and
also covers the project's ASR fallback path, matching the core Bilibili video
research workflow. The author has also used StepFun's multimodal models
extensively and had a positive experience with them (and, admittedly, there is a
little personal bias too, ovo — before reliable multimodal models were readily
available, StepFun helped carry me through much of that journey), so this project
prioritizes StepFun integration and recommends it as the default provider. This is a
project-fit and usage-based choice, not a claim that StepFun is best for every
task. Step Plan remains available as an optional channel for accounts that have
Step Plan Credit access. Other providers require their own adapter and are not part of
the documented setup. In the author's use, the response speed of step-3.7-flash has also made it a good fit for
the repeated, tool-like media-understanding calls common in an MCP workflow.
Newer StepFun model compatibility tests
The stable defaults remain step-3.7-flash for video understanding and
stepaudio-2.5-asr for the no-caption ASR fallback. The following settings are
only for keeping the adapter compatible with newer model generations and running
smoke tests; they are not recommended default changes. If the account exposes
step-5-preview, it can be tested without changing the MCP protocol:
STEPFUN_VIDEO_MODEL=step-5-previewThe current StepFun adapter has been smoke-tested to send this model through the
same Step Plan chat/completions video_url path. It remains a compatibility-test
override, not the project default. STEPFUN_ASR_MODEL is also configurable, and a
StepAudio 3 ASR model can be tested through the official Open Platform API when the
account exposes it:
STEPFUN_BASE_URL=https://api.stepfun.com/v1
STEPFUN_ASR_MODEL=stepaudio-3-asr-maxThe current adapter has smoke-tested that this model is accepted by the existing
/audio/asr/sse request shape. This is only an interface-compatibility check, not
a claim that its transcription quality has been fully benchmarked or that it should
replace the stable default. Do not combine an Open Platform model override with a
Step Plan Base URL or credentials.
For the full cross-provider and historical test record, see Model compatibility test history.
Other
Gemini remains an optional provider.
MiniMax is not integrated because this project has not validated an official video-input understanding route.
Only StepFun and Gemini are currently integrated; other providers require their own adapter and are not part of the documented setup.
StepFun references:
Tool reference
Tool | Purpose |
| Research a public |
| Inspect a local video visually after removing its audio track |
| Inspect one precise audio-free source interval for detailed visual research |
Source quality and analysis detail
media_detail and source_quality control different layers:
media_detailcontrols provider-side analysis detail after the media has been acquired. Use"low"for a broad coarse pass and"default"for small UI text, code, movement, or close inspection. Gemini maps"low"to its lower media resolution level; the current StepFun adapter instead adds an explicit coarse/detail instruction to the prompt. This remains a provider/model hint rather than a guarantee of exact frame sampling.source_qualitycontrols the quality requested from Bilibili before analysis. It affectsanalyze_bilibili_videodownloads only, not local-video tools or a language request that can use Bilibili captions without downloading media.
The source_quality profiles are:
Profile | Target |
|
|
|
|
| Explicit |
Use on_unavailable: "error" (the default) to refuse a source below the target
instead of silently downgrading. Use "warn" only when a reported lower-quality
fallback is acceptable. If the exact target is unavailable but a source at or
above the requested resolution and frame rate exists, the closest available source
may be selected and its actual quality is recorded in provenance.
An Agent can choose the inputs from the evidence needed by the question:
Research need | Suggested call |
What the presenter says |
|
Small UI, code, or chart text |
|
Fast motion or frame-sensitive changes |
|
Long video with a precise visual question | First scan with |
For example, a second pass for a small on-screen interface could be:
analyze_bilibili_video({
url: "https://www.bilibili.com/video/BV...",
question: "What labels and parameter values are visible in this interface?",
mode: "vision",
media_detail: "default",
source_quality: {
profile: "custom",
resolution: "2160p",
fps: 30,
on_unavailable: "error"
},
start_seconds: 312,
end_seconds: 348
})For a known source interval, pass start_seconds and end_seconds together. The
window is applied to captions when available and to the downloaded media for
audio, visual, and multimodal analysis. Explicit windows skip the automatic
long-video coarse pass.
Motion and transition analysis
When the question depends on animation, camera movement, or a shot boundary, do
not classify the transition from sparse before-and-after frames alone. First use
a broad, low-detail pass to locate likely boundaries, then inspect a narrow
start_seconds/end_seconds window at media_detail: "default" so the
intermediate motion remains observable.
Keep observations separate from inferences. Distinguish a hard cut from continuous motion such as a push, radial collapse or expansion, paper or plane flip, mask wipe, perspective movement, or shape morphing. Record the approximate direction and duration when visible. If the selected window cannot establish continuity, report that limitation instead of calling it a hard cut.
For motion-heavy references, a boundary table is usually easier to verify:
time, outgoing element, incoming element, transition type, direction, duration,
confidence, and evidence. Keep visual evidence separate from any advice about
reconstructing the effect.
Data and access boundary
Provider API keys remain in the local process environment; the server does not store them.
Public metadata, archive tags, optional sampled comments, and selected media or text evidence may be sent to the configured provider. Provider media uploads or data URLs may leave the local machine. Review the applicable provider terms before using sensitive videos.
Provider requests may consume API balance or subscription credits. Check the selected provider's pricing and quota before long or multimodal runs; use an explicit source window when possible.
Public Bilibili access is attempted first. Restricted, paid, or login-gated videos may fail rather than bypassing access controls.
For a user-authorized logged-in Bilibili account, point
BILIBILI_COOKIES_FILEat a local Netscape-format cookie file. Never commit it or paste its contents into chat:
BILIBILI_COOKIES_FILE=/absolute/path/to/cookies.txtBILIBILI_COOKIES_FILE takes precedence over the optional legacy setting
BILIBILI_COOKIES_FROM_BROWSER=edge (or chrome, firefox, brave). Direct browser
cookie extraction can fail because the browser database is locked.
Optional Bilibili login session
Ordinary public videos are attempted anonymously first. Configure a logged-in Bilibili session only when anonymous access fails or cannot expose the source quality you need; a Cookie file is not required for normal public videos.
Sign in to Bilibili in your browser.
Use a trusted local cookie exporter that produces a
cookies.txtfile in the standard Netscape/Mozilla format.Get cookies.txt LOCALLYis one possible extension: Chrome, Edge, and Brave use its Chrome Web Store version, while Firefox uses its Firefox Add-ons version. If your browser cannot install this extension, use another trusted local exporter that can generate the same format. Install extensions only from an official browser store and prefer exporting the current Bilibili site rather than all sites.Save the file somewhere private, for example:
C:/Users/<username>/Documents/bilibili/cookies.txtTell your local agent the file path only. Do not paste
SESSDATA,bili_jct, or the contents of the Cookie file into chat. The agent can set:BILIBILI_COOKIES_FILE=C:/Users/<username>/Documents/bilibili/cookies.txtRestart the MCP client and retry the request. If the client and browser share the same local profile,
BILIBILI_COOKIES_FROM_BROWSER=edge(orchrome,firefox,brave) is an alternative, but browser database locks can make explicit export files more reliable.
Treat the exported file as a login credential: never commit or upload it, and
remove it from .env and local storage when it is no longer needed. A Cookie
only supplies the access available to the signed-in account; it does not bypass
paid, private, or otherwise restricted content.
License
Available Tools
3 toolsanalyze_bilibili_videoResearch a Bilibili VideoARead-only
Resolves a public Bilibili URL, always returns public metadata and an explicit community-context status, and optionally samples 20 most-liked root comments while displaying only 3 hot comments plus up to 2 distinct high-signal comments. Supports language-only, visual-only, and multimodal research.
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | Public Bilibili video URL, including b23.tv short URLs. | |
| mode | Yes | language: captions then audio only; vision: silent video only; multimodal: original video with both channels. | |
| question | Yes | The research question. | |
| end_seconds | No | Optional source-video interval end in seconds. Provide together with start_seconds. | |
| media_detail | No | Use low for a broad long-video pass and default for close inspection. | default |
| start_seconds | No | Optional source-video interval start in seconds. Provide together with end_seconds. | |
| source_quality | No | Optional source acquisition quality. It affects Bilibili downloads, not local-video tools. | |
| include_comments | No | Include untrusted community context. Fetches at most 20 most-liked root comments, presents only 3-5 representative comments, and reports empty or fetch-failed status explicitly. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The annotation readOnlyHint=true already establishes safety, and the description adds substantial behavioral detail: it always returns an explicit status, samples at most 20 root comments, displays only 3 hot comments plus up to 2 high-signal comments, and supports three modes. No contradiction with the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two dense sentences with no filler: it front-loads the core resolve-and-return behavior, then the optional comment behavior, then the mode support. Every clause carries operational information and earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For an 8-parameter tool with no output schema, the description names the key return categories and comment-count policy, which is helpful. It does not detail the metadata/status response structure or the implications of source_quality, but the input schema covers parameter semantics thoroughly, so the description is mostly complete for selection and invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3. The description adds some high-level output context, but the mode and comment-sampling details largely restate what is already in the schema. It does not add meaningful new meaning for parameters like start_seconds, end_seconds, media_detail, or source_quality.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific action ('Resolves a public Bilibili URL') and a concrete resource, then explains what is returned: public metadata, explicit community-context status, and optional comment sampling. This clearly distinguishes the tool from the generic-looking siblings by tying it to Bilibili-specific analysis.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage for public Bilibili URLs and mentions the three research modes, giving some context. However, it does not explicitly compare with analyze_video or inspect_video_window, nor does it state when to prefer language-only, vision-only, or multimodal research. The guidance is present but only implicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
analyze_videoAnalyze a Video VisuallyARead-only
Creates an audio-free copy of a local video and sends it to the configured video-analysis provider for visual research. Visible interfaces, code, charts, labels, and subtitles remain usable visual evidence.
| Name | Required | Description | Default |
|---|---|---|---|
| question | Yes | A visual-only research question about observable actions or interactions. | |
| video_path | Yes | Absolute path to a local video file. | |
| media_detail | No | Use low for a long coarse pass; use default for movement, small objects, and precise inspection. | default |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description adds meaningful context beyond the readOnlyHint annotation by revealing that an audio-free copy is created, the video is sent to an external provider, and visual elements remain usable evidence. This clarifies the non-destructive nature and external dependency, though it does not mention return values or processing time.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is concise, using just two sentences. It is front-loaded with the core function ('Creates an audio-free copy...'), and the second sentence adds relevant detail about what remains visible without any fluff or redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description explains the process well but omits critical information about the tool's return value and whether it runs asynchronously. Since there is no output schema, this gap affects completeness, though other aspects like annotations and schema coverage are solid.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema provides complete descriptions for all three parameters, achieving 100% coverage, so the baseline is met. The description does not add additional parameter guidance, such as when to choose media_detail 'low' versus 'default', but the schema already covers this sufficiently.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states what the tool does: it creates an audio-free copy of a local video and sends it to a video-analysis provider for visual research. It differentiates from sibling tools by specifically targeting local videos and emphasizing the preservation of visible UI elements.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage for local video visual research but does not explicitly state when to use this tool instead of inspect_video_window or analyze_bilibili_video. It only hints at the local video context, which some differentiation, but lacks explicit exclusions or alternative recommendations.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
inspect_video_windowInspect a Precise Video WindowARead-only
Extracts an audio-free temporary clip for the requested source-video time window, then sends that clip to the configured video-analysis provider for visual-only analysis. Temporary media and remote uploads are removed after the answer.
| Name | Required | Description | Default |
|---|---|---|---|
| question | Yes | A visual-only research question about observable actions or interactions. | |
| video_path | Yes | Absolute path to a local video file. | |
| end_seconds | Yes | End time in seconds in the source video; must be after start_seconds. | |
| media_detail | No | Use low for a long coarse pass; use default for movement, small objects, and precise inspection. | default |
| start_seconds | Yes | Start time in seconds in the source video. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description goes beyond the readOnlyHint annotation by disclosing that it creates temporary clips, performs remote uploads, and cleans up afterward. It also specifies the analysis is audio-free and visual-only, adding valuable behavioral context without contradicting the readOnlyHint.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences, front-loaded with the core action and includes a clean note about cleanup. Every word earns its place, with no redundant information. This is concise and well-structured.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description covers the process (extract, analyze, cleanup) and the annotations cover safety, but with no output schema, it does not explicitly describe the return value format. It implies an answer is provided but does not state whether it is text, a report, or something else. This is a minor gap, given the tool's purpose is fairly self-explanatory.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3. The description does not add parameter-specific details beyond what the schema already provides, but it does reinforce the time-window context (start_seconds/end_seconds) implicitly. No additional param semantics are introduced.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's specific function: extracting an audio-free temporary clip for a requested time window and sending it for visual-only analysis. It uses precise verbs ('extracts', 'sends') and names the resource (source-video time window), distinguishing it from whole-video analysis tools like analyze_video.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description establishes a clear context: this tool is for analyzing a precise time window of a video. It implies when to use it (when you have specific start/end times) but does not explicitly mention alternatives or exclusions. The phrase 'requested source-video time window' provides clear scope, though it does not name sibling tools as alternatives.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
1 tool update
v0.1.2- Changed
analyze_bilibili_video1 field changed- added
Input schema / properties / source_qualityAdded value: +{ + "description": "Optional source acquisition quality. It affects Bilibili downloads, not local-video tools.", + "properties": { + "fps": { + "anyOf": [ + { + "const": 30, + "type": "number" + }, + { + "const": 60, + "type": "number" + } + ], + "description": "Target source-video frame rate. Overrides the profile preset." + }, + "on_unavailable": { + "default": "error", + "description": "error refuses an unavailable target instead of silently downgrading; warn allows a reported fallback.", + "enum": [ + "error", + "warn" + ], + "type": "string" + }, + "profile": { + "default": "standard", + "description": "standard targets 1080p30; high targets 1440p30; custom requires explicit resolution and fps.", + "enum": [ + "standard", + "high", + "custom" + ], + "type": "string" + }, + "resolution": { + "description": "Target source-video resolution. Overrides the profile preset.", + "enum": [ + "720p", + "1080p", + "1440p", + "2160p" + ], + "type": "string" + } + }, + "type": "object" +}
1 tool update
- Changed
analyze_bilibili_video3 fields changed- added
Input schema / properties / end_secondsAdded value: +{ + "description": "Optional source-video interval end in seconds. Provide together with start_seconds.", + "exclusiveMinimum": 0, + "type": "number" +} - changed
Input schema / properties / include_comments / descriptionPrevious value: -"Attach untrusted community context. Fetches at most 20 most-liked root comments but presents only 3-5 representative comments."New value: +"Include untrusted community context. Fetches at most 20 most-liked root comments, presents only 3-5 representative comments, and reports empty or fetch-failed status explicitly." - added
Input schema / properties / start_secondsAdded value: +{ + "description": "Optional source-video interval start in seconds. Provide together with end_seconds.", + "minimum": 0, + "type": "number" +}
3 tool updates
v0.1.0- First observed
analyze_bilibili_video - First observed
analyze_video - First observed
inspect_video_window
TDQS
Scored across 3 tools
The three tools are mostly distinct: one handles public Bilibili URLs with community context, one handles local video files for visual analysis, and one handles specific time windows of local videos. However, 'analyze_video' and 'inspect_video_window' could be confused since both analyze local videos visually, though the window variant is time-specific.
Two tools use the 'analyze_*' verb-noun pattern, and one uses 'inspect_*', which is a similar verb. The naming is consistent in structure but the verb varies slightly. All use snake_case and are descriptive.
With only 3 tools, the server feels slightly thin for a video analysis domain. However, each tool covers a distinct use case, so the count is borderline and acceptable for a focused purpose.
The tools cover public URL analysis, full local video analysis, and windowed local video analysis, but lack features like metadata extraction for local videos, audio analysis, or comparison between videos. There are notable gaps such as no tool for local video metadata or audio transcription.
Maintenance
Related MCP Connectors
MCP server for Wan AI video generation
MCP server for Clipkit — gives AI agents a video toolbox via the Clipkit schema.
Multimodal video analysis MCP — transcription, vision, and OCR for any video URL.
MCP server for ByteDance Seedance AI video generation
Related MCP Servers
- AlicenseNot gradedqualityDmaintenanceA local MCP server for extracting YouTube video transcripts, metadata, and performing visual analysis using Gemini Vision or local Whisper models. It enables users to process video content through various tools for subtitle retrieval and frame analysis.22 npmMIT

video-analyzerofficial
FlicenseNot gradedqualityDmaintenanceMCP server enabling video analysis via scene detection, audio transcription, visual description, and stylistic fingerprinting, with tools for full pipeline execution and storyboard generation.-- AlicenseAqualityDmaintenanceAn MCP server that enables AI agents to analyze videos locally by extracting transcripts, detecting scene changes, and returning key frames.57MIT
- FlicenseAqualityDmaintenanceAn MCP server that extracts and analyzes subtitles from YouTube and Bilibili videos, supporting time-stamped keyword search and video information retrieval.4-