JIZURA MCP
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@JIZURA MCP夜明けの色を/覚えてる をダークHUD、BPM120でプレビューして"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
JIZURA MCP & Video Studio
字面一 JIZURA ONE STOP EDITION(hirazisora氏によるfork版)をPlaywright経由で自動制御するMCP(Model Context Protocol)サーバー、およびYouTubeやローカル動画を取り込んでキネティックタイポグラフィ(リリック演出)をリアルタイムに重ねて編集・書き出せる専用WebUIアプリケーションです。
📖 HTML 利用ガイド(ビジュアル解説)
より詳細な図解・ステップ解説付きの HTML 形式ガイドを用意しています。ブラウザでそのまま閲覧できます:
WebUI起動中:
http://localhost:3000/guide.htmlローカルファイル:
docs/index.htmlまたはpublic/guide.htmlをブラウザで開く
Related MCP server: lyric-studio
🌟 主な機能と特徴
JIZURA Video Studio (WebUI)
YouTube動画自動ダウンロード: URLを入力するだけで
yt-dlpにより動画を取得・配置ローカル動画/音声対応: MP4, MOV, WebM, MP3等をドラッグ&ドロップで即座に読み込み
プレイヤートリマー & タイムスタンプ同期: 動画を再生しながら
[mm:ss.ss]形式のタイムスタンプを歌詞にワンクリック挿入🎬 テロップ配置 & スケール調整(被写体・顔避け機能):
「👇 下部テロップ(字幕風)」「🎯 中央(全画面)」「👆 上部」の配置切替
50%〜100%のスケール調整と、縦オフセット微調整スライダー
実写動画やキャラクター動画で「顔が文字や装飾パネルで隠れてしまう」問題を完全に防止
完全透過アルファオーバーレイ合成: JIZURA公式の「透過PNG(ZIP・背景なし)」とFFmpegを連携させ、背景を完全に抜いた高品質な文字演出を動画の上に合成
AIエージェント連携(MCPサーバー)
Claude Desktopや各種MCPクライアントから自然言語でJIZURAのUIを全自動操作
歌詞の流し込み、演出スタイル選択、BPM設定、チェックボックス切替、プレビュー取得、動画エクスポートに対応
🚀 クイックスタート
必要環境
Node.js (v18以降)
FFmpeg (macOS:
brew install ffmpeg)yt-dlp (macOS:
brew install yt-dlp※YouTube取得時に使用)
インストール
git clone https://github.com/RimgO/jizura-mcp.git
cd jizura-mcp
npm install
npm run buildWebUIの起動
npm run ui
# ブラウザで開く: http://localhost:3000動作確認テスト
npm run verify🎬 WebUI の基本操作フロー
動画の読み込み
「YouTubeから取得」でURLを入力、または「ローカル動画をアップロード」にMP4ファイルをドラッグ&ドロップします。
タイムライン同期
動画を再生し、歌い出しのタイミングで「⏱️ タイムスタンプ挿入」を押すと、再生位置が
[00:12.50]のように歌詞へ自動挿入されます。
スタイルとテロップ配置の選択
演出スタイル(ノワール・クロマ、ダークHUD、墨と朱、モノ・RGBなど)を選択します。
被写体の顔を遮らないために、「👇 下部テロップ (推奨)」および「標準 (65%)」がデフォルトで有効になっています。
プレビュー & 書き出し
「プレビュー生成」で静止画を確認できます。
「動画を書き出す」をクリックすると、JIZURAで透過PNG連番が書き出され、FFmpegで元動画とアルファ合成された完成MP4が生成・ダウンロードされます。
🤖 MCPクライアントへの登録(Claude Desktop等)
claude_desktop_config.json(macOS: ~/Library/Application Support/Claude/claude_desktop_config.json)に以下を追加します:
{
"mcpServers": {
"jizura": {
"command": "node",
"args": ["/絶対パス/to/jizura-mcp/dist/index.js"]
}
}
}開発中はビルドなしで実行可能:
{
"mcpServers": {
"jizura": {
"command": "npx",
"args": ["tsx", "/絶対パス/to/jizura-mcp/src/index.ts"]
}
}
}チャットでの指示例
「この曲に合う歌詞を書いて、JIZURAでダークHUD風のタイポグラフィにしてプレビューを見せて」
「サビの部分だけ大きく強調して、BPM 120で動画を書き出して」
🛠️ 提供MCPツール一覧
ツール | 説明 | 主な引数 |
| ブラウザ(Playwright Chromium)を起動してJIZURAを開く |
|
| 画面上の操作可能なボタン/チェックボックス/入力欄一覧を取得 | - |
| JIZURA記法の歌詞テキストを入力欄に流し込む |
|
| 表示ラベル名でボタンをクリック(「おまかせで作る」「シャッフル」等) |
|
| チェックボックスをON/OFF(「アイテム枠表示」等) |
|
| 入力欄に値を設定(曲名、アーティスト、BPM等) |
|
| 現在のプレビュー画面をBase64 PNGで取得 | - |
| MP4または透過PNG(ZIP)を書き出して指定パスに保存 |
|
| ブラウザを終了してリソースを解放 | - |
📝 JIZURA 歌詞記法リファレンス
記法 | 効果 | 例 |
| 改行ごとに次の演出カットへ切り替え |
|
| 1行の中でカットを分割 |
|
| サビ用・文字と装飾プレートを巨大化して強調 |
|
| 控えめに小さく表示(動画の邪魔をしない) |
|
行末 | 画面揺れ(シェイク)やフラッシュ演出 |
|
| 複数行を画面に残したまま順に重ねて表示 |
|
| LRC形式タイムスタンプ(ビート同期) |
|
💡 被写体(顔など)を邪魔しないコツ 歌詞に
*強調歌詞*を入れると、JIZURAはサビ演出として画面中央に巨大なパズルピースや帯プレートを出します。スッキリした文字だけの演出にしたい場合は*を外し、通常の改行や/で書くか、WebUIの「下部テロップ」モードをご利用ください。
📄 ライセンス & 著作権
Author: RimgO
License: MIT License
Copyright (c) 2026 RimgOクレジット
字面一 JIZURA ONE STOP EDITION (hirazisora fork): https://hirazisora.github.io/JIZURA/
JIZURA 原案・オリジナル: hirano 氏
Available Tools
10 toolsjizura_clickボタン等をクリックB
表示テキストでボタンを探してクリックする。例: 'おまかせで作る'、'シャッフル'、'次の案'、'MP4 を書き出す'、'テーマ設定'、'適用'。
| Name | Required | Description | Default |
|---|---|---|---|
| name | Yes | ボタンの表示テキスト(部分一致) | |
| exact | No | trueで完全一致のみ |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description must disclose behavioral traits. It only states the click action, omitting side effects, post-click behavior, error handling, or whether it waits for the element.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One front-loaded sentence with illustrative examples; no filler. Efficient and easy to parse.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple click tool with no output schema, the description covers purpose and parameter examples. However, with no annotations, it lacks behavioral details an agent might need to invoke correctly, such as what happens after clicking.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema fully documents both parameters, but the description adds concrete examples of button text values for the 'name' parameter, which helps an agent pick appropriate values. It does not explain the 'exact' parameter, which remains schema-only.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb (click) and resource (a button found by display text), with concrete examples of button labels. It clearly distinguishes from sibling tools that set checkboxes or fill fields, though it does not explicitly name an alternative.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No indication of when to use this tool versus siblings like jizura_set_checkbox or jizura_fill, nor any prerequisites or exclusions. The examples illustrate parameter values, not usage contexts.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
jizura_closeブラウザを閉じるB
開いているJIZURAのブラウザセッションを終了する。
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden. It states the session is terminated but says nothing about side effects—whether unsaved data is lost, what happens to open pages, auth requirements, or reversibility. For a terminal/close-style action this is a meaningful gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with no filler, so nothing needs trimming. It is appropriately minimal for a zero-parameter action, though equally terse.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple 0-param, no-output-schema tool, the description covers the core purpose. However, the one thing an agent most needs before closing a session—side effects such as lost unsaved state—is absent, leaving it just adequate.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so the baseline is 4. The description correctly implies no parameterization is needed, and there is no schema surface left undocumented.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
State a specific verb (終了する = terminate) and resource (JIZURA browser session), so an agent knows exactly what the tool does and can distinguish it from its counterpart jizura_open. It stops short of explicitly contrasting itself with that sibling, but the action is unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is only implied: the agent can infer it should be called when done with the browser session, and that jizura_open is the inverse. There is no explicit when-to-use statement, no prerequisites, and no note on whether other session tools must be called first.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
jizura_export動画・画像を書き出すB
書き出しボタンを押してダウンロードを待ち、ローカルに保存する。buttonNameは 'MP4 を書き出す' / '連番PNG(ZIP)' / '透過PNG(ZIP・背景なし)' のいずれか。
| Name | Required | Description | Default |
|---|---|---|---|
| savePath | Yes | 保存先のローカルパス | |
| buttonName | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It usefully discloses the async flow (press, wait for download, save locally) and that it writes to a local path, but it omits critical behavioral details such as overwrite semantics for savePath, timeout/error handling, and what happens if the export button is unavailable.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, front-loaded with the action sequence and followed by the critical buttonName enumeration. There is no wasted phrasing, and the structure directly supports correct invocation.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a UI-automation export tool with no annotations, no output schema, and only 50% schema description coverage, the description covers the core flow and the required buttonName values. However, it lacks essential context about failure modes, return behavior, and savePath overwrite semantics, so it is only minimally complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 50%: savePath has a schema description, but buttonName has none and no enum. The description compensates for buttonName by enumerating the three valid values ('MP4 を書き出す', '連番PNG(ZIP)', '透過PNG(ZIP・背景なし)'), which is a meaningful addition. It does not add constraints or format details for savePath beyond what the schema already says.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific compound action: press the export button, wait for the download, and save it locally. It is clearly an export operation and distinguishes itself from primitive UI siblings like jizura_click by describing the full save workflow, though it does not explicitly contrast itself with those siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explains the mechanical steps but gives no guidance on when to use this tool versus alternatives, no prerequisites, and no exclusions. It does not state, for example, that it should be used instead of jizura_click when a download is expected.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
jizura_fillテキスト入力欄に値を入れるB
ラベルまたはplaceholderで入力欄を探して値を入れる(BPM、開始秒、ファイル名など)。
| Name | Required | Description | Default |
|---|---|---|---|
| name | Yes | ||
| value | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries full behavioral burden. It usefully discloses the locating mechanism (label or placeholder), but says nothing about failure behavior when no field matches, whether existing values are overwritten, or whether input events are dispatched.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence that states the mechanism first and appends examples. Efficient with no wasted wording, though it could not be shorter without losing the field-locating detail.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a two-parameter UI automation tool with no annotations and no output schema, the description covers the core action and target fields but omits error/edge behavior and interaction side effects, leaving meaningful gaps for an agent.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must compensate. It partially does: 'name' is implied to be the label/placeholder identifier and 'value' is illustrated with examples, but exact-match semantics and format expectations remain unstated.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('find the input field by label or placeholder and enter a value') and gives concrete examples (BPM, start seconds, file name). It is distinguishable from siblings like jizura_set_checkbox and jizura_click, though it never names them explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied by the field-locating mechanism (label/placeholder), but there is no explicit when-to-use guidance, no mention of prerequisites, and no routing against alternatives such as jizura_set_checkbox or jizura_set_lyrics.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
jizura_load_audio曲を読み込むC
「曲を読み込む」ボタンからローカルの音声ファイルを読み込ませる。
| Name | Required | Description | Default |
|---|---|---|---|
| filePath | Yes | ローカルの音声ファイルパス(mp3/wav等) | |
| buttonName | No | 既定は「曲を読み込む」 |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral disclosure burden. It mentions the button mechanism but omits what happens on invalid files, whether the current track is replaced, permission needs, or UI side effects like opening a file dialog.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single, efficient sentence with the action and mechanism front-loaded. It is appropriately sized for the tool's simplicity, though it could benefit from one more clause about prerequisites.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool is simple and the schema is complete, with no output schema needed. However, for a UI automation tool the description omits prerequisite state such as whether jizura_open must be called first, leaving a minor gap in context.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so both parameters are already documented in the schema. The description adds no additional meaning about filePath or buttonName beyond what is structurally provided, making the baseline 3 appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: loading a local audio file via the named button. This distinguishes it from generic siblings like jizura_click or jizura_fill, though it does not explicitly name an alternative.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides no when-to-use or when-not-to-use guidance. It only says what the tool does, leaving the agent to infer context such as whether the app must be open first or whether this replaces the current audio.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
jizura_openJIZURAを開くA
https://hirazisora.github.io/JIZURA/ をPlaywrightのChromiumで開き、操作可能なセッションを開始する。他のjizura_*ツールを使う前に一度呼ぶ。
| Name | Required | Description | Default |
|---|---|---|---|
| headless | No | trueでヘッドレス起動(既定はfalse=画面表示あり) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description must carry the behavioral burden. It discloses the browser engine (Playwright Chromium) and that the call establishes a stateful, manipulable session, which is meaningful. However it omits whether calling it twice is safe, whether the session must be cleaned up via jizura_close, and any failure/rate-limit behavior — notable gaps for a state-establishing tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences with zero filler: the action/URL comes first, the prerequisite-ordering rule second. Nothing could be removed without losing information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema and no annotations, but for a session-opener the description covers what the agent needs: what is opened, how, and when to call it. Minor gap: the session lifecycle (repeated calls, cleanup via jizura_close) is left to inference from sibling names.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
There is only one parameter (headless) and the schema documents it at 100% coverage, including the default (false = visible window). The description adds nothing about it, so the baseline of 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('開き' / open) plus the exact target resource (the JIZURA URL loaded via Playwright Chromium) and the effect ('操作可能なセッションを開始する'). This clearly separates it from the action siblings (click, fill, snapshot), which operate on an already-open session.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly states the prerequisite ordering: '他のjizura_*ツールを使う前に一度呼ぶ' (call once before using the other jizura_* tools). It does not state any when-not condition or mention jizura_close as the lifecycle counterpart, but the invocation context is unambiguous.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
jizura_screenshotプレビューのスクリーンショットA
現在の画面をPNG画像として取得する。アレンジ結果を目視確認したいときに使う。
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden, and it does disclose the output format (PNG). However it says nothing about scope of 'current screen', whether it requires the window to be open/focused, or latency/cost considerations for a capture operation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences with zero waste: the operation is front-loaded, followed immediately by the trigger condition. Every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter, no-output-schema tool the description covers what it does and when to use it. The one real gap is disambiguation from jizura_snapshot, which an agent in this family of tools would benefit from.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so there is nothing for the description to compensate for; the baseline of 4 applies. No parameter-related meaning is missing.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a concrete verb and resource: captures the current screen as a PNG image. That is specific and actionable, but it does not distinguish itself from the similarly named sibling jizura_snapshot, which an agent could reasonably confuse with a screen-capture operation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives an explicit usage condition: use it when you want to visually confirm the arrangement result. That is clear context for when to call it, though it names no exclusions or alternatives (e.g. snapshot vs screenshot) to disambiguate the sibling tools.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
jizura_set_checkboxチェックボックス/トグルを設定C
ラベル名を指定してチェックボックスのON/OFFを設定する。例: '和風の演出も使う', '追加分の演出も使う', '前景との重なりを避ける'。
| Name | Required | Description | Default |
|---|---|---|---|
| name | Yes | ||
| checked | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are supplied, so the description carries the full disclosure burden for a mutation tool. It does not state idempotency, what happens if the label name is not found, whether the change persists across snapshot/export, or error behavior for an invalid label.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Compact and front-loaded: the action and its key are stated first, followed by examples. No filler, though the examples occupy most of the length without adding usage or parameter-constraint information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a two-parameter mutation with no annotations and no output schema, the description is minimally adequate — it says what it does and gives target examples. It omits error/failure semantics and any indication of how to obtain valid label names, which matters since the labels are UI-dependent.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It does clarify that 'name' is a label name (not a free-form key) and gives three concrete example labels, and that 'checked' corresponds to ON/OFF. That covers both parameters at a basic level, but gives no format/length/constraint details or how labels are discovered.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (設定する/set) and resource (チェックボックスのON/OFF) plus the selection key (ラベル名). An agent can tell what it does, but nothing distinguishes it from sibling tools like jizura_click or jizura_fill, which could plausibly also target UI toggles.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides example label names, which hints at intended targets, but never says when to use this versus jizura_click/jizura_fill, nor any prerequisites (e.g., that the label must already exist in the current UI). No when/when-not guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
jizura_set_lyrics歌詞を設定A
歌詞入力欄にJIZURA記法のテキストを流し込む。
JIZURA歌詞記法(jizura_set_lyricsのtextに使う):
1行が1フレーズ。空行はやや間を空ける。
"/" でカットの切れ目を指定(例: 夜明けの色を/覚えてる)
"||"(全角縦線で囲んだ空白)… 文字なし・演出のみのカット
"{" と "}" で囲んだ複数行 … 最後のカットが消えるまで順に重ねて表示
"{-" と "-}" で囲んだ複数行 … 順に残して表示し、歌詞同士の重なりを避ける
"\n" … 1カット内の改行(文字として\nを書きたい場合は "\n")
"強調したい歌詞" … 強調(自動表示エリアを大きくし前景より手前に表示。サビなどに)
"
ささやきたい歌詞" … 抑制(表示・動きを小さく。Bメロや静かな部分に)行末の "!" … フラッシュと揺れ(盛り上がる瞬間に)
"歌詞|ルビや注釈" … 対応演出時の小さい注釈文字
"[01:23.45]歌詞" … LRC形式のタイムスタンプをそのまま使用可能
"#" で始まる行 … コメント(無視される)
記法文字自体を歌詞として出したい場合は直前に "\" を置く(例: \* \~ \{ \} \/ \| \! \# \[)
自然文の指示(例:「サビは強調して、Bメロはささやくように」)は、この記法に沿ってテキストを組み立ててから渡すこと。
| Name | Required | Description | Default |
|---|---|---|---|
| text | Yes | JIZURA記法済みの歌詞テキスト |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries full behavioral burden. It discloses how notation affects visual output (e.g., *emphasis* enlarges and brings forward, ~whisper~ shrinks), which is useful. However, it does not state whether existing lyrics are overwritten or appended, whether an open session is required, or any auth/permission needs. Core mutation semantics remain undisclosed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with the tool's action, then presents a well-organized bullet list of notation rules. Every line defines a distinct notation element or an escape rule, so the length is justified by the tool's domain-specific input format. There is no filler or redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool whose main complexity is the custom input notation, the description is highly complete: it thoroughly documents the parameter format. However, it omits any statement about mutation behavior (replace vs. append existing lyrics) or session prerequisites. Without annotations or an output schema, that missing behavioral context keeps it from being fully complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, but the schema description only says 'JIZURA記法済みの歌詞テキスト' (lyrics text in JIZURA notation). The description provides an extensive notation reference defining cuts, overlays, emphasis, timestamps, escapes, and more—far beyond what the schema conveys. This is exactly the additional meaning the parameter dimension rewards.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a clear verb and resource: '歌詞入力欄にJIZURA記法のテキストを流し込む' (pour JIZURA-notation text into the lyrics input field). This is specific enough for an agent to understand it sets lyrics, but it does not explicitly differentiate from sibling tools like jizura_fill, which also fills input fields. No alternative or exclusion is named.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides input-preparation guidance: convert natural-language instructions (e.g., 'make the chorus emphasized') into the JIZURA notation before passing. This implies when to use the tool (when you have lyrics or instructions to render), but it never states when to choose this tool over siblings or any preconditions/exclusions. Usage is only implied.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
jizura_snapshot画面上の操作可能要素を一覧取得A
現在表示されているボタン・チェックボックス・テキスト入力欄などの名前(ラベル)一覧を返す。jizura_click / jizura_set_checkbox / jizura_fill に渡すnameの確認・調整に使う。
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the burden, and it does disclose that it returns a list of names for currently displayed operable elements. However, it does not explicitly state read-only/non-destructive behavior, rate limits, or how the optional limit affects the result set.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tightly written sentences. The return value and scope come first, followed by the intended use case, with no redundant or wasted wording.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple snapshot tool, the core return and usage are covered. But the description leaves the limit parameter unexplained and does not describe the return structure beyond “list of names,” which is a gap given no output schema and no annotations.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The only parameter, limit, is undocumented in both the schema and the description. The description does not explain what limit controls, its default, or its interaction with the returned list, so it adds no semantic meaning beyond the schema’s type and constraints.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: returns a list of names/labels for currently displayed operable elements such as buttons, checkboxes, and text inputs. It also names the sibling tools (jizura_click, jizura_set_checkbox, jizura_fill) it feeds, making its role distinct from screenshot or interaction tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly states the usage context: use it to check or adjust the name passed to jizura_click, jizura_set_checkbox, and jizura_fill. It does not state when not to use it or contrast it with jizura_screenshot, but the intended workflow is clear.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
10 tool updates
v0.1.0- First observed
jizura_click - First observed
jizura_close - First observed
jizura_export - First observed
jizura_fill - First observed
jizura_load_audio - First observed
jizura_open - First observed
jizura_screenshot - First observed
jizura_set_checkbox - First observed
jizura_set_lyrics - First observed
jizura_snapshot
TDQS
Scored across 10 tools
Each tool has a clearly distinct role: session lifecycle (open/close), inspection (snapshot/screenshot), generic interaction by control type (click/set_checkbox/fill), specialized lyrics input, audio loading, and export. The only mild overlap is fill versus set_lyrics, but the descriptions clearly scope set_lyrics to the JIZURA notation lyrics field while fill handles generic fields.
All tools use the same jizura_ prefix followed by consistent snake_case action/resource names. The verbs are predictable (open, close, click, set_checkbox, fill, set_lyrics, load_audio, export), and snapshot/screenshot are still clear within the same convention.
Ten tools is well-scoped for a browser-driven web app controller. There are no redundant wrapper tools, and each operation corresponds to a distinct capability needed to drive the JIZURA workflow.
The surface covers session start/stop, UI inspection, button/checkbox/input interaction, specialized lyrics entry, local audio loading, screenshot verification, and export. Minor gaps could include reading current field values or checkbox states directly, but the core create-and-export workflow is complete.
Maintenance
Related MCP Connectors
- tonpitOAuthcom.tonpit
Video production studio for AI agents: AI media, motion graphics as code, timeline and export.
AI editor to build, animate & export layered short-form video projects via one tool catalog.
Generate images, GIFs, videos, and PDFs from HTML, URLs, or templates — from your AI agent.
AI-powered browser automation — navigate, click, fill forms, and extract data from any website.
Related MCP Servers
- FlicenseNot gradedqualityAmaintenanceEnables Claude to control a full-stack video editor by issuing commands to add clips, text, animations, and render MP4 videos, with changes reflected in real-time in the browser UI.-
- FlicenseNot gradedqualityCmaintenanceEnables AI assistants to create lyric videos by adding images, audio, styled text, and timed lyric lines, then rendering the final MP4.-
- AlicenseNot gradedqualityCmaintenanceEnables AI agents to programmatically scaffold Remotion projects, analyze audio for beat-synced scenes, synthesize voiceovers with word-level timecodes, preview frames, and render finished videos.10 npmMIT
- AlicenseNot gradedqualityAmaintenanceEnables AI agents to create polished product demo videos by controlling a real Chromium browser, recording actions, and rendering 1080p MP4s with narration, captions, and styled overlays.1MIT