guaardvark
This server is the MCP interface for Guaardvark, a self-hosted local AI studio, exposing tools for content generation, code work, media creation, knowledge retrieval, social outreach, and system operations.
Content generation: Create WordPress-compatible SEO CSV rows (basic, RAG-enhanced, or bulk) and generate new files or CSVs from descriptions.
Code tooling: Analyze, read, search, list, and verify code; run codegen to modify uploaded files; get repository maps, dependency graphs, and AST node source.
Web & files: Process uploaded documents, fetch/analyze URLs, and search the web.
Memory & knowledge: Save and search long-term memory; search the RAG knowledge base, list documents, view outlines, read sections, and summarize the corpus.
Media playback: Play music, control playback, adjust volume, and check playback status.
Media generation: Generate images, animations, videos, music videos, and Film Crew productions; edit, inpaint, outpaint, and remove backgrounds from images; poll generation status.
Social outreach: Check outreach status, list drafts, draft comments/shares, reject drafts, and request human-approved publishing.
System ops: Map the codebase, inspect GPU state, read logs, check swarm status, and view self-improvement status.
Provides a Discord bot and lets agents draft Discord outreach comments/shares or request publishing to a Discord webhook, all behind human approval.
Drafts Facebook outreach comments and share posts through the supervised, approval-gated outreach queue.
Drafts social outreach comments and share posts for Reddit, including subreddit-targeted shares, that wait for human approval before posting.
Generates WordPress-compatible CSV content with SEO-optimized copy for a client and topic, including enhanced variants with brand tone, industry, and local SEO context.
Drafts YouTube outreach comments and shares, with optional automatic scouting of video titles and descriptions, pending human approval.
https://github.com/user-attachments/assets/c6d9d18b-cfff-4ae2-8220-dc7f329fee5d
Guaardvark
The self-hosted AI studio. Coding agents and 20-agent swarms in isolated git worktrees, screen agents with their own real desktop, self-tuning RAG, continuous voice chat — and a full media pipeline: video, image, full-song music, neural voice. One install, one GPU, everything on your machine. Your machine. Your data. Your rules.
Works with your coding agent. Claude Code, Cursor, Codex, OpenClaw and Gemini CLI drive every flow above through the built-in MCP server and fifteen agent skills: "make a music video from this song", "film this script", "train a LoRA of this character", "swarm this refactor" — the agent queues the job on your GPU and polls it to the finished file.
Runs on Linux with one NVIDIA card (16 GB for video). Apple Silicon is supported with GPU features arriving through Metal (what works today is in INSTALL.md); Windows through WSL2 is being verified.
Install (one command, then open the Studio):
curl -fsSL https://guaardvark.com/install.sh | bashAdd it to Claude Code (two lines, no clone):
/plugin marketplace add guaardvark/guaardvark
/plugin install guaardvark@guaardvarkSee the VERSION file for the current release · guaardvark.com · Quick Start for manual install options.
For the exhaustive feature list, models, surfaces, and plugin details, see CAPABILITIES.md. This README focuses on the marquee experience, quick start, and what makes Guaardvark different.
What's in the box
See it | ||
Media studio | 11 local video models across five families (Wan 2.2, CogVideoX, LTX, HunyuanVideo, MiniMax H3 with native audio), image generation, full-song music, neural voice with consent-gated cloning, 4K/8K upscaling | |
Directors | A beat-synced music-video director, a 5-role Film Crew, an auto-editing video editor — and the walkthrough director that produced this README's own video series | |
Coding agent & code intelligence | Monaco editor, AST-aware analysis and dependency graphs, System Mapper: a live constellation of the whole codebase | Ep 14 |
Agent swarms | Up to 20 parallel coding agents in isolated git worktrees with dependency-aware merging; fully-local backend via Ollama | |
Screen agents | A real Ubuntu/XFCE desktop of their own, vision + closed-loop servo clicking, live VNC viewer on any page | Ep 4 |
Knowledge | Hybrid RAG on pgvector with cross-encoder reranking, layout-aware document parsing with page-level citations, retrieval that shows its chunks and scores, Autoresearch that tunes retrieval overnight | Ep 3 |
Voice & channels | Continuous voice chat, a three-tier chat brain, Discord bot, supervised outreach, MCP in both directions, a 25-module CLI | Ep 2 |
Self-running platform | Self-improvement behind guardian review and kill switches, rules engine, jobs & scheduling, schema-aware backups, GPU orchestrator, multi-machine Interconnector |
The aardvark (/ˈɑːrd.vɑːrk/; Orycteropus afer) is a medium-sized, burrowing, nocturnal mammal native to Africa. The aardvark is the only living member of the genus Orycteropus, the family Orycteropodidae and the order Tubulidentata. It is found over much of the southern two-thirds of the African continent, avoiding areas that are mainly rocky. A nocturnal feeder, the aardvark subsists on ants and termites (myrmecophagy) by using its sharp claws and powerful legs to dig the insects out of their hills, and its long snout to sniff out food. It digs a burrow in which to live and rear its young.
— Wikipedia, CC BY-SA
The Guaardvark (/ˈɡwɑːrd.vɑːrk/; Workstationus selfhosticus) is a burrowing, nocturnal AI system native to consumer hardware. The Guaardvark is the only living member of the repository github.com/guaardvark/guaardvark, the family LocalAI, and the order AutonomousAgents. It is found across most of the modern desktop, avoiding regions that are mainly cloud. A nocturnal feeder, the Guaardvark subsists on prompts and unstructured data (promptophagy) by using its sixty-odd tools and a swarm of parallel coding agents to dig bugs out of their codebases, and a retrieval index to sniff out knowledge in the dark. It is fiercely territorial about its single GPU, admitting one process to the card at a time and evicting any language model found loitering there. It digs isolated git worktrees in which to work, and rears its images, video, music, and cloned voices entirely on your own machine.
Related MCP server: omnicinema-mcp
▶ The Walkthrough Series — every feature, on camera
Short, unscripted-feeling screen recordings of the real system doing real work — narrated by a voice the system cloned itself (that's Episode 7). Twelve episodes are live: the first series covers the twelve subsystems, and a second series picks up what shipped since.
|
|
|
|
|
|
|
|
|
|
|
|
▶ Watch the full playlist — Episode 1 (the full tour) and Episode 10 (the video editor) are on the way.
More demos
The beat-synced music video — watch on YouTube. One style prompt and a short narrative, then go: Guaardvark wrote every shot prompt, generated the storyboards, rendered the clips, and assembled the cuts to the beat it detected in the song. (Full disclosure: the glitch effect was the one manual touch, added in Shotcut; the song was made in Suno — Guaardvark's own music generation is being wired into this pipeline.)
Full visual gallery (dashboard, video generator, swarm planner, agents, plugins, media library, etc.) is available on guaardvark.com.
Marquee Capabilities
Generation & Editing (all local, no cloud APIs)
Text-to-Video / Image-to-Video (Wan 2.2 5B + 14B MoE, CogVideoX-5B, LTX-2.3, LTX-2.5) with an in-process batch queue, quality tiers, frame interpolation, prompt enhancement, and one-click jump to ComfyUI for custom workflows.
Audio Studio (Audio Foundry plugin): ACE-Step 3.5B music (vocals or instrumental, Suno-style chips + LLM polish), Stable Audio Open FX/ambience, Chatterbox + Kokoro neural TTS, Piper fallback, consent-gated voice cloning.
Image gen (Stable Diffusion + batch + face/anatomy controls) + powerful 4K/8K GPU upscaling (Real-ESRGAN family, HAT-L, NMKD, Foolhardy, two-pass, video frame-by-frame).
Built-in Video Editor (Shotcut-lite 3-lane timeline: video/text/audio, real
ffmpeg drawtextoverlays, drag-and-drop from media library, visual trims, undo, keyboard shortcuts).
Agents, Automation & Swarms
AgentBrain three-tier router (Reflex <100 ms pattern match, Instinct single-shot, Deliberation full ReACT).
Real-desktop screen agents (Xvfb + XFCE
:99, Gemma4 vision + closed-loop servo, 45+ deterministic recipes, live per-iteration reasoning stream in chat, draggable VNC viewer everywhere).Swarm: parallel agents in isolated git worktrees (Claude Code or fully local Cline/OpenClaw via Ollama), Flight Mode (offline), dependency-ordered merge, cost tracking, up to 20 concurrent.
Film Crew: 5 specialized agents that turn a logline into a finished video (script → casting with LoRAs → shots → keyframes → edit).
Self-improvement engine (test → agent fix → verify → broadcast) with guardian review and kill switches.
Supervised social outreach (Reddit fully working; others drafting+review ready) with persona, grading, cadence, audit, and global kill switch.
MCP server + client integration (Claude Desktop, Cursor, etc.).
Knowledge, Code & Workflow
Strong RAG (hybrid keyword + vector search on pgvector, cross-encoder reranking, AST code chunking, layout-aware PDF/DOCX parsing, entity extraction, RAG Autoresearch, per-project isolation).
Monaco code editor + Code Analyzer + per-repo indexing + dependency graphs + System Mapper (constellation view of the whole codebase).
Full desktop-grade file/project/client/website/notes/media management with cross-links and recursive indexing.
Task scheduler (Celery beat), Rules & Prompts (portable bundles), Interconnector for multi-machine clusters (master/client, approval gates, learning broadcast).
10+ managed plugins with health checks, port orphan cleanup, and a real GPU Memory Orchestrator.
Platform & Ops
Everything stays on your machine by default. Flight Mode is real and end-to-end tested.
Plugin system + resource orchestrator so big models don't fight for VRAM.
Backup/restore (granular or full, schema-migration aware), advanced settings surfaced in UI, live GPU/CPU monitoring.
CLI (
llx/ PyPIguaardvark), browser UI, and MCP.
See CAPABILITIES.md for the complete enumerated list (models, exact tool counts, plugin manifests, page surfaces, etc.).
Why local?
Cloud platforms | Guaardvark | |
Where your data lives | Their servers | Your machine. Period. |
Per-token / per-minute fees | Always on the meter | Free. Generate all night if you want. |
Content policy | Their rules | Your rules. |
Custom models / LoRAs | Whatever they expose | Any GGUF, any LoRA, any embedding model |
Works offline | No | Yes. Flight Mode tested end-to-end. |
Agents drive a real desktop | Sandboxed browsers | Real Ubuntu/XFCE on your hardware |
Swarms of parallel agents | Per-task billing scales nastily | 20 agents in parallel; only cost is power |
Multi-machine clusters | "Talk to sales" | Built-in. Master/client, approval gates |
Lock-in | Migrate at your own risk | It's your computer. Move it whenever. |
How Guaardvark compares
The local-AI ecosystem has excellent tools for every slice: chat UIs, RAG second-brains, node-graph media pipelines, coding agents, assistant gateways. Guaardvark's bet is different — one install on one GPU that is the whole studio, with a single GPU orchestrator arbitrating all of it.
Capability | Chat UIs | RAG apps | Node graphs | Coding agents | Assistant gateways | Guaardvark |
Local chat + RAG | core | core | — | — | via tools | core (Ep 3) |
Media production (video · image · music · voice) | — | — | image/video graphs | — | via connected tools | core, with director engines (Eps 5–9) |
Agents on a real desktop | — | — | — | — | browser/tool use | core (Ep 4) |
Parallel coding swarms | — | — | — | usually one agent | — | up to 20 in git worktrees |
Self-improvement behind human gates | — | — | — | — | — | core (Ep 11) |
One-GPU resource arbitration | — | — | — | — | — | core (Ep 12) |
Integration / plugin ecosystem breadth | varies | varies | enormous | growing | enormous | smaller — 10 first-party plugins, plus MCP both ways |
Hosted / mobile option | often | often | often | often | often | none, by design — it's your machine |
Columns describe the typical shape of each category, not any single project — several projects exceed their category in places. The Guaardvark column links to walkthrough episodes where you can watch the claim happen.
If all you need is one slice, use the excellent specialist: a chat UI like Open WebUI, a node graph like ComfyUI (Guaardvark hands off to it with one click), a RAG workspace like AnythingLLM. Guaardvark is for when you want the whole studio on one box.
Agent-driven media production, side by side
A newer category: the coding agent runs the studio. Facts checked 2026-09-11 from each project's repository; stars move, the shape does not.
OpenMontage | Nomi | Maestro | Comfy MCP | Promptus / LocalForge / SimpliGen | Guaardvark | |
What it is | 12 video pipelines driven by Claude Code, Cursor, Codex | Desktop video workbench with 25 MCP tools | Local video, image, music, voice with a director mode | Official MCP for ComfyUI | One-click local image + video apps, $30–$97 one-time | The whole studio, driven by your agent or the Studio UI |
Generation runs | mostly cloud APIs, local models optional | your ComfyUI or cloud providers | local (a Wan2GP fork) | your ComfyUI, or Comfy Cloud on subscription | local | local |
Agent driving it | yes (skills + CLI) | yes (MCP) | no (in-app planner) | yes (generation only) | no | yes (MCP + skills) and the built-in agent brain |
Music, voice, voice clone | via cloud TTS/Suno | — | music + voice | audio nodes | — | ACE-Step songs, three TTS engines, consent-gated clone |
Film crew, music-video director | pipelines, storyboard board | storyboard + timeline | director mode | — | — | 5-role Film Crew, beat-synced director, video editor |
Coding swarm, screen agents, outreach, RAG | — | — | — | — | — | core |
LoRA training, upscaling, add any HF model by URL | — | — | LoRA browser | — | model manager | core, in the Studio |
OS | mac, Linux, Windows | mac, Windows | NVIDIA via Pinokio | any | Windows, mac (LocalForge: Linux too) | Linux; Apple Silicon partial (Metal); WSL2 in verification |
License | AGPL-3.0 | AGPL-3.0 | WanGP non-commercial | open source | proprietary | MIT |
Repository stars, 2026-09-11 | 57k | 0.5k | 0.5k | (part of ComfyUI, 133k) | — | 0.2k |
Every cell is a claim you can check in the linked repositories; corrections welcome in an issue.
What Makes This Different
Security, Privacy & Local Guarantees
Everything runs locally by default. No telemetry or cloud phoning home unless you explicitly enable the Interconnector (master/client with approval workflows).
Flight Mode — fully offline operation with automatic network detection and local-model fallback. Swarm and agent tasks have been validated end-to-end without internet.
Self-improvement safety — three modes (Scheduled / Reactive / Directed). Every proposed code change can be reviewed by "Uncle Claude" (Anthropic API guardian) before application. Codebase lock toggle + Pending Fixes queue for human staging/approval. Fixes can be broadcast to connected family members.
MCP server uses a strong default-deny policy (
backend/mcp/config.py): desktop control, agent execution, system/shell, browser automation, and test execution tools are hidden by default. Only safer tools + read-onlyguaardvark://outputs/resources are exposed unless you explicitly allowlist.Outreach is supervised by default (drafts queue; nothing posts without explicit Approve). Kill switch, per-platform cadence limits, full JSONL audit trail, and persona enforcement.
Voice cloning requires an explicit consent prompt. Reference clips stay under your control.
WordPress connectivity ships with security disclaimers and is treated as opt-in/beta until a final hardening pass.
Your data, models, LoRAs, and generated media never leave the machine unless you choose to push them.
AgentBrain — Three-Tier Neural Routing
Every message is routed through a three-tier decision engine that picks the fastest path to the right answer. Reflexes fire in under a millisecond. Instinct handles single-shot requests in one LLM call. Deliberation spins up a full ReACT reasoning loop when the problem demands it.
Tier | Name | Latency | LLM Calls | When It Fires |
1 | Reflex | <100ms | 0 | Greetings, farewells, media controls — pattern-matched, no inference |
2 | Instinct | 1–3s | 1 | Single-shot questions, web searches, image generation, vision tasks |
3 | Deliberation | 5–30s | 3–10 | Multi-step research, analysis chains, complex agent tasks |
Automatic escalation — Tier 2 can signal complexity and hand off to Tier 3 mid-response.
Agent-screen gating — vision/desktop tools are only in scope when the virtual screen is active.
BrainState singleton + warm-up thread for zero-overhead routing and fast first-token times.
Autonomous Screen Agents
Guaardvark agents control a real Ubuntu desktop (Xvfb + XFCE at 1000×1000) — exactly what the model would see if you VNC'd into the box from another machine. Same Applications menu, same desktop icons, same taskbar. Agents see the screen through vision models, move the mouse, click buttons, type text, navigate browsers, and verify their own actions.
Real XFCE session — not a custom widget panel.
xfce4-sessionruns on the virtual display via a scrubbed environment, with isolatedXDG_DESKTOP_DIRandXDG_CONFIG_HOMEso the agent's desktop, file manager, and configs never collide with the user's. Vision models recognize the layout instantly because it's standard Ubuntu.Unified vision brain — Gemma4 sees the screen, decides the next action, and emits click coordinates (native
box_2d) in a single inference call. Per-model scale factors are tracked and updated by the self-improvement loop.Closed-loop servo targeting — three-attempt adaptive strategy: ballistic move → single correction with crosshair overlay → full corrections with zoom-cropped analysis around the cursor
Live per-iteration reasoning stream — every Think step (action, target, full reasoning, pivots when the loop gets stuck) streams into chat in real-time. No more 30-second blackouts followed by a single "completed" line. The trail persists in history so you can audit any run.
45+ deterministic recipes — browser navigation, tabs, scroll, search, find, zoom, copy/paste — all execute instantly from a JSON recipe library, bypassing the vision loop entirely. Recipes carry optional
preconditions(visibility checks) so they're skipped cleanly when their UI isn't on screen.Obstacle detection — handles popups, permission dialogs, and notification bars with automatic thinking model escalation
Self-QA sweep — agent navigates every page of its own UI and reports what's working and what's broken
Live agent monitor — real-time SEE/THINK/ACT transcript of every decision the agent makes
Integrated screen viewer — draggable, resizable VNC viewer on any page with popup window mode
Supported Vision Models
Model | Role | Coordinate System | Notes |
Gemma4 (e4b) | Sees + decides + clicks | box_2d normalized to 1000, | Unified brain — vision, reasoning, and coordinates in one call |
Moondream | Fallback eyes | 1024px internal width | For text-only chat models (llama3, ministral-3) that need external vision |
Swarm Orchestrator — Parallel Agent Execution
Launch multiple AI coding agents in parallel, each working in an isolated git worktree on its own branch. Results merge back with dependency-ordered conflict detection, optional test validation, and full cost tracking.
Two backends — Claude Code (cloud, cost-tracked at $0.015/$0.075 per 1K tokens) and Cline/OpenClaw (fully local via Ollama, zero cost)
Flight Mode — fully offline operation. Auto-detects network state, falls back to local models, serializes file conflicts automatically. No prompts, no internet required.
Git worktree isolation — each task gets its own branch and working directory. All worktrees share the
.gitdirectory (lightweight). Automatically excluded fromgit status.Dependency-aware merging — topological sort ensures foundational changes land first. Dry-run conflict detection before real merge. Test suite validation before integration.
Built-in templates — REST API scaffold, refactor-and-extract, test coverage expansion, Flight Mode demo
Up to 20 concurrent agents — configurable limit with automatic slot management
Live dashboard — real-time status, per-task logs, cost breakdown, elapsed time, disk usage
Self-Improving AI
The system runs its own test suite, identifies failures, dispatches an AI agent to read the code and fix the bugs, verifies the fix, and broadcasts the learning to other instances. No human in the loop.
Three modes — Scheduled (every 6 hours), Reactive (triggered by repeated 500 errors), Directed (manual tasks)
Guardian review — Uncle Claude (Anthropic API) reviews code changes for safety before applying, with risk levels and halt directives
Verification loop — re-runs tests after every fix to confirm it worked
Pending fixes queue — stage, review, approve, or reject proposed changes
Cross-machine learning — fixes propagate to all connected instances via the Interconnector
RAG That Actually Works
Chat grounded in your documents. Upload files, build a knowledge base, and ask questions. The AI reads and understands your content — not just keyword matching.
Hybrid retrieval — Postgres full-text keyword + vector semantic search, fused with a per-query weighting that leans keyword-ward for identifier-like queries and semantic-ward for prose
Cross-encoder reranking — a reranker reads the query and passage together and reorders the candidates, which a bi-encoder cannot do; it is admitted against free VRAM and falls back to CPU rather than competing with image or video generation
Smart chunking — code files get AST-informed chunking (a function stays one chunk instead of being split mid-body), prose gets semantic splitting
Layout-aware document parsing — PDFs, DOCX and PPTX are parsed for reading order, section headers and page positions, so a retrieved passage can cite the page it came from. (Scanned documents need OCR, which is not installed by default — they report that rather than indexing as empty.)
Grounded citations — every chunk carries its source file, section breadcrumb and page, and that context is written into the embedded text as well, so a passage lifted out of the middle of a document still says where it came from
Postgres-backed vector store — embeddings live in pgvector alongside the rest of your data, with an ANN index and a persisted full-text index; they are covered by the same database backups
Index profiles — the same documents can be projected more than one way (fewer, larger passages for a small local model; finer-grained ones for an external client), switched in Settings
Corpus-level summaries — a recursive summarisation pass answers "what are the themes across all of this", which passage search structurally cannot
Multiple embedding models — switch between lightweight (300M) and high-quality (4B+) via UI
Entity extraction — automatic entity and relationship indexing
Per-project isolation — each project has its own knowledge base and chat context
Retrieval that shows its work — live retrieval tests display the actual chunks and scores behind an answer (Episode 3), instead of just asserting one. Every query can also return a trace of which retrieval legs actually ran, so a degraded answer is distinguishable from a bad one
Autoresearch — retrieval that tunes itself. An autonomous optimization loop runs overnight experiments on your corpus: it proposes changes to chunking and retrieval parameters, evaluates them with an LLM-as-judge harness, keeps wins, and reverts regressions — bounded by a wall-clock budget, a run ledger, and a circuit breaker (Episode 11). Your retrieval gets better while you sleep, and the morning report says exactly what changed and why.
Even the file manager is better
A client's Linux desktop player refused to play a video — the distro was missing the right codec plugin. They dropped the same file into Guaardvark's file desktop and it just played. No codec pack involved: the backend range-streams files inline (download_document in backend/api/files_api.py) into the browser's own decoders, which ship with H.264/VP9 support regardless of what the desktop has installed. The same mechanism means PDFs start painting before they finish downloading and video seeking works from the first byte.
Code Intelligence & the System Mapper
Monaco code editor with multi-file tabs and an AI assistant pane.
AST-aware code intelligence — repository maps, dependency graphs, and structure-aware code search (
get_repository_map,read_ast_node,search_code— also exposed over MCP).System Mapper — a live, force-directed constellation of the entire codebase computed from real imports (1,300+ modules on camera in Episode 14), with lifecycle tagging (active / dormant / auto-loaded / test / script / config), ranked findings you can dispatch to the self-improvement agent, and findings whose remedy is mechanical staged as an exact proposal for your review.
Guarded self-coding — every AI code write funnels through a single verified exact-replacement gate, behind the codebase lock.
Model Context Protocol (MCP)
Guaardvark speaks MCP both ways — exposes its tools to any MCP client (Claude Code, Cursor, Grok, Claude Desktop, Zed, Gemini, etc.) and can call tools from connected external MCP servers.
One-command setup —
python -m backend.mcp installdetects the agent clients on your machine and writes theguaardvarkserver entry into their configs (existing files are backed up, other entries untouched).python -m backend.mcp doctordiagnoses a broken setup: server self-test, a real stdio handshake, and a scan of client configs for stale paths.Claude Code plugin — two lines, no clone:
/plugin marketplace add guaardvark/guaardvarkthen/plugin install guaardvark@guaardvark. It asks for the path of your Guaardvark checkout, wires the MCP server from there, and loads every skill below as/guaardvark:<skill>.Agent skills —
.agents/skills/ships one skill per flow (images, video, music video, Film Crew, voice, music, upscaling, Cast/LoRA training, Hugging Face model onboarding, swarm, knowledge, code, outreach, ops) in the Agent Skills format, so Claude Code, Cursor, Codex and OpenClaw know which tool or route to call for each job.python -m backend.mcp install --skillslinks them into~/.claude/skills; other agents read.agents/skills/from the checkout. Start withsetup.As a server —
python -m backend.mcp(stdio, the default) orpython -m backend.mcp http(streamable HTTP on127.0.0.1:8788/mcp; loopback-only by default since there is no auth yet). Strong default-deny policy (seebackend/mcp/config.py): categories such asdesktop,agent_control,system,browser,test_execution, andmcpmeta-tools are denied by default. Dozens of safer tools (chat, RAG, files, generation, memory, etc.) plus read-onlyguaardvark://outputs/resources are exposed —python -m backend.mcp list-toolsprints the live list. Generation tools queue by default over MCP and hand back a batch id (get_generation_statusreads it); every call runs on a worker thread under an enforced timeout (GUAARDVARK_MCP_TIMEOUT, 120 s; 30 min when a caller asks to wait for a render). Verified end-to-end by an initialize/tools-list handshake in the smoke tests.As a client —
mcp_connect/mcp_execute+ live tool inventory so the chat LLM can discover and use tools from other MCP servers by name.Audit logging, timeouts, and circuit breakers are built in.
Outreach System — Supervised AI for Social-Media Engagement
A supervised, auditable framework for drafting and posting authentic comments on Reddit, Discord, Twitter/X, and Facebook — using your own indexed knowledge as the source of truth for citations and context. The point isn't volume. It's keeping up with engagement on your own products and topics, with the agent handling the legwork.
How it works:
Discover — the agent scouts target threads either by URL (you paste one into the New Draft modal) or by walking platform-specific entry points (subscribed subreddits, Discord channels, Twitter feeds, Facebook groups).
Context — for each candidate post, the agent fetches the OP body and top comments. Reddit goes through the JSON API (fast, no scrape). Discord, Twitter, and Facebook go through the agent's logged-in Firefox session over CDP/BiDi, with a vision-model fallback when DOM selectors drift after a platform redesign.
Draft — your local LLM composes a reply grounded in the thread context plus citations from your indexed documents (clients, projects, products, examples — whatever you've fed the knowledge base).
Grade — every draft is scored against a relevance + quality rubric. Anything below threshold is dropped before it reaches the queue. Generic "great post!" replies don't survive grading.
Review — drafts land in a queue. In supervised mode (the default), nothing posts without your approval. Edit, save, approve, reject — your call on each one.
Post — approved drafts post via the logged-in browser session (Reddit/YouTube servo) or Discord API, cadence-gated. Natural language from chat (
/outreach …) orllx outreach "…"runs recon+draft; posting still needs approve while supervised. Twitter/Facebook drafting works; auto-post for those platforms is not wired.
Three layers of safety:
Kill switch at the system level. Flip it off and every outreach pipeline — drafting, queueing, posting — stops mid-flight. Nothing escapes.
Supervised mode is the default. Drafts queue, never auto-post. You approve each one explicitly.
Cadence gates — at most 1 post per 30 minutes per platform, configurable. Prevents bot-shaped behavior and respects platform anti-spam expectations.
Audit log — every action (scout, draft, grade, approve, reject, post, fail) is recorded in a JSONL audit trail with timestamps, draft IDs, and outcomes. Exportable for compliance or post-hoc review.
Persona system — a single configurable persona (voice, expertise areas, citation style, what to never say) shapes every draft for consistency. Your replies sound like you, not like an LLM.
Manual draft mode — paste a thread URL, the agent auto-scouts the context, the LLM seeds a draft, you edit and save. Full human control with the agent doing the legwork (scouting, context-fetching, citation suggestion).
On-demand passes — instead of waiting for the cron, fire a pass for a specific platform or subreddit on demand from the UI. Useful for active engagement around a launch or a thread you spotted.
Why it's not spam — outreach is anchored on your own knowledge base. Citations point at YOUR documentation, YOUR examples. The system grades drafts for genuine relevance and refuses to engage when it can't add value. The cadence gate keeps the volume human-paced. Supervised mode keeps the human in the loop. The result is closer to "an assistant that helps you keep up with engagement on your own products and topics" than "an outbound bot."
Film Crew — End-to-End Production Pipeline
Five specialized agents collaborate to turn a one-line idea into a finished video. Built on the Swarm Orchestrator, so every role runs in parallel where possible and merges back deterministically.
Role | What It Does |
Screenwriter | Generates the script + scene breakdown from a logline |
Casting | Assigns characters to LoRAs (via the LoRA Trainer plugin) or stock characters |
Cinematographer | Produces a shot list with camera moves, framing, and lens choices |
Storyboard | Generates keyframe images for every shot via the image pipeline |
Editor | Assembles the generated clips into a finished video via the Video Editor |
The LoRA Trainer plugin ships alongside — train character/environment/prop LoRAs from reference images on your local GPU (bf16, ~46 MB per LoRA) and route them automatically to the Casting agent.
Music Video — Beat-Synced, Automatic
Give it a song (.mp3 / .wav), a style prompt, and a short narrative — Guaardvark does the rest:
Audio analysis & beat detection — the track is analyzed for tempo/beats so cut timing follows the music instead of an arbitrary clock.
Director — an LLM writes a distinct prompt for every cut (no mechanical repetition across a long song), keyed to your style + narrative.
Storyboards → video — a keyframe still is generated per cut (SDXL/FLUX, optional character LoRAs for identity), then animated with the chosen image-to-video model (Wan 2.2 I2V, etc.).
Beat-timed assembly — clips are stretched/filled to land on the beat (
clip stretch, fill methods) and assembled into the final cut, with RIFE frame interpolation for smoothness.Honest about the edges — native filters/transitions/effects aren't in yet (the demo's glitch effect was added manually in Shotcut); that's on the near-term roadmap.
Linux & macOS: The final assembly step needs melt (MLT) from Shotcut. ffmpeg is pre-installed by the platform bootstrap. Full commands (brew/apt/flatpak/snap) are in plugins/video_editor/README.md.
Video Generation Pipeline
State-of-the-art video generation running entirely on your GPU. No cloud APIs, no per-minute billing, no content restrictions.
Model | Type | Max Duration | Native Resolution | VRAM |
Wan 2.2 TI2V-5B (default) | Text + Image-to-Video | ~5s (up to 121 frames @ 24fps) | 1280x704 | ~11GB |
Wan 2.2 (14B MoE) | Text-to-Video | 5s (81 frames @ 16fps) | 832x480 | 11GB |
Wan 2.2 14B I2V | Image-to-Video | 5s (81 frames @ 16fps) | 832x480 | 11GB |
CogVideoX-5B | Text-to-Video | 6s (49 frames @ 8fps) | 720x480 | 16GB |
CogVideoX-5B I2V | Image-to-Video | 6s (49 frames @ 8fps) | 720x480 | 16GB |
LTX-2.3 Distilled FP8 | Text + Image-to-Video | ~10s (161 frames @ 16fps) | 768x512 | ~14GB |
LTX-2.5 Distilled Int8 | Text + Image-to-Video | ~10s (161 frames @ 16fps) | 768x512 | ~14GB |
HunyuanVideo 13B (GGUF Q5) | Text-to-Video | ~3s (73 frames @ 24fps, up to 129) | 848x480 | ~11GB |
HunyuanVideo 13B I2V (GGUF Q5) | Image-to-Video | ~3s (73 frames @ 24fps, up to 129) | 848x480 | ~11GB |
MiniMax H3 (pruned Int8) | Text, first-frame, last-frame and first+last-frame → Video with its own stereo soundtrack (dialogue, ambience, score) | ~7s offered (175 frames @ 24fps; the model trains to 15s) | 864x480 default, 1344x768 max | 16GB card; measured 6.5 min for a 5s clip at 864x480, 20 steps, on a 16 GB RTX 40-series card |
MiniMax H3 Reference (pruned Int8) | Up to 9 images, 3 clips and 3 audio files → Video with soundtrack (identity, motion, voice, editing) | same | same | same |
Resolution options — 512px, 576px, 720px, 1280px, 1920px (1080p), and custom dimensions (aligned per model)
Quality tiers — Fast (10 steps), Standard (30), High (40), Maximum (50); a model declares the fewest steps it renders well at and a preset below that floor is raised to it (Wan and MiniMax H3: 20). MiniMax H3 also offers turbo speed profiles (8 and 4 steps) through its distilled LoRAs, installed from Manage Video Models.
MiniMax H3 — a video model that generates picture and sound in one pass. Its prompt is compiled into the model's structured format (numbered shots with cut times, speaker ids, tagged dialogue) by Guaardvark, and the Film Crew renders each scene as one spoken window on it. Licensed under the MiniMax H3 Community License, which names the EU, UK, South Korea and USA as territories that need MiniMax's application form; the Video Models modal shows the license and the link, and posts carrying H3 clips add a "Generated with MiniMax H3" line.
Frame interpolation — 1x raw, 2x doubled FPS, 2x + upscale for cinema-quality output
Prompt enhancement — Cinematic, Realistic, Artistic, Anime, or raw
Low VRAM mode — reduces resolution, frames, and inference steps to keep 16GB cards inside budget (mutually exclusive with High consistency); video generation itself needs a 16GB-class card — see docs/HARDWARE.md
Batch processing — queue multiple videos from a prompt list via an in-process worker (one batch at a time; ComfyUI primary, offline CogVideoX fallback)
ComfyUI integration — one-click launch to the node editor for custom workflows; Wan/LTX require ComfyUI. LTX-2.5 needs ComfyUI ≥ 0.32.0 and a one-time license accept on Lightricks/LTX-2.5 (
HF_TOKENin.env); after download, generation stays local.
Audio Studio — Music, FX, and Neural Voice
Three audio backends in one plugin with shared GPU-arbitration so they don't trample each other or fight Ollama for VRAM.
Music generation — ACE-Step v1 (3.5B) for full songs with vocals or instrumental-only mode. Suno-style chip-prompt UX (Genre / Mood / Instrument) with optional LLM "Polish" pass that translates plain English into ACE-Step's tag vocabulary plus a paired negative prompt. ~10 GB VRAM at fp16.
FX Lab — Stable Audio Open for sound effects and short ambient pieces. Light, fast, runs alongside other models.
Neural Voice — Chatterbox as the primary TTS backend, Kokoro as a fast fallback, Piper for narration with 6 voice profiles included. Used for chat narration, voiceover for videos, and the voice-chat conversational mode.
Voice Cloning — opt-in, gated behind an explicit consent prompt before any clone is created or used. Reference clips are kept under your control; the system never auto-clones from incidental audio.
Built-in audio player — generated WAVs and MP3s open in an in-app player modal instead of triggering a browser download. Documents page surfaces audio rows with prompt, model, duration, and a waveform.
Suno export — bulk-export a Suno library into the local DocumentsPage for use with the other generators.
Video Editor — Shotcut-lite Timeline
A built-in non-linear editor for stitching generated clips, layering text, and rendering finished videos — without leaving the app.
Lane | Holds | Source |
Video | one clip per timeline (multi-clip tracking on the roadmap) | Media Library — drag-and-drop |
Text | unlimited overlays, draggable on the preview, properties-panel for size/color/rotation | Add-Text button + properties editor |
Audio | one music or voice clip | Media Library — Audio tab |
Visual trim slider — Material UI range slider bound to source duration, two thumbs for start/end, live monospace readout. No more typing seconds into number inputs.
Tabbed icon-grid library — three tabs (Video / Audio / Images) with counts in the tab labels. 36px tiles, drag from tile to matching timeline track.
Real text overlay rendering — backend uses
ffmpeg drawtext(9 named positions, optional outline + translucent box, proper escaping for colons/quotes/commas). Original is preserved.Keyboard shortcuts — space to play/pause, arrow keys to scrub,
tto add text,delto remove selected,cmd+zfor one-step undo.JobOperationGate — render path checks the gate before grabbing the GPU, so a render won't trample an active video generation or upscaling job.
Standalone Video Text Overlay tool — for the simple one-off case where you don't need a timeline.
Linux & macOS prerequisites: See plugins/video_editor/README.md ("Linux & macOS Setup") for melt + Shotcut install (ffmpeg is already handled by core platform scripts on brew/apt).
GPU Image Upscaling — 4K and 8K Output
Upscale images and video frames to 4K (3840px) or 8K (7680px) with GPU-accelerated super-resolution — eight models from HAT-L (maximum-quality restoration) to anime-tuned Real-ESRGAN variants, two-pass mode, FP16/BF16 precision with torch.compile, frame-by-frame video upscaling, and an optional watch folder. The full model table is in CAPABILITIES.md.
For the complete, enumerated reference (every tool category, exact model support, plugin manifests, page routes, RAG details, self-improvement internals, vision pipeline, dependency reconciler, backup format, advanced settings, etc.) see CAPABILITIES.md.
The sections above cover the experience and differentiators. The rest of this README focuses on getting started, requirements, architecture notes, operations, and contributing.
Quick Start
Python 3.12 is required for the ML stack. Ubuntu 26.04 ships Python 3.14 by default —
./start.shinstalls 3.12 automatically (deadsnakes or uv). Manual installs: use a 3.12 interpreter only.
curl -fsSL https://guaardvark.com/install.sh | bashThis clones to ~/guaardvark (override with GUAARDVARK_HOME=/path) and launches ./start.sh. Re-running it updates an existing install. Prefer doing it by hand? Same thing:
git clone https://github.com/guaardvark/guaardvark.git
cd guaardvark
./start.shFirst run handles everything: Python 3.12, venv, Node dependencies, PostgreSQL, Redis, Ollama, Whisper.cpp, database migrations, frontend build, and all services. Requires your system password once for PostgreSQL setup (and optionally apt packages on fresh Linux installs).
Service | URL (defaults; see |
Web UI | |
API | http://localhost:5000 (macOS: 5055) |
Health Check | http://localhost:5000/api/health (macOS: 5055) |
./start.sh # Full startup with health checks
./start.sh --fast # Reuse venv + node_modules as they are: no installs, no frontend build, no preflight
./start.sh --test # Health diagnostics
./start.sh --plugins # Start all enabled plugins
./start.sh --external-ollama # You run Ollama yourself; never started or stopped by these scripts
./stop.sh # Stop Guaardvark (and only the Ollama that start.sh launched)
./stop.sh --keep-ollama # Stop Guaardvark, leave Ollama running whoever started it
./stop.sh --all # Also stop your own `ollama serve` and the systemd serviceOllama already running before ./start.sh is adopted, not restarted, and left running on
./stop.sh. To make either policy permanent, flip the switches in Settings → Product Profile →
Ollama, or set GUAARDVARK_OLLAMA_KEEP_RUNNING=1 / GUAARDVARK_OLLAMA_EXTERNAL=1 in .env.
Pick a profile
The first start asks what Guaardvark is for here. Creator lists the media workflow — image,
video, audio, Film Crew, LoRA, upscaling — and leaves agents, the knowledge index, outreach and
automation installed but out of the way; Workstation is everything. Either is a starting
point, not a ceiling: switch in Settings → Product Profile, or ./start.sh --profile creator.
Details in backend/profiles/README.md; building a distribution of
your own is docs/EXTENSIONS.md.
Install via PyPI
pip install guaardvarkThe package is the guaardvark command. It talks to a running backend on the configured port, or starts one with start.sh from a checkout it finds through GUAARDVARK_ROOT or the current directory; with no checkout it stops with "Guaardvark installation not found". An MCP client can start the server with guaardvark mcp serve from a pip install plus a checkout.
CLI
Typer + Rich + prompt_toolkit. The PyPI package and the command are guaardvark (llx is a deprecated alias). Tab completion works with or without a leading /; /help imagine shows one command; unknown commands suggest a close match.
guaardvark # Interactive REPL
guaardvark status # System dashboard
guaardvark chat "explain this codebase" # Chat with RAG context
guaardvark search "query" # Semantic search
guaardvark files upload report.pdf # Upload and index
guaardvark plugins list # ComfyUI / Ollama / …
guaardvark gpu status # VRAM and owner lock
guaardvark mcp install --client cursor # Wire Guaardvark into Cursor
guaardvark completion zsh # Shell completion scriptConfig: ~/.guaardvark/cli.json (legacy ~/.llx/config.json is still read). Themes: default, teal, musk, hacker, vader, guaardvark, day, auto. Short terminals get a compact aardvark banner.
REPL Slash Commands (examples)
/imagine <prompt> Generate an image (inline preview in Kitty/iTerm)
/video <prompt> Generate a video from text
/voice <text> Text-to-speech (plays locally)
/agent [on|off|shot] Screen-agent mode; shot = desktop screenshot
/web [images|chat] Open the web UI on the real frontend port
/ingest <path> Index files or directories for RAG
/plugins list|start GPU / service plugins
/gpu status|release VRAM and owner lock
/audio tts|music|sfx Audio Foundry
/swarm run <prompt> Parallel agents in worktrees
/lessons begin|end Lesson pearls
/skills List SKILL.md files
/help [query] Full command reference, or one commandRequirements
Dependency | Version | Notes |
Python | 3.12 only | Backend. 3.13/3.14 not yet supported — the ML stack (numpy<2.0, mediapipe, basicsr/gfpgan) has no wheels for them. |
Node.js | 20+ | Frontend build |
PostgreSQL | 14+ | Auto-installed |
Redis | 5.0+ | Auto-installed |
Ollama | latest | Local LLM inference |
CUDA GPU | 8GB+ VRAM | 16GB recommended for video generation |
Which tier is your machine? See docs/HARDWARE.md for what runs CPU-only, on 8–12 GB, on the 16 GB design target, and with 24 GB+ of headroom.
GPU Memory Guide
Feature | Minimum | Recommended |
Chat + RAG | 4GB | 8GB |
Image generation | 6GB | 12GB |
Wan 2.2 video | 16GB* | 16GB |
CogVideoX-5B video | 16GB | 20GB |
Upscaling | 0.5GB | 2–4GB |
* Wan's weights fit in ~11GB, but the generation preflight requires a 16GB-class card for every current video family — see docs/HARDWARE.md.
Making It Fast
Chat and agent latency are dominated by a few settings, all in Settings unless noted. Defaults favor visibility while you learn the system; flip these once you trust it:
Thinking mode off — extended reasoning (
/thinking, or the chat-thinking default under Settings) adds a long deliberation pass to every turn. Off, simple turns answer in a second or two.Developer toggles off — RAG Debug, Verbose Logging, and LLM Debug each add per-request work. Leave them off outside debugging sessions.
Pick one reliable model and stay on it — every model switch evicts and reloads weights on the GPU (seconds to a minute). A single mid-size model that stays resident beats a bigger one that thrashes.
Mind the VRAM neighbors — renders wait politely for the card, but idle services holding VRAM (voice models, image pipelines) slow everything's admission. The Plugins page shows who's holding what.
Architecture (simplified)
Browser / CLI (PyPI: guaardvark) / MCP Client (Claude Desktop, Cursor, etc.)
| HTTP + WebSocket / stdio MCP
v
Flask (~90+ API modules, auto-discovered) + GraphQL + Socket.IO
|
+-- AgentBrain (3-tier routing: Reflex → Instinct → Deliberation)
|
Service Layer (many modules; plugin sidecars for heavy GPU work)
|-- Agent Executor (ReACT + ~60 tool classes + BrainState)
|-- Screen Control (See-Think-Act-Verify + live reasoning stream)
|-- RAG + Autoresearch + Entity extraction
|-- Self-Improvement (detect/fix/verify/broadcast + guardian)
|-- Generation (image/video/audio/voice/content)
|-- Swarm + Film Crew (isolated worktrees + 5-role pipeline)
|-- Servo + Vision Pipeline
|-- System Mapper / Repo intelligence (AST dependency graphs)
|-- GPU Memory Orchestrator + Plugin runner (CUDA sidecar safety)
\-- Interconnector (multi-machine sync + cluster)
|
+---+---+---+---+---+
v v v v v v
PostgreSQL Redis Ollama Agent Display (:99, on-demand) ComfyUI / Audio Foundry (plugins)
CeleryNotes:
Many components (blueprints, tool registry, plugins) are discovered or declared at runtime.
Exact counts drift between releases; see source and CAPABILITIES.md.
backend/mcp/config.pycontrols the default-deny policy for the MCP server.
Frontend: React 18 · Vite · Material-UI v5 · Zustand · Apollo Client · Monaco Editor · Socket.IO
Core models & engines: Gemma4 / Llama-family / Moondream (vision) · Stable Diffusion · Wan 2.2 / CogVideoX · ACE-Step / Chatterbox / Kokoro / Piper · Real-ESRGAN family + HAT · Whisper.cpp
Roadmap (high-level signals)
See the more detailed view in the project plans and CAPABILITIES for status.
Near term / in flight
Polish + full platform support for supervised outreach (Discord, X, Facebook posting).
Stronger tier-gated memory and conversation context.
Continued video/music pipeline unification and Film Crew robustness.
Plugin GPU auto-orchestration (intent-driven start/stop based on route + VRAM).
Repo intelligence surfaces and more AST-precise agent tools.
Longer term / research
Singing voice cloning (Applio-style) with consent + watermarking.
Cluster metrics + better multi-node UI bridge.
Video editor multi-clip + advanced timeline UX.
Embeddings-backed semantic memory recall (via the gpu_embedding plugin).
Not on the roadmap
Cloud-by-default or SaaS-hosted primary experience. Local-first is the product.
Release & Docs Maintenance (for contributors)
Version source of truth: root
VERSIONfile.backend/app.py, the CLI, and setup.py read it. Avoid hard-coding the version string in README.md, CAPABILITIES.md, or README_zh.md.On release: verify that public screenshots in
docs/screenshots/are up to date, spot-check counts (blueprints via discovery, exposed MCP tools, plugin manifests, CLI catalog), and make sure the "See VERSION" line and CAPABILITIES link are current. Most visual assets live in a separate non-public directory.npm run buildinfrontend/before trusting JSX-related docs or claiming UI completeness (the production Rollup build is strict).AGENTS.md + CLAUDE.md + GROK.md are the orientation files for AI coding sessions in this workspace.
Support the Project
Guaardvark is built with love by a solo developer. If it's useful to you:
Ko-fi (zero fees!)
Star the repo if you find it interesting — it helps with visibility.
Questions, install trouble, or feedback: support@guaardvark.com. Press, partnerships, and business: info@guaardvark.com.
Get Involved
Guaardvark is open source (MIT) and built in public. Whether you want to try the bot, ship a small PR, or hang out with other local-AI builders — here is the short path.
1. Join the community
Where | What |
Discord | The Discord bot ships as a plugin — connect it to your own server for local chat, |
GitHub Issues | Bugs, features, and labeled starter work |
GitHub Discussions | Longer-form questions if enabled |
2. Run it (≈ two commands)
git clone https://github.com/guaardvark/guaardvark.git && cd guaardvark
./start.shWeb UI → http://localhost:5173 · API → http://localhost:5000 (macOS: 5055)
Details: INSTALL.md · agent mental model · full feature list: CAPABILITIES.md
3. Pick a good first issue
Starter issues carry the good first issue label, each with acceptance criteria and a clear out of scope list. When none are open, the safe zones below are the best place to start.
We aim to review serious PRs within 24–48 hours.
4. Safe vs high-risk contribution zones
Safe (great first PRs) | Ask first / high risk |
Agent recipes ( | Agent loop, servo, vision targeting |
Docs, INSTALL, mental-model guides | Self-improvement auto-apply paths |
CLI polish & offline commands | MCP default-deny / security policy |
UI copy, empty states, error messages | Core GPU fork/CUDA plugin runner |
Tests for pure helpers | Production auth / credential handling |
Full setup, style, and PR expectations: CONTRIBUTING.md
For AI coding agents and heavy contributors: read AGENTS.md (required reading order), CLAUDE.md, and GROK.md. They document the self-coding chokepoint (guarded_code_service.py::apply_exact_replacement), project conventions, dead-code handling, and verification habits.
5. Other ways to help (no code required)
Star the repo and share a short demo (screen agent, Film Crew, or Discord
/imagine)Report install friction with GPU model + logs from
logs/(an issue, or email support@guaardvark.com)Suggest recipes or workflows you wish worked out of the box
License
MIT License — Copyright (c) 2025-2026 Albenze, Inc.
"Guaardvark"™ and the Guaardvark logo are trademarks of Albenze, Inc. The MIT License covers the code, not the name; see TRADEMARK.md for what you may do with the name without asking.
Available Tools
51 toolsanalyze_codeBRead-only
Analyze code files for structure, patterns, best practices, and potential improvements
| Name | Required | Description | Default |
|---|---|---|---|
| file_path | Yes | Path to the code file to analyze | |
| analysis_type | No | Type of analysis: 'full', 'structure', 'security', 'performance', 'style' | full |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The readOnlyHint annotation already tells the agent this is a safe read operation. The description adds useful detail about what the analysis covers, but it does not disclose return format, whether the analysis is synchronous, or any limitations such as file size or language support. With annotations covering the safety profile, the description adds some context but not a robust behavioral picture.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single concise sentence with no filler. It front-loads the main action and lists focus areas efficiently. It loses one point because it is slightly generic and could include a brief note about output or usage without becoming verbose.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description should at least hint at what the tool returns, but it only says what is analyzed. The readOnlyHint and 100% parameter coverage lower the burden, but the missing return-value context and lack of usage guidance make the description adequate rather than fully complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so both file_path and analysis_type are already documented in the schema. The description does not add any extra meaning about parameter semantics beyond what the schema provides, so the baseline score of 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb and resource: 'Analyze code files,' and enumerates the analysis dimensions (structure, patterns, best practices, improvements). This clearly distinguishes it from sibling read-only tools like read_code, but it does not explicitly state what makes it different from more focused analysis tools such as read_ast_node or get_dependency_graph.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is provided on when to use this tool versus alternatives. It does not mention that read_code should be used for simply reading a file, or that search_code is better for finding specific symbols. The description implies analysis use but gives no explicit context, exclusions, or alternative recommendations.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
analyze_websiteBRead-only
Analyze a website URL to extract content, SEO information, structure, and provide insights
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | Website URL to analyze (with or without protocol) | |
| query | No | What the user wants to know from the page. When given, the content excerpt is the stretch of the page about it instead of the top of the page (which is often navigation). | |
| analysis_type | No | Type of analysis: 'full', 'seo', 'content', 'structure' (default: 'full') | full |
| include_metadata | No | Include metadata analysis (meta tags, Open Graph, etc.) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already indicate readOnlyHint=true, so the description doesn't need to prove safety. The description adds value by listing what the analysis covers, but it does not disclose behavioral details such as whether the tool fetches the page live, how it handles errors, or what exactly is returned.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, focused sentence with no fluff. It front-loads the primary action and resource, and every phrase contributes to understanding the tool's purpose.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description gives a broad overview of outputs but no output schema exists, so the agent must infer what shape the insights take. With four parameters and no usage guidance, the description is adequate but not fully complete for an agent deciding when and how to invoke it.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so every parameter is already documented. The tool description doesn't add meaningful detail beyond saying the analysis includes content, SEO, structure, and insights, which loosely maps to analysis_type and include_metadata but doesn't explain parameter-specific behavior beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action ('Analyze'), the resource ('website URL'), and the main outputs ('content, SEO information, structure, and insights'). It is specific enough to understand what the tool does, though it does not explicitly distinguish it from sibling tools like fetch_url or web_search.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives no guidance on when to use this tool versus alternatives. It does not mention when to choose analyze_website over fetch_url, web_search, or analyze_code, and provides no exclusions or prerequisites.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
codegenBDestructive
Analyze uploaded code files and generate complete, modified versions with requested changes
| Name | Required | Description | Default |
|---|---|---|---|
| language | No | Programming language (auto-detected from extension if not specified) | auto |
| input_file | No | Path to input file to analyze and modify (optional) | |
| instructions | Yes | Modification instructions or code generation request | |
| idempotency_key | No | Optional. Send the same key when retrying this call: the retry waits for or returns the first run instead of starting a second one. | |
| output_filename | Yes | Output filename for generated code | |
| preserve_structure | No | Preserve original file structure and formatting |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true and readOnlyHint=false, so the agent knows this is a mutating operation. The description adds that it generates 'complete, modified versions' and mentions 'preserve_structure' in the schema, but doesn't disclose what gets overwritten, whether the original file is modified in place, or any side effects. With annotations covering the destructive nature, a 3 is appropriate.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single sentence that front-loads the core action and outcome. It's concise and readable, though it could be slightly more specific about the 'complete, modified versions' phrasing. No wasted words.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a code-generation tool with 6 parameters and no output schema, the description is adequate but not complete. It doesn't explain what the output looks like, whether the original file is modified, or how the 'preserve_structure' option behaves. The annotations cover the destructive hint, but an agent would benefit from knowing the return value or side effects.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all 6 parameters. The description adds no parameter-level detail beyond what the schema provides. Baseline 3 is correct when the schema does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a clear verb ('Analyze' and 'generate') and resource ('uploaded code files'), and mentions producing 'complete, modified versions with requested changes'. It distinguishes itself from read-only siblings like analyze_code and read_code by emphasizing generation/modification. However, it doesn't explicitly name a sibling alternative, so it's clear but not fully differentiated.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage: when you have uploaded code files and need modified versions. It doesn't explicitly state when not to use it or name alternatives like analyze_code for analysis-only or generate_file for new files. The context is implied rather than explicit, so it's adequate but has gaps.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
edit_imageA
Edit an existing image using a natural-language instruction. Use this when the user has attached/uploaded an image (or names one) and asks to add, remove, or change something in it, e.g. 'put a cowboy hat on this character'. Preserves the original subject and only applies the requested edit. If the user did not attach an image, ask them to attach one. Do NOT use this to make a brand-new image from scratch — use generate_image. For a new scene that keeps a face from an attached photo, use generate_identity.
| Name | Required | Description | Default |
|---|---|---|---|
| image | No | Path, URL, or reference of the image to edit. Usually omit this — the image the user just attached is used automatically. | |
| model | No | Image model/backend. Default follows /imagemodel (Settings). 'qwen-image-edit' or 'auto' uses Qwen-Image-Edit when installed; 'kontext' uses FLUX.1 Kontext; other downloaded models use img2img. | auto |
| steps | No | Diffusion steps (more = higher fidelity, slower). Default 28. | |
| instruction | Yes | The edit to perform, e.g. 'put a cowboy hat on this character', 'change the shirt to red'. | |
| idempotency_key | No | Optional. Send the same key when retrying this call: the retry waits for or returns the first run instead of starting a second one. | |
| reference_image_2 | No | Optional second reference (another person or style). Qwen-Image-Edit only. | |
| reference_image_3 | No | Optional third reference. Qwen-Image-Edit only. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare readOnlyHint=false and destructiveHint=false, and the description is consistent with both — no contradiction. Beyond the annotations, it adds a useful behavioral guarantee: 'Preserves the original subject and only applies the requested edit,' clarifying this is a targeted modification rather than a regeneration. It does not cover auth/rate-limit behavior, but the annotation safety profile lowers the bar.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Five sentences, front-loaded with the core purpose, followed by usage context and then alternative routing. Each sentence earns its place, though 'Preserves the original subject' is slightly redundant with the edit semantics and the exclusions could be tightened without losing routing value. Appropriately sized overall.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 7-param tool with full schema coverage, the description fully covers the primary invocation (attached image + instruction), the no-image fallback, and sibling routing. Minor gaps: it never mentions return format or whether edits are asynchronous, and model/reference-image decisions are left entirely to the schema. Acceptable but not exhaustive.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so baseline 3 applies. The description reinforces that the instruction is natural-language and that the flow assumes an attached image, but it adds no parameter-level meaning beyond what the schema already documents (image auto-use, model backends, steps, reference images, idempotency key). It does not need to compensate for any coverage gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Edit an existing image using a natural-language instruction') and explicitly differentiates from siblings: 'Do NOT use this to make a brand-new image from scratch — use generate_image' and 'For a new scene that keeps a face from an attached photo, use generate_identity.' This cleanly disambiguates it from the many image-related sibling tools (generate_image, inpaint_image, remove_background, outpaint_image).
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit when-to-use conditions ('user has attached/uploaded an image... asks to add, remove, or change something'), explicit when-not conditions ('Do NOT use this to make a brand-new image from scratch'), names the alternatives (generate_image, generate_identity), and even covers the failure case: 'If the user did not attach an image, ask them to attach one.' Nothing is left to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
fetch_urlARead-only
Fetch a specific URL and return its page title, meta description, and main text content (up to ~2000 chars). Use this for ANY question about a specific webpage or domain — e.g. 'what's on example.com', 'read https://site.com/page', 'tell me about acme-example.ai'. For open-ended searches without a specific URL, use web_search instead.
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | URL or bare domain to fetch (e.g. 'https://example.com', 'example.com', 'www.example.com'). Protocol is optional — https:// will be added automatically if missing. | |
| query | No | What the user wants to know from the page, in their words. When given, the returned text is the ~2000-character stretch of the page about it; without it, the top of the page, which is often navigation. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only declare readOnlyHint=true; the description adds meaningful behavioral context: truncation at ~2000 chars, the conditional behavior of the query parameter, and the caveat that without a query the top of a page is often navigation. No contradiction with annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences with no wasted words: purpose is front-loaded, usage examples earn their place, and the alternative is named at the end. Ideal length for the information conveyed.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 2-parameter read-only tool with no output schema, the description covers the return shape (title, meta description, main text), the query behavior, and the decision boundary vs web_search. An agent can correctly select and invoke the tool without further clarification.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so baseline is 3. Both url and query are already documented in the schema (including optional protocol handling and what query does). The main description does not add parameter-level detail beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific action ('Fetch a specific URL') with a concrete result (page title, meta description, main text content up to ~2000 chars). It explicitly scopes the tool to known URLs/domains and contrasts with web_search for open-ended queries, distinguishing it from siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit when-to-use ('ANY question about a specific webpage or domain') with concrete example queries and an explicit when-not-to-use with a named alternative ('For open-ended searches without a specific URL, use web_search instead.'). The decision boundary is fully spelled out.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
generate_animationA
Generate a short looping GIF or frame-morph MP4 from a text prompt with motion description, via Stable Diffusion img2img. Use when the user asks to animate, create a GIF, or make a looping frame morph. For a cinema clip from a video model use generate_video instead.
| Name | Required | Description | Default |
|---|---|---|---|
| format | No | Output format: 'gif', 'mp4', or 'both'. Default: 'both'. | both |
| frames | No | Number of frames to generate (2-24). Default: 8. More frames = smoother but slower. | |
| motion | Yes | What moves or changes between frames (e.g. 'walking forward', 'waving hand', 'clouds drifting'). | |
| prompt | Yes | Detailed description of the scene to animate. | |
| strength | No | How much each frame changes from the previous (0.1=subtle, 0.3=moderate, 0.5=dramatic). Default: 0.20. | |
| idempotency_key | No | Optional. Send the same key when retrying this call: the retry waits for or returns the first run instead of starting a second one. | |
| vision_steering | No | Use vision model to guide frame evolution (slower but more coherent). Default: false. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already indicate readOnlyHint=false and destructiveHint=false, so the agent knows this is not a read-only or destructive operation. The description adds useful context about Stable Diffusion img2img and looping output, but it does not disclose async behavior, status polling, or what artifact is returned. That leaves a moderate behavioral gap for a generation tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is three tight sentences: first states the core function, second gives usage triggers, third gives the exclusion. No filler or repetition.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 7-parameter generation tool with no output schema, the description covers purpose, usage, and alternative well. However, it does not describe how the result is returned or whether generation is asynchronous, which would be useful for a tool that creates artifacts. The sibling get_generation_status partially compensates.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3 and the schema documents every parameter. The description loosely mirrors the motion and prompt fields but adds no extra semantics beyond what the schema already provides.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool generates a short looping GIF or frame-morph MP4 from a text prompt with motion description, explicitly naming the output type and generation method. It also differentiates itself from generate_video, so the agent can distinguish it from adjacent tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It provides explicit triggers: 'Use when the user asks to animate, create a GIF, or make a looping frame morph.' It also names the alternative for cinema clips: 'use generate_video instead.' This is textbook when-to-use guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
generate_bulk_csvBDestructive
Generate bulk CSV files with hundreds of pages efficiently using concurrent processing
| Name | Required | Description | Default |
|---|---|---|---|
| topic | Yes | Main topic or subject for content generation | |
| client | No | Client name for personalized content | |
| filename | Yes | Output CSV filename (e.g., 'output.csv') | |
| quantity | Yes | Number of CSV entries/pages to generate (50-1000+) | |
| project_id | No | Project ID for RAG context | |
| word_count | No | Target word count per entry | |
| idempotency_key | No | Optional. Send the same key when retrying this call: the retry waits for or returns the first run instead of starting a second one. | |
| concurrent_workers | No | Number of concurrent generation workers |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true and readOnlyHint=false, so the agent knows this is a mutating operation. The description adds the behavioral trait of concurrent processing and efficiency, which is useful context. However, it doesn't disclose what gets destroyed, whether the operation is idempotent (though the idempotency_key parameter hints at retry behavior), or any side effects beyond file generation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single sentence that front-loads the core purpose and key differentiator (concurrent processing). It is concise and free of filler, though it could arguably be slightly more specific about the output format or scale.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with 8 parameters, 100% schema coverage, and no output schema, the description is adequate but not complete. It doesn't explain return values, error conditions, or the relationship between quantity and pages. The idempotency_key parameter is documented in the schema, but the description doesn't reinforce when retries are needed. Given the destructiveHint annotation, more context about side effects would improve completeness.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all 8 parameters. The description adds the context that 'quantity' relates to 'hundreds of pages' and that 'concurrent_workers' is the mechanism for efficiency, but it doesn't add meaning beyond what the schema provides. Baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('generate'), a resource ('bulk CSV files'), and a key characteristic ('hundreds of pages efficiently using concurrent processing'). It distinguishes itself from the sibling 'generate_csv' by the 'bulk' and 'concurrent processing' qualifiers, though it doesn't explicitly name the sibling.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage for large-scale CSV generation with concurrent processing, which hints at when to use it over generate_csv. However, it does not explicitly state when to use this tool versus alternatives, nor does it mention any exclusions or prerequisites.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
generate_csvBDestructive
Generate a CSV file based on user specifications and data structure instructions
| Name | Required | Description | Default |
|---|---|---|---|
| filename | Yes | Output CSV filename | |
| row_count | No | Number of data rows to generate | |
| idempotency_key | No | Optional. Send the same key when retrying this call: the retry waits for or returns the first run instead of starting a second one. | |
| include_headers | No | Whether to include column headers | |
| data_description | Yes | Description of the data to generate (columns, rows, content type) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare destructiveHint=true, so the description should clarify what destructive behavior occurs (e.g., overwriting an existing file). The description only says 'Generate a CSV file' and does not disclose side effects, file overwriting behavior, or any other behavioral traits. With a destructive hint, the description carries a higher burden and fails to address it.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single sentence with no wasted words. It is front-loaded with the verb and resource, though it could be slightly more informative without becoming verbose.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the destructiveHint annotation and the existence of overlapping siblings like generate_bulk_csv and generate_file, the description is incomplete. It does not explain the destructive nature, when to prefer this tool over alternatives, or what the output looks like. The schema covers parameters, but the description fails to provide the behavioral and selection context an agent needs.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all five parameters. The description adds no parameter-level meaning beyond what the schema provides, so the baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Generate') and resource ('a CSV file'), and mentions it is based on user specifications and data structure instructions. It is clear enough to distinguish from siblings like generate_bulk_csv and generate_file, though it doesn't explicitly name those alternatives.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage for generating CSV files from user specifications, but it does not explicitly state when to use this tool versus generate_bulk_csv or generate_file. The context signals show siblings that overlap, so the lack of explicit routing guidance is a gap.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
generate_enhanced_wordpress_contentCRead-only
Generate WordPress CSV content with RAG-enhanced business intelligence and topic constraints
| Name | Required | Description | Default |
|---|---|---|---|
| topic | Yes | Content topic | |
| client | Yes | Client/company name | |
| row_id | Yes | Unique row ID | |
| website | No | Client website URL | |
| industry | No | Client industry for context | |
| location | No | Geographic location for local SEO | |
| brand_tone | No | Brand voice/tone | professional |
| primary_service | No | Client's primary service offering | |
| target_audience | No | Target audience for content | |
| secondary_service | No | Client's secondary service offering |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description states 'Generate WordPress CSV content', which implies a write/create operation, while the annotations declare readOnlyHint=true. This is a direct contradiction. Additionally, the description offers no other behavioral disclosures such as side effects, permissions, or output characteristics, making it unreliable.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, efficient sentence with the core action and object front-loaded. It contains no fluff or repetition, making it easy to parse quickly. However, the brevity comes at the cost of explanatory detail, which is acceptable for conciseness but not for completeness.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool has 10 parameters and a close sibling, yet the description fails to explain what 'enhanced' means, what 'topic constraints' are, or how it differs from generate_wordpress_content. With no output schema, the agent is left without information about the CSV format or content structure. For a generation tool of this complexity, the description is insufficient.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all parameters are already documented. The description adds a vague reference to 'RAG-enhanced business intelligence' and 'topic constraints' that loosely connects to parameters like topic and client, but does not clarify specific parameter semantics beyond what the schema provides. Baseline 3 is appropriate given the high schema coverage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly identifies a specific verb 'Generate' and a resource 'WordPress CSV content', and adds 'RAG-enhanced business intelligence and topic constraints' to indicate a specialized approach. However, it does not explicitly differentiate from the sibling generate_wordpress_content beyond the word 'enhanced', leaving some ambiguity about what makes it distinct.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use this tool versus its siblings like generate_wordpress_content or generate_csv. There is no mention of conditions, prerequisites, or alternatives, leaving the agent to infer usage from the name alone.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
generate_fileADestructive
Create a brand-NEW output file from a description, written under data/outputs/files. It generates from the description ALONE and never reads any existing file. Do NOT use it to improve, refactor, modify, or produce a new version of an existing or uploaded file — it cannot see that file and would fabricate. For that, use codegen with input_file= (grounded copy) or edit_code (in-place repo change).
| Name | Required | Description | Default |
|---|---|---|---|
| filename | Yes | Relative output filename with extension; nested paths are allowed (e.g., 'frontend/src/page.jsx') | |
| file_type | No | File type hint (code, document, data, config) | auto |
| save_to_disk | No | Whether to save the file to disk | |
| idempotency_key | No | Optional. Send the same key when retrying this call: the retry waits for or returns the first run instead of starting a second one. | |
| content_description | Yes | Description of what the file should contain |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description meaningfully adds behavior beyond the annotations by disclosing that the tool never reads existing files and would fabricate if asked to modify one, which is a critical trust boundary. It also names the output directory. It does not clarify overwrite behavior or the save_to_disk=false case, but the destructiveHint annotation already signals the write-side risk.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and front-loaded with the main purpose. Each sentence earns its place: core function, read-never guarantee, when-not-to-use warning, and replacement routing. There is no filler, repetition of schema defaults, or unnecessary elaboration.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a write tool with no output schema and five parameters, the description covers the output location, generation source, and sibling differentiation, which is most of what an agent needs. The remaining gaps are whether a filename collision overwrites an existing file and what happens when save_to_disk is false, but the schema and annotations cover much of that context.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the description adds useful meaning by binding filename to the data/outputs/files directory and clarifying that content_description is the sole source of generation. This is modest but genuine added context beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The first sentence uses a specific verb, Create, with a clear resource, brand-NEW output file, and a concrete output location, data/outputs/files. The qualifiers brand-NEW and from description ALONE immediately distinguish this tool from codegen and other file generators, so an agent can tell what it does without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly states when not to use the tool: do not use it to improve, refactor, modify, or produce a new version of an existing or uploaded file. It also names two alternatives, codegen with input_file and edit_code, with conditions attached, making the routing decision explicit and actionable.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
generate_imageA
Generate an image from a text prompt. Returns the URL of the generated image. Use when the user asks to create, generate, draw, or visualize an image. For a trained Cast character, ALWAYS pass subject_ids as a separate array of numeric Cast Library IDs (e.g. subject_ids=[26] for Batman 2). Do NOT bury subject_ids inside the prompt string. Putting [batman_2] only in the prompt without subject_ids will NOT load the LoRA. Over this MCP server, omitted arguments default to wait_for_result=false.
| Name | Required | Description | Default |
|---|---|---|---|
| model | No | Model to use. Default 'auto' — recommended; the system auto-picks the best downloaded model for the prompt (usually Z-Image-Turbo or SDXL). With subject_ids, base is taken from the character's train family (Z-Image/SDXL/FLUX). Only override when the user names a specific model: 'krea2-turbo', 'zimage-turbo', 'sd-xl', 'sdxl-turbo', 'realistic-vision', 'epic-realism'. | auto |
| steps | No | Optional sampling steps. The server may raise it to the model's minimum and will say so. | |
| style | No | Image style: 'realistic', 'artistic', 'anime', 'photographic', 'digital-art'. Default: 'realistic'. | realistic |
| width | No | Image width in pixels. Default: 1024. Options: 512, 768, 1024. | |
| height | No | Image height in pixels. Default: 1024. Options: 512, 768, 1024. | |
| prompt | Yes | Scene/action description only (pose, lighting, setting). Do not embed JSON here. For cast characters put identity in subject_ids, not as the whole prompt body. If quoting on-image text, put EXACT words in double quotes — e.g. title "BATMAN". | |
| subject_ids | No | Optional. Numeric Cast Library subject IDs with trained LoRAs to lock identity (e.g. [26]). Separate parameter — never nest this inside prompt. Loads LoRA + trigger + vision bible. Required for consistent characters. | |
| idempotency_key | No | Optional. Send the same key when retrying this call: the retry waits for or returns the first run instead of starting a second one. | |
| wait_for_result | No | True (the default in chat): render now and return the image inline. False: queue the render as an image batch and return its batch id at once; poll get_generation_status for the file. Use False from a remote agent or when the GPU may be busy. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only say readOnly=false, destructive=false, so the description carries the main burden. It adds behavioral detail: the server default is wait_for_result=false, and merely placing [batman_2] in the prompt without subject_ids 'will NOT load the LoRA.' This is meaningful behavior beyond the annotations, though side effects like rate limits or cost are not mentioned.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is four tight sentences: core operation first, return value, use trigger, then the two critical caveats. Every sentence earns its place; the repetition about not burying subject_ids is emphatic and valuable, not filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With 9 parameters but full schema coverage, the description doesn't need to restate param details. It covers the only required input prompt, the output (URL), the key LoRA rule, and the default async behavior. The polling flow for async results is already explained in the wait_for_result schema description, so nothing essential is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, setting the baseline at 3. The description adds real semantic value by insisting subject_ids be a separate numeric array, giving a concrete example (subject_ids=[26]), and warning not to bury it in the prompt. It also reinforces the wait_for_result default in the context of the MCP server, which complements the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with 'Generate an image from a text prompt,' a specific verb and resource, and follows with the output format ('Returns the URL of the generated image'). This clearly separates it from sibling editing tools like edit_image, inpaint_image, or video generation, even without naming them.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives an explicit trigger condition: 'Use when the user asks to create, generate, draw, or visualize an image.' It also explains the key prerequisite for cast characters (always pass subject_ids). It does not name alternatives or state when not to use it, so it stops short of a full exclusion guide.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
generate_music_videoA
Start a music-video project from a song and a visual style. Uploads or attaches the song, writes unique cut prompts, and stops at the approval gate — it does not spend GPU rendering clips. Use when the user asks to make a music video. Pass song as a document id or a path to an audio file.
| Name | Required | Description | Default |
|---|---|---|---|
| name | No | Project name. Defaults to the song filename. | |
| song | Yes | Document id or filesystem path of the song (mp3/wav/flac/ogg). | |
| i2v_model | No | Optional I2V model id. Default: the active video model. | |
| style_prompt | Yes | Visual style for the Director (mood, palette, movement). | |
| idempotency_key | No | Optional. Send the same key when retrying this call: the retry waits for or returns the first run instead of starting a second one. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Adds significant context beyond annotations: it stops at an approval gate and does not spend GPU rendering, telling the agent this is a setup step, not a final render. It also states it uploads/attaches the song and writes cut prompts. Given annotations only indicate non-read-only and non-destructive, this behavioral detail is valuable.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two concise sentences with clear structure. The first sentence covers purpose and key behavior, the second gives usage trigger and parameter guidance. No unnecessary words, front-loaded with the most critical information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Covers the essential context: what it does, when to use, and a key limitation (no rendering). It does not describe the return value or what happens after the approval gate, but given no output schema and the tool's moderate complexity, this is not a critical gap. Could be improved by mentioning the output or next steps, but it is reasonably complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so parameters are already fully documented. The description repeats the song format instruction but adds no new meaning for other parameters like style_prompt or idempotency_key. It does not explain the idempotency behavior beyond the schema. Minimal added value, consistent with baseline for high schema coverage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action ('Start a music-video project from a song and a visual style') and details the workflow (uploads/attaches song, writes cut prompts, stops at approval gate). It explicitly distinguishes itself from siblings by noting it does not spend GPU rendering clips, which separates it from video generation tools like generate_video or generate_animation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides an explicit trigger: 'Use when the user asks to make a music video.' It also clarifies scope by noting it stops before rendering, implying other tools handle rendering. However, it does not name alternative tools explicitly, but the instruction is clear enough for an agent to route correctly.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
generate_videoA
Queue a video clip from a text prompt using a local video model. Returns immediately with a batch id and Studio deep-link so long jobs are not false-timeouted. Use when the user asks to create, generate, or make a video. Pick model='minimax-h3-int8' (or audio=true) for a clip with its own soundtrack and spoken dialogue; give first_image / last_image to animate between frames; reference_images and reference_audio lock a person, look or voice on the reference build. For short looping frame-morph animations use generate_animation instead.
| Name | Required | Description | Default |
|---|---|---|---|
| audio | No | Require a model that generates its own soundtrack (MiniMax H3). Fails on a silent family. | |
| model | No | Video model id from the registry (e.g. wan22-5b, minimax-h3-int8). Default: the installed default. | |
| style | No | Prompt style: cinematic, realistic, artistic, anime, 3d_animation, stop_motion, hand_drawn, western_cartoon, none. A model may not offer every style; its capability record lists prompt_styles. | |
| prompt | Yes | Description of the video to generate (scene, subject, motion, style, any spoken lines). | |
| duration_s | No | Clip length in seconds; clamped to the model's declared range. Overrides duration_frames. | |
| last_image | No | Document id or path of the last frame; needs a model with first+last-frame mode. | |
| first_image | No | Document id or path of the first frame (image-to-video). | |
| aspect_ratio | No | 16:9, 9:16, 1:1, 4:3, 3:2, 21:9 or 3:4; must be one the model declares. | |
| speed_profile | No | A speed profile the model declares (e.g. turbo-8). | |
| duration_frames | No | Number of frames (legacy). Default 49. Clamped to the model's longest clip. | |
| idempotency_key | No | Optional. Send the same key when retrying this call: the retry waits for or returns the first run instead of starting a second one. | |
| reference_audio | No | Document id or path of a voice or music reference; needs the reference build. | |
| wait_for_result | No | If true, poll until the clip finishes (or timeout). Default false: enqueue and return a Studio Video Gen / Jobs link immediately. | |
| reference_images | No | Document ids or paths of reference images (identity, look); needs the reference build. | |
| num_inference_steps | No | Inference steps. Omit to use the model's default; a value below the model's floor is raised to it. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only declare readOnlyHint=false and destructiveHint=false, so the description carries the real burden and does: it discloses async queueing, immediate return with batch id and deep-link, avoidance of false timeouts, and idempotency behavior. It stops short of stating auth/permission requirements or failure modes for silent-model audio requests, which keeps it just under a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core action and return behavior, then usage, then model-selection guidance, then the sibling alternative. Dense for three sentences but every clause carries selection value; no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 15-parameter async generation tool with no output schema, the description covers the return contract (batch id, Studio deep-link), the sync/async toggle, and how to pick models and references. Nothing essential to correct invocation is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so 3 is the baseline. The description goes beyond the schema by tying parameters to outcomes: model='minimax-h3-int8' or audio=true yields a soundtrack/dialogue clip, first_image/last_image animate between frames, and reference_images/reference_audio lock a person or voice on the reference build. That relational guidance is genuinely additive.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Queue), resource (video clip), and input (text prompt) plus the execution model (local video model). Explicitly distinguishes itself from the sibling generate_animation, so an agent can route without opening a schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives a clear trigger ('user asks to create, generate, or make a video'), names the alternative tool with the condition that selects it ('short looping frame-morph animations use generate_animation'), and tells the agent how to choose a model/settings for audio, frame interpolation, or identity locking.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
generate_wordpress_contentARead-only
Generate a WordPress-compatible CSV row with SEO-optimized content for a specific topic and client
| Name | Required | Description | Default |
|---|---|---|---|
| topic | Yes | Content topic or subject to write about | |
| client | Yes | Client/company name for content personalization | |
| row_id | Yes | Unique row ID for the CSV entry | |
| website | No | Client website URL for context | |
| word_count | No | Target word count for the content section |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, so no external mutation is expected. The description adds that the output is a CSV row with SEo-optimized content, but provides no further behavioral detail about how the content is generated, what format the row takes, or any limitations. It doesn't contradict the annotations, and with the annotation covering side effects, a 3 is appropriate.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, front-loaded sentence with no filler. It states the core purpose immediately and is appropriately concise for a tool of this simplicity.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description usefully identifies the output as a CSV row, but doesn't specify the row's column structure, whether it returns a string or file, or how it differs from the bulk/file-generation siblings. For the moderate complexity and the rich sibling context, this is adequate but has clear gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3. The description echoes 'topic' and 'client' from the schema but adds little beyond that; it doesn't explain how row_id, website, or word_count affect the generated CSV row. The schema itself carries most of the parameter meaning.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific verb ('Generate'), resource ('WordPress-compatible CSv row'), and two defining characteristics ('SEO-optimized content,' 'for a specific topic and client'). It clearly states what the tool produces, but it doesn't differentiate from its close sibling 'generate_enhanced_wordPress_content' or 'generate_bsv', so it misses top-tier sibling distinction.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage: when you need a single WordPress-compatible CSV row for a given topic and client. However, it doesn't state when to prefer this over alternatives like 'generate_enhanced_wordPress_content' or 'generate_bulk_csv', nor does it mention any exclusions or preconditions. This is implied usage, not explicit guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_dependency_graphARead-only
Retrieve the file-level import dependency graph for a given folder ID. This returns a JSON string mapping files to the files they import. Use this tool to trace dependencies and understand how files interact.
| Name | Required | Description | Default |
|---|---|---|---|
| folder_id | Yes | The integer ID of the Code Repository folder. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, and the description's 'Retrieve' and 'returns' are consistent. The description adds the output format (JSON string) which is useful, but it does not disclose any other behavioral traits such as performance limits, depth of analysis, or error conditions. Given the annotation coverage, a score of 3 is appropriate.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences with no filler. The primary action is front-loaded, and the return format is mentioned immediately. Every sentence contributes essential information without redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With a single parameter, a read-only annotation, and no output schema, the description adequately explains the return value ('a JSON string mapping files to the files they import'). It does not cover edge cases or error handling, but for a simple read-only tool this is sufficient. A 4 reflects good completeness given the low complexity.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema covers the only parameter (folder_id) with a description, achieving 100% coverage. The description only restates 'given a folder ID' without adding any new meaning, syntax, or constraints beyond what the schema already provides, so it meets the baseline but adds no extra value.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool retrieves a file-level import dependency graph for a folder ID, specifying the output as a JSON string mapping files to their imports. It is specific about the resource and action, but it does not explicitly differentiate from sibling tools like get_repository_map or map_codebase, so it misses the full distinction criterion.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides a clear usage context: 'Use this tool to trace dependencies and understand how files interact.' However, it does not mention when not to use it or suggest any alternative tools, so it lacks explicit exclusions or sibling routing.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_document_outlineARead-only
Show the internal structure of one indexed document — its sections or pages, in order, with passage counts. Use this after list_documents to see what is inside a document before retrieving from it.
| Name | Required | Description | Default |
|---|---|---|---|
| source_filename | Yes | The document filename, as shown by list_documents. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description goes beyond the readOnlyHint annotation by describing what the tool returns: sections or pages in order, with passage counts. This tells the agent what kind of structural data to expect without needing an output schema. It adds useful behavioral context beyond annotation metadata.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two tight sentences: the first states what the tool does, and the second gives context for when to use it. There is no filler, and the most important information is front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
This is a simple one-parameter read-only tool, and the description plus schema fully cover what an agent needs to call it correctly: what the tool returns, how to identify the document, and the recommended workflow context. No output schema is necessary because the return shape is described in plain language.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema already documents source_filename with 100% coverage, including 'as shown by list_documents.' The description adds only the idea of 'one indexed document,' which is marginal beyond the schema. Since the schema carries the burden, baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: 'Show the internal structure of one indexed document — its sections or pages, in order, with passage counts.' It clearly distinguishes this from list_documents, which lists documents, and from retrieving content from a document. The scope is unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly instructs 'Use this after list_documents to see what is inside a document before retrieving from it,' giving clear sequencing and context for when this tool is appropriate. It does not explicitly name or exclude siblings like read_document_section, but the 'before retrieving' phrasing implies the distinction.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_generation_statusARead-onlyIdempotent
Report the state of a queued generation: an image batch (ImageBatch_...) from generate_image with wait_for_result=false or the batch image route, or a video batch from generate_video. Returns status, progress, and the URL of each finished file. Use after a queued generate call, or when the user asks whether a render is done.
| Name | Required | Description | Default |
|---|---|---|---|
| batch_id | Yes | The batch id a generate tool returned (e.g. ImageBatch_09-11-2026_132620_013). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and idempotentHint=true, so the description does not need to restate safety. It adds value by disclosing the return shape: 'status, progress, and the URL of each finished file', which is particularly useful given there is no output schema. It also implies polling-friendly behavior by indicating the tool can be called to check completion.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, front-loaded with the core purpose, followed by the source types, return values, and usage guidance. Every clause is necessary and there is zero redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a one-parameter status-check tool with no output schema and read-only/idempotent annotations, the description covers what the tool does, what it returns, and when to use it. It also correctly handles both image and video batch sources. No critical information missing for an agent to invoke it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and the schema already explains batch_id as 'The batch id a generate tool returned (e.g. ImageBatch_09-11-2026_132620_013).' The tool description does not add meaning beyond the schema, only referencing the parameter in passing. Baseline 3 is appropriate since the schema carries the full semantic load.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names the exact verb-resource pair ('Report the state of a queued generation'), specifies the two source categories (ImageBatch_... and video batches), and details the return contents (status, progress, file URLs). It clearly differentiates from sibling status tools by identifying the specific generation calls it covers.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides explicit usage context: 'Use after a queued generate call, or when the user asks whether a render is done.' It does not name alternatives or exclusions, but for this tool the correct usage is clear from the generation-source context. Sibling status tools like swarm_status or media_status serve different domains, so no confusion is likely.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_repository_mapARead-only
Retrieve the PageRank-based architectural repository map for a given folder ID. This map shows the most important functions and classes in the codebase and their relationships. Use this tool to get a high-level understanding of a Code Repository.
| Name | Required | Description | Default |
|---|---|---|---|
| folder_id | Yes | The integer ID of the Code Repository folder. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The readOnlyHint already indicates this is a safe read. The description adds meaningful context beyond the annotation by explaining that the map is PageRank-based, prioritizes important functions/classes, and shows relationships, which helps set expectations about output style.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two concise sentences with the core action and resource front-loaded. Every sentence earns its place and no filler or repetition exists.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple, read-only, one-parameter tool, the description gives enough information to understand what the tool does and when to use it. It could be strengthened by clarifying how it differs from map_codebase and get_dependency_graph, but it is not incomplete for basic invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The single parameter folder_id is fully described in the schema with 100% coverage, so the description adds no new semantic detail. The description merely restates 'given folder ID' without enriching the parameter meaning.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly identifies the resource: a PageRank-based architectural repository map for a given folder ID, showing important functions/classes and relationships. It is specific and actionable, though it does not explicitly distinguish itself from sibling tools like map_codebase or get_dependency_graph.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description states a clear use case: get a high-level understanding of a Code Repository. It provides context for when to call the tool, but does not mention exclusions or alternatives, so it stops short of full guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
inpaint_imageA
Change or remove something in an attached photo from a natural-language instruction ('remove the coffee cup', 'replace the sky with sunset'). Uses Qwen-Image-Edit when installed, else FLUX Kontext. For extending the canvas use outpaint_image. For a brand-new scene of a person's face use generate_identity.
| Name | Required | Description | Default |
|---|---|---|---|
| image | No | Path of the photo. Usually omit — the attached image is used. | |
| steps | No | Diffusion steps. Default 20. | |
| instruction | Yes | What to change or remove. | |
| idempotency_key | No | Optional. Send the same key when retrying this call: the retry waits for or returns the first run instead of starting a second one. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=false and destructiveHint=false, so the description doesn't need to restate those. It adds a meaningful behavioral detail: the underlying model is Qwen-Image-Edit when installed, else FLUX Kontext. This helps the agent anticipate variability in output quality or availability. It doesn't contradict the annotations. The only gap is no explicit statement about what the function returns, but that is not a contradiction.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is tight: three sentences. The core purpose is front-loaded, the model detail is secondary, and the sibling routing comes last. There is zero filler; every sentence adds distinct value. This is exemplary for a tool description.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For an image-editing tool with no output schema, the description covers purpose, model selection, and routing to alternatives. It does not explicitly state what the tool returns (presumably the edited image), but that is reasonably inferred. Given the clear purpose and the existence of sibling tools for related tasks, an agent has enough context to decide when and how to call this tool. Minor omission: no note on idempotency behavior, but that is documented in the parameter schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so every parameter (image, steps, instruction, idempotency_key) already has a description. The tool description itself does not add parameter-level detail beyond the schema; it only gives example instructions. Baseline 3 is appropriate when the schema carries the semantic load, and the description does not need to compensate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb-resource pair ('Change or remove something in an attached photo') and gives concrete examples ('remove the coffee cup', 'replace the sky with sunset'). It explicitly names sibling tools (outpaint_image, generate_identity) and what they do differently, so an agent can immediately tell this tool apart without inspecting schemas.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It provides explicit routing: 'For extending the canvas use outpaint_image. For a brand-new scene of a person's face use generate_identity.' This tells the agent when NOT to use this tool and names the alternatives, which is the strongest form of usage guidance. The examples also clarify the kind of instructions to supply.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
inspect_gpuARead-onlyIdempotent
Inspect live GPU state: nvidia-smi, the exclusive lock (Ollama vs video), orchestrator model slots, and which plugins are running. Use when the user says 'debug GPU issues', 'what's using VRAM', 'GPU status', or 'OOM'.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and idempotentHint=true, covering the safety profile. The description adds value by specifying exactly what live state is inspected (nvidia-smi, lock, model slots, plugins), which goes beyond the annotations and gives the agent a clear picture of the tool's scope.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two concise sentences. The first states the purpose and enumerates the inspection targets, the second gives usage triggers. Everything earns its place with no fluff.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter, read-only inspection tool with safety annotations, the description fully covers the purpose and when to use it. The inspection targets imply the nature of the output, and the annotations handle safety, so nothing essential is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters, so schema coverage is trivially 100%. Per the guidelines, a baseline of 4 is appropriate for 0-parameter tools; the description needs no parameter explanation, and it doesn't waste words on them.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose: 'Inspect live GPU state' and enumerates the specific aspects (nvidia-smi, exclusive lock, orchestrator model slots, plugins running). This is specific and distinguishes it from any sibling tool that might touch GPU-related concerns.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly lists trigger phrases: 'debug GPU issues', 'what's using VRAM', 'GPU status', or 'OOM'. This gives an agent concrete cues for when to invoke the tool, and no alternative tool in the sibling list covers GPU inspection, so no exclusions are needed.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_code_filesARead-only
List files and directories to understand project structure. Returns a formatted tree view of the directory contents. Use this to explore the codebase and find relevant files.
| Name | Required | Description | Default |
|---|---|---|---|
| directory | No | Relative path from project root (default: 'frontend/src') | frontend/src |
| max_depth | No | Maximum directory depth to show (default: 5) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The readOnlyHint annotation already declares this as a safe read operation. The description adds useful behavioral context by stating it returns a formatted tree view of directory contents, which is not evident from the schema or annotations alone.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two focused sentences convey the action, output format, and usage intent with no wasted words. The main purpose is front-loaded, making it easy for an agent to quickly understand the tool.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple read-only list tool with fully documented parameters and no required arguments, the description is complete. It covers what the tool does, what it returns, and when to use it. The absence of an output schema is mitigated by the explicit 'formatted tree view' statement.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Input schema has 100% coverage, with directory and max_depth fully described including defaults. The description adds no new parameter details, but it also does not need to given the schema's completeness, so the baseline score of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('List files and directories') and the output form ('formatted tree view'). It clearly differentiates from sibling tools like read_code, search_codebase, and list_code_repositories, which serve different purposes.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly describes the intended use case: 'Use this to explore the codebase and find relevant files.' It provides clear context for when to invoke this tool, though it does not name alternatives or give exclusion criteria.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_code_repositoriesARead-only
List all Code Repository folders that have been marked as such (is_repository=True) and analyzed. Returns id, name, path, and whether repo_metadata is available. Use this first when the user refers to 'the uploaded code', 'guaardvark upload folder', 'the code repo in data/uploads/Code', or similar to discover the folder_id(s) needed for get_repository_map, get_dependency_graph, read_ast_node, etc. This helps the agent get a full picture of available code repositories before analyzing or editing.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare readOnlyHint=true, and the description adds useful behavioral context by noting that only folders with is_repository=True and analysis completed are returned. It also discloses the exact output fields: id, name, path, and repo_metadata availability, which goes beyond the annotation's safety hint.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is concise and front-loaded, beginning with the core listing behavior and output before moving to usage scenarios. The example phrases are helpful, though the description is slightly longer than strictly necessary and includes a minor typo ('guaardvark'), which keeps it from a perfect score.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter read-only list tool, the description provides everything needed: the filtering condition, the returned fields, the triggering user language, and the purpose of the returned folder ids. No output schema is present, so including the return fields is especially valuable. The description is complete for the tool's complexity.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters, so there are no parameter semantics to document. Per the baseline for zero-parameter tools, the description does not need to compensate for missing schema documentation, and it correctly focuses instead on what the tool returns and when to call it.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'List all Code Repository folders', and clearly scopes the result to those 'marked as such (is_repository=True) and analyzed'. It also names downstream tools like get_repository_map, get_dependency_graph, and read_ast_node, which distinguishes this tool from sibling listing tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives explicit context for when to use this tool: 'Use this first when the user refers to...' and offers concrete example phrases like 'the uploaded code'. It does not explicitly state when not to use it or compare it directly to a sibling like list_code_files, but the guidance is clear enough to route an agent correctly.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_documentsARead-only
List the documents currently in the knowledge base, with how many passages each contributes. Use this to find out what the knowledge base actually contains before searching it. Supports paging and a name filter.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | How many documents to return (default 40, max 200). | |
| offset | No | Skip this many documents — use to page through a long list. | |
| name_contains | No | Only list documents whose filename contains this text. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, and the description is consistent. It adds useful context about the current knowledge-base state, passage counts, and paging/filter support. No side effects or auth details need disclosure for a simple read-only listing.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, each purposeful: the core action, the recommended use case, and supported options. Information is front-loaded and there is no redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple read-only list tool with fully documented optional parameters and no output schema, the description gives enough to decide when to call it and what to expect (documents plus passage counts). The absence of exact return structure is a minor gap given the low complexity.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema describes all three parameters with 100% coverage, so the baseline is 3. The description only recaps paging and name filtering without adding new parameter semantics or format details.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Clearly identifies the verb (list), resource (documents in the knowledge base), and a distinguishing detail (per-document passage counts). This separates it from sibling search and read tools, so an agent can accurately select it without inspecting other schemas.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides explicit usage context: 'Use this to find out what the knowledge base actually contains before searching it.' This makes the intended scenario clear, though it does not explicitly name alternatives or state when not to use it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
map_codebaseARead-only
Run the System Mapper (same snapshot as /system-map) and return stats plus ranked findings. Use when the user says 'use the system mapper', 'map the codebase', 'what's wrong with this repo', or 'constellation findings'.
| Name | Required | Description | Default |
|---|---|---|---|
| root | No | Code root. Defaults to GUAARDVARK_ROOT. Must stay inside that tree. | |
| limit | No | Max findings to return (default 15). | |
| refresh | No | Ignore the 5-minute disk cache and recompute. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The readOnlyHint annotation already establishes the operation is safe. The description adds useful behavioral context: it runs the same snapshot as /system-map and returns specific output types. No contradictions with annotations are present.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single sentence packs the core action, output, and usage triggers with no filler. It is appropriately concise and front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the read-only annotation, fully described optional parameters, and explicit usage examples, the description covers the essential invocation context. It does not detail the exact structure of 'stats' or 'ranked findings', but that is not required for correct selection and invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so each parameter (root, limit, refresh) is already documented in the schema. The description does not add extra parameter-level nuance, fitting the baseline of 3.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the specific action ('Run the System Mapper') and resource, and specifies the output ('stats plus ranked findings'). It also lists distinct trigger phrases that help separate it from sibling tools like analyze_code or get_repository_map.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives explicit usage triggers: 'Use when the user says...' with concrete phrases. It does not mention when not to use it or name alternatives, but the trigger guidance is strong enough for an agent to select it appropriately.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
media_controlB
Control media playback. Actions: pause, stop, next (skip to next track), previous (go back), toggle (play/pause toggle).
| Name | Required | Description | Default |
|---|---|---|---|
| action | Yes | Action to perform: pause, stop, next, previous, toggle | |
| player | No | Specific player name (default: auto-detect) | |
| idempotency_key | No | Optional. Send the same key when retrying this call: the retry waits for or returns the first run instead of starting a second one. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=false and destructiveHint=false, but the description adds no extra behavioral context such as effects on the current player, prerequisites (e.g., an active media session), or error behavior when no player is available. The action names imply mutation, but nothing beyond that is disclosed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two short sentences with zero filler. The main purpose is front-loaded first, and the action list follows immediately in a compact, scannable format. Every word contributes.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple tool with 100% schema coverage and clear annotations, the description is mostly sufficient for basic invocation. The main gaps are absence of usage boundaries versus sibling media tools and no behavioral notes about player autodetection or failure modes, leaving the agent to infer edge-case behavior.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, but the description enriches the action parameter beyond the schema by clarifying the meanings: next means "skip to next track", previous means "go back", and toggle means "play/pause toggle". The player and idempotency_key parameters are already well documented in the schema, so the description adds meaningful value where needed.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description gives a clear verb and resource: "Control media playback," then enumerates the specific actions (pause, stop, next, previous, toggle). This is unambiguous about what the tool does, though it does not explicitly differentiate itself from sibling tools like media_play, media_volume, or media_status.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The action list strongly implies use for playback transport control (pause, skip, toggle) rather than volume or status queries. However, the description never explicitly says when to prefer this tool over media_play or media_status, so the guidance is implied rather than stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
media_playA
Play music by searching for songs, artists, albums, or genres. If user says 'play some music' or similar generic request without a specific artist/song, use query='music' to play all music. If no query is given at all, resumes current playback. Launches VLC with matching music files from the configured music directory.
| Name | Required | Description | Default |
|---|---|---|---|
| query | No | Search query for music (e.g. 'Alice in Chains', 'jazz'). Use 'music' for generic 'play some music' requests. Omit entirely to resume paused playback. | |
| shuffle | No | Shuffle the playlist | |
| directory | No | Play all music in a specific directory path | |
| idempotency_key | No | Optional. Send the same key when retrying this call: the retry waits for or returns the first run instead of starting a second one. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations are minimal (readOnlyHint=false, destructiveHint=false), so the description carries the burden of behavioral disclosure. It states that VLC is launched and that playback resumes when no query is given, which are side effects. It does not mention potential interruption of current playback or error behavior, but for a play command, these are implied. The description adds meaningful context beyond the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is three sentences, front-loading the core purpose and then covering usage patterns and the underlying mechanism. Every sentence earns its place without excessive detail. It could be slightly tighter (e.g., merging the VLC mention into the first sentence), but it remains efficient and readable.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with four optional parameters and no output schema, the description covers the main scenarios (specific search, generic request, resume) and the underlying action (launching VLC). It does not explain what happens when no matching music is found or how directory interacts with query, but these are edge cases the schema partially addresses. Overall, an agent has enough to invoke the tool correctly in most cases.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3. The description repeats the key guidance for the query parameter (use 'music' for generic, omit to resume), which is already in the schema. It adds no new meaning for shuffle, directory, or idempotency_key, all of which are adequately described in the schema. Thus, the description provides marginal added value.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool plays music by searching for songs, artists, albums, or genres, which is a specific verb-resource pairing. It also distinguishes itself by explaining edge-case behaviors (generic 'play some music' and resume without query). While it doesn't explicitly name sibling tools, the focus on launching playback and the mention of VLC makes its role distinct from media_control (likely for pause/skip) and media_volume.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives explicit guidance for generic requests (use query='music') and for resuming playback (omit query). It implies when to use this tool versus alternatives by its focus on starting playback, but it doesn't explicitly state 'use media_control for pause' or list exclusions. Still, the context is clear enough for an agent to select it correctly.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
media_statusARead-only
Get current playback status: what's playing, track info (title, artist, album), player state, and volume level.
| Name | Required | Description | Default |
|---|---|---|---|
| player | No | Specific player name (default: auto-detect) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already provide readOnlyHint=true, and the description reinforces a non-mutating read operation with no contradiction. It adds value by specifying the return contents (title, artist, album, player state, volume level), which is useful behavioral context beyond the schema. It does not discuss edge cases like no active player, but the read-only nature is well covered.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is one concise sentence with the core action front-loaded and the detailed contents following immediately. Every word contributes useful information, with no redundancy or filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple read-only status tool, the description covers the main return areas and aligns with the optional player parameter and readOnlyHint. It is slightly thin on edge-case behavior (e.g., what happens if no player is active, or how auto-detection behaves), but given the tool's simplicity, the description is largely sufficient.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% for the single optional 'player' parameter, so the schema already documents it. The description does not add any additional semantics about the parameter, which fits the baseline of 3 for high schema coverage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb-resource pair ('Get current playback status') and enumerates exactly what the tool reports: track info, player state, and volume level. It is clearly distinct from siblings like media_play, media_control, and media_volume, so an agent can identify it as the status-query tool without ambiguity.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implicitly tells the agent to use this when playback status is needed, rather than controlling or changing media. However, it does not explicitly state when not to use it or name alternative tools, so the usage guidance is inferred rather than stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
media_volumeA
Get or set the system audio volume level (0-100). Supports absolute ('50'), relative ('+10', '-10'), 'mute', and 'unmute'.
| Name | Required | Description | Default |
|---|---|---|---|
| level | No | Volume level: absolute (e.g. '50'), relative ('+10', '-10'), 'mute', or 'unmute'. Omit to get current volume. | |
| idempotency_key | No | Optional. Send the same key when retrying this call: the retry waits for or returns the first run instead of starting a second one. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond annotations, the description discloses that the tool both reads and mutates system volume and enumerates accepted modes ('absolute', 'relative', 'mute', 'unmute'). The annotations already signal non-read-only and non-destructive behavior, so the description adds useful context without needing to restate everything.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences front-load the core operation and immediately provide concrete syntax examples. There is no filler or redundant explanation.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple tool, the description covers operation, volume range, and accepted value formats; idempotency_key is fully documented in the schema. The only minor gap is that the return value for 'get' is not explicitly described, but it is reasonably inferable.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents both parameters. The description largely restates the level parameter's allowed formats without adding new meaning, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific resource ('system audio volume level') and explicit verbs ('Get or set'), supported by a 0-100 scale. This clearly distinguishes media_volume from siblings like media_play, media_control, and media_status.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The intended use is clear: adjust or read system volume. It doesn't explicitly name alternatives or exclusions, but no sibling appears to handle volume, so this is sufficient context for selecting the tool.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
outpaint_imageA
Expand an attached photo in one or more directions and fill the new area so it matches the scene. Use when the user says extend, expand the canvas, or outpaint. Prefer Qwen-Image-Edit when installed.
| Name | Required | Description | Default |
|---|---|---|---|
| top | No | Pixels to add on the top. | |
| left | No | Pixels to add on the left (multiples of 8). | |
| image | No | Path of the photo. Usually omit — the attached image is used. | |
| right | No | Pixels to add on the right. | |
| steps | No | Diffusion steps. Default 20. | |
| bottom | No | Pixels to add on the bottom. | |
| instruction | No | Optional fill direction, e.g. 'continue the forest to the left'. | |
| idempotency_key | No | Optional. Send the same key when retrying this call: the retry waits for or returns the first run instead of starting a second one. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations are only non-readonly and non-destructive hints, which provide little safety detail. The description adds that new area should match the scene, but it does not disclose whether the original image is modified, whether the operation is asynchronous, cost/time implications, or what the output will be. There is no contradiction with annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is very concise and front-loaded with the core action. Every sentence earns its place: purpose, usage trigger, and implementation preference. There is no filler or redundant repetition of schema data.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema and a non-trivial image-generation tool, a sentence about what the tool returns or whether it runs asynchronously would round out the description. The full parameter schema, clear trigger phrases, and attached-image handling make the core operation understandable enough for an agent to invoke it.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema already has 100% parameter description coverage, so top, left, right, bottom, steps, instruction, and idempotency_key are well-documented there. The description only adds the general notion of one or more directions and an attached photo, which does not materially clarify parameter behavior beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool expands an attached photo in one or more directions and fills the new area to match the scene. This distinguishes it from image-editing siblings like inpaint_image and edit_image, which address different modifications.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly says to use the tool when the user says extend, expand the canvas, or outpaint, which gives clear invocation triggers. It does not name alternative sibling tools or provide exclusion criteria, but the guidance is sufficient for basic routing.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
outreach_draft_postA
Draft a social outreach comment or share post (does NOT post). Platforms: reddit, discord, facebook, twitter, youtube. For mode='comment' you must supply either thread_context (the OP/comment/video description) or target_url (we'll scout it — for YouTube URLs, scrapes the video title + description). For mode='share' supply share_target (e.g. 'r/SideProject') and optionally share_link (defaults to guaardvark.com). The draft lands in the queue at status='drafted' for human approval — nothing posts until the user approves it in the OutreachPage UI.
| Name | Required | Description | Default |
|---|---|---|---|
| mode | No | 'comment' or 'share' | comment |
| tone | No | Optional tone preset: default, engaging, technical, casual, formal, humorous | |
| platform | Yes | One of: reddit, discord, facebook, twitter, youtube | |
| share_link | No | Link to share; defaults to https://guaardvark.com | |
| target_url | No | URL of the thread; if thread_context is missing we scout it | |
| feature_hint | No | Override auto-detected feature angle (e.g. 'video_gen', 'rag') | |
| include_link | No | Comment mode only. When true, the persona includes a guaardvark.com link where it fits naturally. The persona still self-grades and may return grade<0.7 if the link would feel forced (the human reviewer would rather hold than ship spam). Defaults to false. | |
| share_target | No | Where the share post goes, e.g. 'r/SideProject' (share mode) | |
| thread_context | No | OP body + top comments concatenated (comment mode) | |
| idempotency_key | No | Optional. Send the same key when retrying this call: the retry waits for or returns the first run instead of starting a second one. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the annotations (readOnlyHint=false, destructiveHint=false), the description reveals key behavioral facts: it does NOT post, it scours target_url (including scraping YouTube metadata), and it places the result in a queue at status='drafted' for human approval. These are meaningful side-effect and workflow details that are not captured in the annotations or schema.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and front-loaded, with the most important constraint ('does NOT post') in the first sentence. Mode-specific requirements and the approval flow are packed into three additional sentences with no filler or repetition.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 10-parameter tool with no output schema, the description covers platform scope, mode constraints, required inputs, defaults, scraping behavior, and post-draft workflow. The only notable gap is that it does not describe the return value shape, which the absence of an output schema leaves to the agent to infer. Otherwise it is complete enough for correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3. The description adds real value beyond the schema by explaining mutually exclusive parameter requirements (thread_context vs target_url for comment mode), defaults for share_link, and mode-specific semantics for share_target and include_link. It does not enumerate every parameter, but the schema already handles those.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'Draft a social outreach comment or share post (does NOT post).' It lists the exact platforms and immediately contrasts with post/publish operations, making the tool's boundary clear. This distinguishes it from siblings like request_publish and outreach_reject_draft without needing to inspect schemas.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives explicit mode-based conditions: comment mode requires thread_context or target_url, while share mode requires share_target and optionally share_link. It also states the draft awaits human approval, clarifying when the tool is appropriate. It does not explicitly name sibling alternatives like request_publish for posting, so a small inference is left to the agent.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
outreach_list_queueARead-only
List social outreach drafts. Defaults to status='drafted' (pending review). Pass status='approved' to see what's queued to post next, or status='posted' for recent history. Returns up to limit rows (default 10).
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | Max rows to return (1-50) | |
| status | No | One of: drafted, approved, posted, rejected | drafted |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The readOnlyHint annotation already signals a safe read operation, and the description reinforces this by saying it lists drafts. It adds useful behavioral context: default status is 'drafted', status values mean different pipeline stages, and results are limited by `limit`. This exceeds the minimal safety signal from annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, each earning its place: the core action, the status semantics, and the limit behavior. The most important information is front-loaded, and there is no filler or repetition of schema content.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple read-only list tool with two well-documented parameters, the description covers the essential invocation details: defaults, status meanings, and row limits. It does not mention ordering, but this is a minor omission given the tool's simplicity and the absence of an output schema to clarify further.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema already covers both parameters thoroughly (100% coverage), so the baseline is 3. The description adds meaningful semantics beyond the schema by explaining what each status represents in the workflow (pending review, queued to post, recent history) and clarifying that `limit` caps returned rows. This adds value without being redundant.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific action ('List') and resource ('social outreach drafts') and clarifies the meaning of each status filter. It clearly distinguishes this tool from mutation-oriented siblings like outreach_draft_post, outreach_reject_draft, and request_publish by framing it as a read-only listing operation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides concrete usage context per status: 'drafted' means pending review, 'approved' is queued to post, 'posted' is history. It stops short of explicitly naming when to prefer this over sibling tools like outreach_status, but the status-based usage guidance is strong enough for an agent to select and invoke it correctly.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
outreach_reject_draftADestructive
Reject an outreach draft by id. Marks the row 'rejected' so it won't post. Use when the user says 'kill that one', 'don't post draft 42', etc.
| Name | Required | Description | Default |
|---|---|---|---|
| id | Yes | SocialOutreachLog row id | |
| idempotency_key | No | Optional. Send the same key when retrying this call: the retry waits for or returns the first run instead of starting a second one. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already indicate destructiveHint=true and readOnlyHint=false, and the description adds useful context beyond that: the row is marked 'rejected' and will not post. This clarifies the actual state change without contradicting the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences deliver the action, the effect, and the user-intent triggers without any filler. The core behavior is front-loaded, and the examples are compact yet helpful.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple mutation tool with annotations and full schema coverage, the description is complete: it names the operation, the target, the resulting state, and the language users might use to request it. Nothing essential is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents both id and idempotency_key. The description only reinforces that the id identifies the draft; it adds no new parameter-level meaning.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Reject'), a specific resource ('an outreach draft'), and the precise effect ('Marks the row 'rejected' so it won't post'). This clearly distinguishes it from posting tools like outreach_draft_post and from read-only outreach tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly gives when to use the tool with natural-language examples like 'kill that one' and 'don't post draft 42'. It does not explicitly mention alternatives or when not to use it, but the usage trigger is concrete and unambiguous.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
outreach_statusARead-only
Get the current state of the social outreach loop (enabled / supervised / cadence per platform). Use when the user asks 'is outreach on?', 'how many posts today?', or 'what's the outreach status?'.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, and the description reinforces the read-only nature by saying 'Get the current state'. It adds useful context about what the state includes, which goes beyond the annotation. No contradictory or hidden side-effect behavior is suggested.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences with no wasted words. The first sentence states the operation and result; the second quickly gives concrete user-phrase triggers. This is efficient and well front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a parameterless, read-only status tool, the description is complete: it names the domain, the state components returned, and the conversational triggers. The sibling-tool context is handled by the explicit 'outreach' scope. No output schema exists, but the described categories are enough to interpret the result.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema has no parameters, so there is nothing for the description to clarify. The baseline for a zero-parameter tool is 4, and the description appropriately focuses on the tool's outcome rather than parameter details.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Get'), a clear resource ('the social outreach loop'), and the exact data points returned ('enabled / supervised / cadence per platform'). This makes it easy to distinguish from sibling outreach tools like outreach_draft_post or outreach_list_queue, which involve different actions.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides explicit trigger examples ('is outreach on?', 'how many posts today?', 'what's the outreach status?'), giving clear guidance on when to call this tool. It does not enumerate exclusions or alternatives, but for a parameterless status tool, the positive use cases are sufficient.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
process_fileBRead-only
Process and extract content from files (PDF, DOCX, CSV, images, Excel, etc.)
| Name | Required | Description | Default |
|---|---|---|---|
| file_path | Yes | Path to file to process |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The readOnlyHint annotation already signals the tool is non-mutating, and the description's 'extract content' wording is consistent with that. The description adds file-format context, but it does not disclose output format, behavior on unsupported files, size limits, or any other operational traits beyond what the annotation and basic action imply.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single concise sentence with no filler. It front-loads the action and resource, then enumerates supported formats. Every word earns its place, and 'etc.' adequately signals the format list is non-exhaustive.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description should explain what the tool returns or yields after processing, but it only says 'extract content,' leaving the result format ambiguous. There is also no mention of error cases or how to handle unsupported file types, so an agent is left guessing about the call's outcome.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% for file_path, so the schema already documents the only parameter. The description adds meaning by listing supported file types, which gives the agent a sense of acceptable input values, but it does not elaborate on path syntax, relative vs. absolute paths, or any additional constraints.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a clear action ('extract content') on a resource ('files') and lists common formats (PDF, DOCX, CSV, images, Excel), which helps identify what the tool handles. However, the verb 'process' is generic, and it does not explicitly distinguish this tool from sibling file-readers like read_code or analyze_code, so it is not fully differentiated.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives no guidance on when to use this tool vs. alternatives. Sibling tools such as read_code, analyze_code, and read_logs overlap in the general 'read/process a file' space, and no conditions or exclusions are provided to help an agent choose correctly.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
read_ast_nodeARead-only
Read the exact source code of a specific class or function from a Python file in a Code Repository folder. This is more precise and token-efficient than reading the entire file. Only supports Python (.py) files currently and requires a repository-relative filepath.
| Name | Required | Description | Default |
|---|---|---|---|
| filepath | Yes | Path to the Python file, relative to the Code Repository folder. | |
| folder_id | Yes | The integer ID of the Code Repository folder. | |
| node_name | Yes | The name of the class or function to extract (e.g., 'MyClass' or 'my_function'). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, and the description adds useful behavioral context: it extracts exact source verbatim, currently supports only Python files, and requires a repository-relative path. It could mention error behavior or return format, but given the read-only annotation and simple operation, the added detail is sufficient.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two tight sentences with no wasted words. The first sentence identifies the action and object; the second adds key differentiators and constraints. Every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple read-only tool with all required parameters documented in the schema, the description supplies the purpose, constraints, and differentiators. 'Read the exact source code' implies the return value is the source snippet itself. The only minor gap is the absence of explicit error-case behavior, so it stops short of a perfect 5.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, with inline descriptions for filepath, folder_id, and node_name. The tool description reinforces the roles of filepath and node_name but does not add new parameter-level meaning beyond what the schema already provides. Baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a clear verb ('Read'), a precise resource ('exact source code of a specific class or function from a Python file'), and a clear scope ('Code Repository folder'). The token-efficiency contrast differentiates it from reading whole files, which is especially relevant given sibling read tools like read_code and process_file.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explicitly says when to use this tool: when you need precise, token-efficient extraction of a specific class or function rather than the entire file. It also provides constraints—Python-only and repository-relative filepath. It does not name specific sibling tools, but the guidance is clear enough for an agent to decide appropriately.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
read_codeARead-only
Read the complete contents of a source code file. Returns file content with line count and character count. Use this to understand existing code before making modifications. Accepts paths relative to the project root and explicit absolute paths for user-referenced external files.
| Name | Required | Description | Default |
|---|---|---|---|
| filepath | Yes | Path relative to project root, or an explicit absolute path for an external text file |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The readOnlyHint annotation already marks the tool as read-only. The description adds value by specifying the return includes line and character counts, and by clarifying path semantics (relative vs absolute for external files). No side effects or contradictions are present.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three dense, front-loaded sentences: action, return value, usage context, and path rules. Every sentence earns its place with no redundancy or filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a one-parameter, read-only tool with no output schema, the description provides purpose, usage context, return information (line count, character count), and path handling. An agent has enough to select and invoke the tool correctly without additional lookups.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3. The description largely repeats the path rule already in the schema ('relative to project root' vs 'absolute path for external text file'). It adds no meaningful new parameter semantics beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description starts with a specific verb and resource: 'Read the complete contents of a source code file.' It clearly differentiates from siblings such as read_logs (logs), search_code (search), and list_code_files (list). Including the return value (line count, character count) further sharpens the tool's identity.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
'Use this to understand existing code before making modifications' provides a clear, actionable when-to-use context. The description does not explicitly mention alternatives or when-not-to-use, but for a simple read tool, this is sufficient guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
read_document_sectionARead-only
Read the actual indexed text of one section or page of a document, without searching. Use after get_document_outline when you know which part you need.
| Name | Required | Description | Default |
|---|---|---|---|
| page_label | No | Page number, for paginated documents such as PDFs. | |
| heading_path | No | Section breadcrumb from get_document_outline (exact or partial match). | |
| source_filename | Yes | The document filename. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The readOnlyHint annotation already covers safety, so the description's job is to add context. It adds that the tool reads 'actual indexed text' and mentions the usage flow with get_document_outline. It does not describe behavior when no page_label or heading_path is provided, or error cases, but given the annotation covers read-only, a 3 is appropriate.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two concise sentences with no redundancy. The main action is front-loaded, and the usage context is given in a single subordinate clause. Every word earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description is complete for a simple read tool: it states what it reads, when to use it, and its relationship to get_document_outline. It lacks clarification on behavior when neither page_label nor heading_path is supplied, which could lead to ambiguity, but the core use case is well-covered.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so all three parameters are already described in the input schema. The description repeats the heading_path source (from get_document_outline) but does not add new meaning beyond that. The baseline for full coverage is 3, and the description adds minimal extra value.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool reads indexed text of a section or page, specifying the verb (read) and resource (section/page of a document). It also differentiates from searching by explicitly saying 'without searching' and mentions the prerequisite tool get_document_outline, making it distinguishable from siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It provides explicit guidance on when to use: after get_document_outline when you know the part you need. It also excludes searching, which helps agents avoid search tools. However, it does not name specific search alternatives (like search_knowledge_base or search_codebase) explicitly, but the 'without searching' exclusion is clear enough.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
read_logsARead-only
Tail a Guaardvark log file under logs/. Use when the user says 'review the logs', 'check backend.log', 'celery errors', or 'what did the last crash say'.
| Name | Required | Description | Default |
|---|---|---|---|
| name | No | Log filename (not a path). Default backend.log. | backend.log |
| lines | No | How many trailing lines (10-400). | |
| query | No | Optional case-insensitive substring filter. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The annotation readOnlyHint=true already declares the operation is safe and non-mutating, and the description adds genuine value by disclosing the 'tail' semantics (reads the trailing portion) and the logs/ location. However, it does not disclose the return format or pagination/truncation behavior, leaving some behavioral detail unspecified.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences with zero waste. The core action is front-loaded in the first sentence, and the second adds high-value usage triggers. Every word earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only tail tool with 0 required params, full schema coverage, and a readOnlyHint annotation, the description covers purpose, location, and when to use it. The only gap is return format, which is minor for a simple log-tail operation and not covered by an output schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% — all three parameters (name, lines, query) are documented in the schema with defaults and constraints. The description adds no parameter-level detail beyond the schema, so the baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Tail') plus a concrete resource and scope ('a Guaardvark log file under logs/'). The trigger-phrase examples ('review the logs', 'check backend.log') sharpen what this tool is for and clearly distinguish it from read-code/document siblings that operate on source files rather than logs.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit when-to-use guidance via natural-language trigger phrases ('review the logs', 'celery errors', 'what did the last crash say'). It does not name an alternative tool or state when not to use it, but among the siblings there is no competing log reader, so the omission is minor.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
remove_backgroundA
Remove the background from an attached photo and return a transparent PNG. Use for product shots, stickers, and cut-outs. Does not invent a new scene — use generate_identity or edit_image for that. The attached image is used automatically if image is omitted.
| Name | Required | Description | Default |
|---|---|---|---|
| image | No | Path of the photo. Usually omit — the attached image is used. | |
| idempotency_key | No | Optional. Send the same key when retrying this call: the retry waits for or returns the first run instead of starting a second one. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description adds behavioral details beyond the annotations: the result is a transparent PNG, the tool does not invent a new scene, and the attached image is used automatically when image is omitted. It could further disclose whether the original file is modified or whether processing is synchronous, but the core behavior is clear.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences front-load the action and output format. The first two sentences are essential; the third repeats schema content but is brief and reinforces the common calling pattern, so the redundancy is mild.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a two-optional-parameter tool with no output schema, the description covers what the tool does, the output format, the intended use cases, the alternative tools, and the default input behavior. No critical operational detail an agent needs to invoke it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents both image and idempotency_key. The description restates the auto-use of the attached image, which is already present in the schema's image property ('Usually omit — the attached image is used'), so it adds little beyond the structured definition.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific verb and resource: 'Remove the background from an attached photo and return a transparent PNG.' It also distinguishes the tool from sibling image-manipulation tools by explicitly disclaiming scene invention and naming generate_identity/edit_image as the alternatives for that need.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'Use for product shots, stickers, and cut-outs' states the intended application, and 'Does not invent a new scene — use generate_identity or edit_image for that' gives an explicit when-not-to-use rule with named alternatives. This routes an agent correctly without needing to inspect siblings.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
request_publishA
Ask to publish a text post to one of the user's social connections (Discord webhook, Bluesky, Mastodon, Telegram). This does NOT post: the request waits on the Approvals page until a person approves, rejects or cancels it, whatever the publish settings say. Omit connection when only one is set up; otherwise pass its id or name.
| Name | Required | Description | Default |
|---|---|---|---|
| body | Yes | The text of the post. | |
| title | No | Title, for platforms that show one. | |
| link_url | No | A link to attach to the post. | |
| connection | No | Connection id, display name, provider or handle. Optional when only one exists. | |
| idempotency_key | No | Optional. Send the same key when retrying this call: the retry waits for or returns the first run instead of starting a second one. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description discloses the most important behavioral trait beyond annotations: 'This does NOT post: the request waits on the Approvals page until a person approves, rejects or cancels it, whatever the publish settings say.' This async, approval-gated behavior is essential context that annotations (readOnlyHint=false, destructiveHint=false) do not convey. No contradiction with annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tightly written sentences with zero filler. The most decision-relevant facts are front-loaded: the non-posting caveat appears immediately after the purpose, before any parameter guidance.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Complete for a request tool: purpose, approval behavior, and parameter usage are all covered. With no output schema, the rule exempts return-value explanation. A minor gap is that it doesn't mention how to check the outcome of an approval, but the Approvals-page reference suffices.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the schema documents all five parameters. The description adds real value on the connection parameter by explaining when to omit it and what values to pass (id or name), going slightly beyond the schema's 'Optional when only one exists.'
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb-resource pair ('Ask to publish a text post to one of the user's social connections') and enumerates the exact target platforms (Discord webhook, Bluesky, Mastodon, Telegram). The critical qualifier that it 'does NOT post' immediately separates it from a direct-publish tool, and it stands apart from the content-generation siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explains the approval workflow clearly and gives concrete parameter guidance ('Omit connection when only one is set up; otherwise pass its id or name'). It does not name alternative tools explicitly, but for a publish-request tool the key usage context (waits for human approval) is fully covered.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
save_memoryA
Save a fact, user preference, or instruction to long-term memory. Use this to remember things the user tells you about themselves, their projects, or how they want you to behave.
| Name | Required | Description | Default |
|---|---|---|---|
| tags | No | List of string tags for categorization (e.g. ['python', 'formatting']). | |
| type | No | Type of memory: 'fact', 'preference', or 'note'. Legacy 'instruction' is normalized to 'note'. | fact |
| content | Yes | The fact, preference, or instruction to remember. Be specific and concise. | |
| importance | No | Importance score from 0.0 to 1.0. Higher means it should be retrieved more often. | |
| idempotency_key | No | Optional. Send the same key when retrying this call: the retry waits for or returns the first run instead of starting a second one. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already indicate write (readOnly=false) and non-destructive behavior. The description goes further by stating data is persisted to long-term memory and can influence future behavior, which is useful context beyond the annotation flags.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two concise sentences front-load the core action and follow with concrete usage context. Every sentence earns its place with no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 5-parameter write tool with no output schema and no nested objects, the description provides sufficient purpose and usage guidance. Parameter details are fully covered by the schema, and the annotation set covers safety behavior.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3. The description reinforces what content to store but does not add significant meaning beyond the schema's own parameter descriptions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Save') and resource ('long-term memory') and enumerates what can be saved: facts, user preferences, or instructions. It is clearly distinct from retrieval-focused siblings like search_memory and search_knowledge_base.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explicitly says to use this when the user tells you things about themselves, their projects, or desired behavior, which gives clear when-to-use context. It does not name alternative retrieval tools or exclusions, but the sibling set makes the intended use obvious.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
search_codeARead-only
Search for code patterns across the project using case-insensitive regex. Returns all matches with file paths, line numbers, and matched content. Use this to find where code patterns exist before making changes.
| Name | Required | Description | Default |
|---|---|---|---|
| pattern | Yes | Text or regex pattern to search for (e.g., 'handleClick', 'Button.*onClick') | |
| file_glob | No | Glob pattern for files to search (default: '**/*.{py,jsx,js,tsx,ts}') | **/*.{py,jsx,js,tsx,ts} |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only declare readOnlyHint=true; the description adds useful behavioral details beyond that, such as case-insensitive regex matching and the exact shape of results (file paths, line numbers, matched content). It is consistent with the read-only annotation and provides meaningful context without contradiction.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences convey the operation, regex behavior, return contents, and a guiding use case with no redundant or filler language. The most decision-relevant detail is front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple, read-only two-parameter search tool with no output schema, the description covers key runtime behavior: regex matching, result contents, and search scope. It omits minor details like regex flavor or default file glob, but those are already listed in the schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already fully documents both parameters with examples and defaults. The description reinforces that the pattern is regex-based but does not add meaningful parameter semantics beyond what the schema provides, so a baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool searches for code patterns using case-insensitive regex and returns file paths, line numbers, and matched content. It is specific about the resource (code) and the operation (search), though it does not explicitly distinguish itself from the similarly named sibling search_codebase.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The final sentence gives a clear intended use ('Use this to find where code patterns exist before making changes'). It does not mention when not to use it or explicitly name alternatives, so it falls just short of the strongest guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
search_codebaseARead-only
Search the current project's source code, which is already indexed: no path or upload is needed, just call it. Ask by meaning or by symbol: 'where is it decided whether a model supports thinking', 'which function cuts retrieved text', 'callers of think_payload'. Returns the matching files with line numbers and the code itself. Use it before saying the code is unavailable, before reading whole files, and instead of shell commands like grep or ls.
| Name | Required | Description | Default |
|---|---|---|---|
| root | No | Checkout to search; leave empty for the current project | |
| limit | No | How many hits to return (default 8) | |
| query | Yes | What to find, in plain words or as a symbol name |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only provide readOnlyHint: true. The description adds valuable behavioral context beyond that: the codebase is already indexed, no path or upload is needed, it supports meaning- or symbol-based queries, and it returns files with line numbers and code. This meaningfully expands on the safety trait.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four sentences with no wasted words: purpose and key trait first, then examples, then return format, then usage guidance. Every sentence earns its place and the most important information is front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, but the description covers the return shape (files with line numbers and code), the indexing/setup-free behavior, and typical usage patterns. It is nearly complete, with the only gap being explicit disambiguation from siblings like 'search_code or read_code.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3. The description adds 'Ask by meaning or by symbol' but this essentially restates the query parameter's schema description ('What to find, in plain words or as a symbol name'). It doesn't add new parameter-level meaning, so a 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: 'Search the current project's source code' and gives concrete examples of query types, such as 'where it is decided whether a model supports thinking' and 'callers of think_payload'. The scope is clear, but it doesn't explicitly differentiate itself from the similarly named sibling 'search_code', so it falls just short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives explicit usage context: 'Use it before saying the code is unavailable, before reading whole files, and instead of shell commands like grep or ls.' This clearly tells an agent when to prefer this tool, though it doesn't name sibling alternatives or explicitly state when not to use it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
search_knowledge_baseARead-only
Search the internal knowledge base for information about the project, architecture, code repositories, or documents. Returns verbatim source passages with their filenames and relevance scores, not a summary — cite the filenames in your answer.
| Name | Required | Description | Default |
|---|---|---|---|
| query | Yes | The specific question or query to search for in the knowledge base. | |
| top_k | No | How many passages to return (1-50). Omit for the configured default. | |
| project_id | No | Optional project ID to scope the search. | |
| filter_type | No | Optional filter on the indexed content type. Real values include 'document', 'text', and 'repository_summary'. Omit to search everything. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The readOnlyHint annotation already communicates that this is a safe read operation. The description adds valuable behavioral detail beyond that: results are verbatim source passages with filenames and relevance scores, not a synthesized summary, and the agent should cite filenames in answers.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences with no filler. It front-loads the action and scope, then immediately states the output format and citation instruction. Every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only search tool with a fully documented schema and no output schema, the description adequately explains what the agent will receive back (passages, filenames, relevance scores). It could be marginally more complete by routing between sibling search tools, but the core call and return behavior are sufficiently covered.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema has 100% description coverage, documenting query, top_k, project_id, and filter_type with meaningful detail including ranges and real filter values. The description does not add parameter-specific semantics beyond what the schema already provides, so the baseline score of 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Search'), a clear resource ('internal knowledge base'), and the kinds of information it covers ('project, architecture, code repositories, or documents'). It also distinguishes itself from sibling code/memory search tools by specifying verbatim passages with filenames and relevance scores, and explicitly notes it is not a summary.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear context about when to use it: when searching the internal knowledge base for project, architecture, repository, or document information. It does not explicitly name sibling alternatives or state when not to use this tool, but the intended scope is reasonably clear.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
search_memoryARead-only
Search your long-term memory for previously saved facts, preferences, or instructions.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | Maximum number of results to return (default 5). | |
| query | Yes | Search query or keyword to look for in memories. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, and the description aligns with that by framing the operation as a search. The description adds the semantic scope ('long-term memory', 'previously saved facts, preferences, or instructions') but does not disclose behavior beyond that, such as result ordering or relevance mechanics.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single, front-loaded sentence that states the action, the resource, and the content scope without any filler. Every word contributes to the agent's understanding.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple, read-only, two-parameter search tool, the description is sufficient: the agent knows what to search, what type of content will be returned, and that the operation is safe. Return format is not described, but for a memory retrieval tool this is reasonably inferable and annotations cover the safety profile.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so both 'query' and 'limit' are already documented in the schema. The description adds no additional parameter detail beyond what the schema provides, so the baseline of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific verb ('Search'), a distinct resource ('your long-term memory'), and the content types searched ('facts, preferences, or instructions'), which clearly differentiates it from codebase and web search tools. It does not explicitly contrast it with a sibling like search_knowledge_base, so it stops short of full differentiation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'previously saved' implies this tool is for retrieving stored memory rather than live or external data, giving some usage context. However, there is no explicit when-to-use guidance, exclusion of alternatives, or note about when search_knowledge_base or web_search would be more appropriate.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
self_improvement_statusARead-only
Report whether self-improvement can run (codebase lock, enabled flag, already running) plus recent runs and PendingFix rows. Use when the user asks if SI is on, why a fix didn't apply, or what pending fixes exist.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, and the description reinforces a read-only reporting behavior. It adds useful context about the kind of information returned, such as lock state, enabled flag, recent runs, and pending fix rows, which goes beyond the annotation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two concise sentences, with the core reporting purpose and key information types front-loaded, followed by direct usage triggers. No filler or redundant content.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter, read-only status tool, the description fully covers what it reports, why it exists, and when to invoke it. The lack of an output schema is not a gap here because the description enumerates the major content areas.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters, so the schema provides no parameter semantics to explain. The description compensates by clarifying what the status report contains, making the tool's meaning clear without needing parameter documentation.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Description clearly states the tool reports self-improvement status, including specific checks like codebase lock, enabled flag, already running, recent runs, and PendingFix rows. This is specific enough to distinguish from sibling status tools like swarm_status or outreach_status.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly provides when to use the tool: when the user asks if SI is on, why a fix didn't apply, or what pending fixes exist. It does not mention alternatives or exclusions, but the use-case framing is clear and actionable.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
start_film_crewA
Start a five-role Film Crew production from a screenplay. The screenwriter begins at once; casting, storyboards and GPU renders wait for you in Studio. Use when the user asks to film a script or start the film crew.
| Name | Required | Description | Default |
|---|---|---|---|
| name | No | Production name. Defaults to the first line of the script. | |
| script_text | Yes | Screenplay or scene list (plain text). | |
| video_model | No | Optional I2V / scene model id. Default: the active video model. | |
| idempotency_key | No | Optional. Send the same key when retrying this call: the retry waits for or returns the first run instead of starting a second one. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description reveals meaningful async behavior beyond the annotations: the screenwriter 'begins at once' while casting, storyboards, and GPU renders are deferred pending user action in Studio. This helps the agent set expectations about side effects and follow-up steps, though it does not detail what the direct API response contains.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences with no filler: the first states the core action and immediate effect, the second gives the exact usage trigger. Every sentence earns its place, and key information is front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 4-parameter tool with no output schema, the description adequately conveys the workflow, asynchronous start, and what the user should expect next. It stops short of describing return values or how to monitor progress, but those are reasonably inferable from the provided workflow context.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents every parameter. The description adds context by framing the tool as screenplay-driven, which maps to script_text, but it does not add information beyond what the parameter descriptions already provide.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific action ('Start'), a specific resource ('five-role Film Crew production'), and the input ('a screenplay'). It clearly distinguishes this workflow-launching tool from video/image generation siblings like generate_video by framing it as a multi-stage production.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly says when to use it: 'Use when the user asks to film a script or start the film crew.' It gives a clear trigger condition but does not mention alternatives or when not to use it rather than a more direct generation tool, which would have made the guidance stronger.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
summarize_corpusARead-only
Get high-level summaries of what the whole knowledge base covers, rather than individual passages. Use for broad questions — overall themes, what a collection is about, how topics relate — which passage search answers poorly.
| Name | Required | Description | Default |
|---|---|---|---|
| level | No | Summary level: 1 is closer to the source, 2+ is broader. Default 1. | |
| limit | No | How many summaries to return (default 8, max 30). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The annotations already declare readOnlyHint=true, so the safety profile is covered. The description adds context about corpus-level aggregation versus passage-level retrieval, but it does not disclose additional behaviors like return shape, whether summaries are precomputed, or how level interacts with output.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two tight sentences with no filler. The core purpose is front-loaded, followed by concrete usage examples and a contrast with the alternative, so every clause earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple read-only tool with two optional parameters fully documented in the schema, the description is complete. It explains what the tool does, when to use it, and how it differs from passage search, and the readOnlyHint annotation covers the mutation concern.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3. The description does not add parameter-level detail beyond the schema, but the schema already explains level, limit, defaults, and max, so no compensation is needed.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states a specific action ('Get high-level summaries') and a distinct resource scope ('the whole knowledge base'). It also explicitly contrasts itself with 'individual passages,' which distinguishes it from passage-level search tools among the siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives clear when-to-use guidance with concrete examples ('broad questions, overall themes, what a collection is about, how topics relate'). It also implies when not to use it ('rather than individual passages') and points to passage search as the weaker alternative, though it does not name a specific sibling tool explicitly.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
swarm_statusARead-onlyIdempotent
Get Swarm Orchestrator status (same as GET /api/swarm/status). Use when the user asks about the coding swarm, worktrees, or running swarm tasks. Returns an honest offline error if the swarm plugin is not running.
| Name | Required | Description | Default |
|---|---|---|---|
| swarm_id | No | Optional swarm id for a single swarm. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already mark the tool as readOnly and idempotent. The description adds useful behavioral context by disclosing the 'honest offline error if the swarm plugin is not running,' which goes beyond the annotations and helps agents anticipate failure modes.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three concise sentences with no filler: purpose, usage trigger, and error behavior are each covered efficiently. Information is front-loaded and every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the low complexity, strong annotations, and complete schema, the description is largely sufficient for selecting and invoking the tool. It does not describe the success response payload, but the endpoint mapping and simple status semantics make this a minor gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% for the single optional swarm_id parameter, so the schema already documents its meaning. The description adds no extra parameter detail, but the baseline of 3 is appropriate because no compensation is needed.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('Get') and resource ('Swarm Orchestrator status'), and even maps it to a concrete endpoint. This clearly identifies what the tool does and distinguishes it from sibling tools like self_improvement_status.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly says when to use the tool: 'when the user asks about the coding swarm, worktrees, or running swarm tasks.' It provides clear context but does not mention exclusions or name alternative tools, so it stops just short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
verify_changeARead-only
Verify that a code change was successful by checking if text exists in file. Use after edit_code to confirm changes were applied correctly. Set should_exist=False to verify that text was successfully removed. Accepts paths relative to the project root and explicit absolute paths for user-referenced external files.
| Name | Required | Description | Default |
|---|---|---|---|
| filepath | Yes | Path relative to project root, or an explicit absolute path for an external text file | |
| should_exist | No | True if text should exist, False to verify deletion (default: True) | |
| expected_text | Yes | Text to check for in the file |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, and the description aligns by describing a non-mutating check. It adds value by explaining the ability to verify deletion via should_exist and the path handling for external files, which goes beyond the annotation's bare safety flag.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences with zero filler. Purpose is front-loaded, followed by usage context and a parameter nuance. Every sentence earns its place, making it efficient and easy to scan.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a verification tool with three parameters and no output schema, the description covers the core scenario (post-edit verification), the inverse case (removal), and path semantics. It doesn't specify return format, but that's not essential for correctness and annotations cover safety.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema covers all parameters with descriptions, but the description enriches meaning: it clarifies filepath accepts relative or absolute paths, and explains the should_exist flag's effect. This adds practical context beyond the schema's bare definitions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific action ('Verify that a code change was successful') with a concrete mechanism ('checking if text exists in file'). It clearly distinguishes itself from siblings by anchoring to edit_code and the should_exist flag, making its role unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly says 'Use after edit_code to confirm changes were applied correctly', giving a precise trigger condition. It also covers the inverse case with should_exist=False. It doesn't mention when not to use it, but the context is clear enough for an agent.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
web_searchARead-only
Search the web via DuckDuckGo — returns a ranked list of titles, snippets, and URLs for a query. Use this for open-ended research or when you need to discover pages about a topic. For fetching a SPECIFIC URL or domain the user already named, use fetch_url instead (it's a direct fetch, no search ranking in between).
| Name | Required | Description | Default |
|---|---|---|---|
| query | Yes | Search query string | |
| max_results | No | Maximum number of results to return |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, so the read-only safety profile is covered. The description adds useful behavioral context beyond that: it uses DuckDuckGo as the backend and returns a ranked list with titles, snippets, and URLs. No contradictions with annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, no filler. The primary purpose and output are front-loaded, and the sibling alternative is mentioned concisely in the second sentence. Every phrase earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
This is a simple two-parameter read-only search tool with no output schema. The description covers what the tool does, what it returns, when to use it, and when not to. No critical context is missing for an agent to invoke it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents the query and max_results parameters. The description adds context about how results are ranked and what they contain, but does not substantially elaborate on parameter behavior beyond the schema. Baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly identifies the specific verb (Search), resource (web via DuckDuckGo), and output (ranked list of titles, snippets, URLs). It also explicitly distinguishes itself from the sibling tool fetch_url, which is the direct-fetch alternative.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explicitly states when to use this tool (open-ended research, discovering pages about a topic) and names the alternative for a different use case (fetch_url for a specific URL/domain). An agent can make the right routing decision directly from the description.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
1 tool update
v2.9.3- Changed
generate_video1 field changed- changed
Input schema / properties / style / descriptionPrevious value: -"Prompt style: cinematic, realistic, artistic, anime, 3d_animation, stop_motion, hand_drawn, western_cartoon, none."New value: +"Prompt style: cinematic, realistic, artistic, anime, 3d_animation, stop_motion, hand_drawn, western_cartoon, none. A model may not offer every style; its capability record lists prompt_styles."
51 tool updates
v2.9.0- First observed
analyze_code - First observed
analyze_website - First observed
codegen - First observed
edit_image - First observed
fetch_url - First observed
generate_animation - First observed
generate_bulk_csv - First observed
generate_csv - First observed
generate_enhanced_wordpress_content - First observed
generate_file - First observed
generate_image - First observed
generate_music_video - First observed
generate_video - First observed
generate_wordpress_content - First observed
get_dependency_graph - First observed
get_document_outline - First observed
get_generation_status - First observed
get_repository_map - First observed
inpaint_image - First observed
inspect_gpu - First observed
list_code_files - First observed
list_code_repositories - First observed
list_documents - First observed
map_codebase - First observed
media_control - First observed
media_play - First observed
media_status - First observed
media_volume - First observed
outpaint_image - First observed
outreach_draft_post - First observed
outreach_list_queue - First observed
outreach_reject_draft - First observed
outreach_status - First observed
process_file - First observed
read_ast_node - First observed
read_code - First observed
read_document_section - First observed
read_logs - First observed
remove_background - First observed
request_publish - First observed
save_memory - First observed
search_code - First observed
search_codebase - First observed
search_knowledge_base - First observed
search_memory - First observed
self_improvement_status - First observed
start_film_crew - First observed
summarize_corpus - First observed
swarm_status - First observed
verify_change - First observed
web_search
TDQS
Scored across 51 tools
Many tools have overlapping purposes: search_code vs search_codebase vs read_code vs read_ast_node for code exploration, edit_image vs inpaint_image vs outpaint_image vs remove_background for image editing, and generate_file vs generate_csv vs generate_wordpress_content for content generation. The detailed descriptions help somewhat, but the boundaries remain blurry across multiple clusters.
Nearly all tools use snake_case and follow a predictable verb_noun or noun_verb pattern within their subgroup (e.g., media_play, outreach_status, generate_image). A few outliers like codegen and start_film_crew are minor deviations, but overall naming is consistent and readable.
With 51 tools spanning unrelated domains such as media playback, GPU diagnostics, code mapping, and social outreach, the set is extremely bloated for a single MCP server. This far exceeds the 3–15 sweet spot and creates choice overload for agents.
The surface covers many domains broadly, but there are notable gaps: an in-place code edit tool (edit_code) is referenced in multiple descriptions but absent from the tool list, and there is no job cancellation or delete operation for generated or indexed content. These missing operations create dead ends for common workflows.
Maintenance
Related MCP Connectors
AI image, video, voice and music generation over MCP, routed to Veo 3.1, Seedance 2.5 and more.
Multi-model AI image and video generator. 14 models behind one OAuth-secured MCP endpoint.
AI video editor for agents and humans: timeline, captions, color, audio and generation as MCP tools.
OCR, transcription, file extraction, and image generation for AI agents via MCP.
Related MCP Servers
AlicenseNot gradedqualityDmaintenanceGenerate videos, images, audio, and 3D models from any MCP-compatible AI agent — Claude, Cursor, ChatGPT, and more.MIT- AlicenseAqualityBmaintenanceLocal MCP server that plans, generates, and assembles production assets (images, audio, video) through multi-agent personas and official APIs, with free-tier budget guard.15MIT
- AlicenseNot gradedqualityBmaintenanceA local-first creative studio MCP server that provides access to 73 curated image and video generation models across 5 providers, enabling users to generate, refine, and manage creative outputs with transparent cost tracking and local file custody.5MIT
- AlicenseCqualityAmaintenanceOpen-source local-first AI media workbench with Canvas, video editor, Assets, and MCP control.29103Apache 2.0
Ep 2 — Chat Brain: one chat box, three speeds
Ep 3 — File Desktop: your files get a desktop, plus RAG that shows its work
Ep 4 — Screen Agent: its own desktop, eyes, and hands
Ep 5 — Image Gen: one prompt, a whole story
Ep 6 — Video Gen: seven models, one GPU
Ep 7 — Voice Clone: consent-gated, self-checking
Ep 8 — Music Video: drop a song, get a film
Ep 9 — Film Crew: script, cast, storyboard, cut
Ep 11 — Self-Repair: it fixes its own code, behind a gate you control
Ep 12 — Command Center: see everything, gate everything, kill everything
Ep 13 — The New Front Door: the Workspaces bar, and everything new since the first series
Ep 14 — System Map: every module, drawn from its real imports; findings that carry their own fix