stitch-mcp-stdio
Allows interaction with Google Stitch AI UI design tool, providing tools for managing projects, screens, design systems, and generating UI designs from text prompts.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@stitch-mcp-stdioGenerate a login screen for a mobile app"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
stitch-mcp-stdio
Stable stdio MCP server for Google Stitch AI UI design tool.
Drop-in replacement for @_davideast/stitch-mcp that eliminates the proxy layer crash (libuv UV_HANDLE_CLOSING assertion failure on Windows and intermittent connection drops on other platforms).
Why this exists
The official @_davideast/stitch-mcp package runs a proxy subprocess that suffers from a libuv bug causing random crashes, especially on Windows. This package uses the same @google/stitch-sdk under the hood but connects via stdio transport directly -- no proxy process, no crash.
Related MCP server: Stitch MCP
Setup
1. Get a Stitch API key
Go to stitch.withgoogle.com -> Profile -> Settings -> API Key.
2. Configure your MCP client
Point any MCP-compatible client (Claude Code, Claude Desktop, Cursor, Gemini CLI, VS Code) at npx stitch-mcp-stdio:
{
"mcpServers": {
"stitch": {
"command": "npx",
"args": ["-y", "stitch-mcp-stdio"],
"env": {
"STITCH_API_KEY": "your-api-key"
}
}
}
}npx -y downloads + runs the latest version without a permanent install. To pin a version use stitch-mcp-stdio@1.0.0.
Alternative: local clone (if you'd rather not use npx)
git clone https://github.com/bluedevilcollectibles/stitch-mcp-stdio.git
cd stitch-mcp-stdio
npm installThen point the client at node /path/to/stitch-mcp-stdio/server.js instead of npx.
Available tools
All 12 tools from the Stitch SDK are exposed:
Tool | Description |
| Create a new Stitch project |
| Get project details |
| List all projects |
| List screens in a project |
| Get screen details |
| Generate a screen from a text prompt |
| Edit an existing screen |
| Generate design variants |
| Create a design system |
| Update a design system |
| List design systems |
| Apply a design system to screens |
How it works
MCP Client (Claude, Cursor, etc.)
|
| stdio (stdin/stdout)
|
stitch-mcp-stdio (this package)
|
| HTTPS (API key auth)
|
stitch.googleapis.com/mcpNo proxy subprocess. No libuv handle management. Just stdio in, API calls out.
Built by
Blue Devil Collectibles -- makers of LiveSeller Pro, inventory and live selling tools for comic shops, card shops, and collectibles dealers.
We built this MCP server to power our AI design pipeline. If you sell on WhatNot, manage pull lists, or run a brick-and-mortar shop, check out what we're building.
License
MIT
Available Tools
16 toolsapply_design_systemADestructive
Applies a design system to a list of screens. Use this tool when the user wants to update one or more screens to match the style of a design system. This tool applies the selected design system's foundational design tokens (colors, fonts, shapes, etc.) to the chosen screens, modifying their appearance to align with the design system.
| Name | Required | Description | Default |
|---|---|---|---|
| assetId | Yes | Required. The asset id of the design system to apply, can be fetched from 'list_design_systems'. Example: '15996705518239280238', without the `assets/` prefix. | |
| projectId | Yes | Required. The project ID of screen instances to edit, example: '4044680601076201931', without the `projects/` prefix. | |
| selectedScreenInstances | Yes | Required. The screen instances to edit, which is available in the Project info, fetched by `get_project`. |
Output Schema
| Name | Required | Description |
|---|---|---|
| projectId | No | The project ID of the generated screen. This is the same as the input project ID. |
| sessionId | No | The session ID of the generated screen. |
| outputComponents | No | The generated output components. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already mark this as destructive and non-read-only. The description adds context by explaining the tool modifies screen appearance via design tokens, but it does not disclose potential overwriting or irreversibility of existing screen styles. It adds some behavioral value without contradicting the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences, front-loaded with the core action and usage guidance. Every sentence earns its place with no redundancy or filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool's purpose, usage trigger, and expected effect are clearly stated. The input schema fully documents all parameters and even explains where to find screen instances, and annotations cover safety traits. An agent has enough information to select and invoke the tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all three parameters are already documented in the schema. The description adds minimal parameter-level meaning beyond naming colors, fonts, and shapes as applied design tokens, which does not need to compensate for any schema gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'Applies a design system to a list of screens.' It clearly describes the effect (applying foundational design tokens to modify appearance), which distinguishes it from sibling tools like create_design_system or update_design_system. The purpose is unmistakable.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly states when to use the tool: 'Use this tool when the user wants to update one or more screens to match the style of a design system.' It provides clear context, though it does not explicitly name alternatives or state when not to use it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
create_design_systemADestructive
Creates a new design system for a project. Use this tool when the user wants to set or update the overall visual theme, style, or branding of the application. This includes configuring:
Color Palette: Presets, custom primary colors, and saturation levels.
Typography: Font families (e.g., Inter, Roboto, etc.).
Shape: Corner roundness for UI elements.
Appearance: Light and dark mode background colors.
Design MD: Free-form design instructions in markdown. This tool establishes the foundational design tokens that apply across all screens in the project.
Instructions for Tool Call:
Call
update_design_systemtool immediately after this tool to apply the design system to the project, and display the design system in the UI.
| Name | Required | Description | Default |
|---|---|---|---|
| projectId | No | Optional. The project ID to create design system for, example: '4044680601076201931', without the `projects/` prefix. If empty, creates a global asset (not associated with any project). | |
| designSystem | Yes | Required. The design system to create. |
Output Schema
| Name | Required | Description |
|---|---|---|
| name | No | Identifier. The resource name of the asset. Format: assets/{asset} |
| version | No | Output only. The version of this asset. 0 indicates unversioned (legacy data). Incremented when the asset content changes. |
| copiedFrom | No | Optional. The resource name of the asset this was copied from, if any. Format: assets/{asset} Tracks the fork history of assets. |
| designSystem | No | Optional. The design system. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already indicate destructiveHint=true and readOnlyHint=false. The description adds valuable context: the tool establishes foundational design tokens that apply across all screens, and it requires a follow-up call to update_design_system to actually apply the system. This goes beyond the annotations without contradicting them.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is well-structured with a clear purpose statement, a bulleted list of configuration areas, and a bolded instruction section. It is slightly verbose but every part earns its place, and the most important operational detail (call update_design_system after) is highlighted.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity (nested DesignSystem object, many theme properties) and the presence of an output schema, the description covers the essential user-facing behavior and the critical follow-up workflow. It does not need to explain return values, and parameter details are in the schema. Missing only a note about global vs. project-level creation, which is present in the schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3. The description provides a high-level grouping of parameters (Color Palette, Typography, Shape, Appearance, Design MD) but does not add syntax, formats, or field-specific details beyond what the schema already documents. It's helpful for conceptual understanding but not necessary for correct invocation.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states 'Creates a new design system for a project' with a specific verb and resource. It lists the types of configuration (color, typography, shape, etc.), but does not explicitly distinguish from sibling tools like update_design_system or create_design_system_from_design_md, and the phrase 'set or update' could imply an update capability.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly says when to use it: 'when the user wants to set or update the overall visual theme, style, or branding of the application.' It also provides a crucial sequencing guideline: call update_design_system immediately after. However, it does not explicitly state when not to use it or name alternatives like create_design_system_from_design_md.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
create_design_system_from_design_mdADestructive
Creates a design system for a project, with user uploaded DESIGN.md file, and displays the design system in the UI.
Instructions for Tool Call:
Should call
upload_design_mdtool first to upload DESIGN.md to a Stitch project.
| Name | Required | Description | Default |
|---|---|---|---|
| projectId | Yes | Required. The project ID to update design system for, example: '4044680601076201931', without the `projects/` prefix. | |
| deviceType | No | Optional. The device type of the screen. | |
| selectedScreenInstance | Yes | Required. The screen instance to create the design system from, which is created by UploadDesignMd RPC. |
Output Schema
| Name | Required | Description |
|---|---|---|
| assetId | No | The asset id of the design system. Example: '15996705518239280238', without the `assets/` prefix. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already indicate destructiveHint=true, so the mutation risk is known. The description adds that it 'displays the design system in the UI', a useful behavioral detail, but doesn't disclose what specifically might be destroyed or other side effects. This adds some value beyond annotations without contradicting them, warranting a 3.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two short sentences, with the first clearly stating purpose and the second providing a critical instruction. Every word earns its place, and it is appropriately front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool has an output schema, so return values are covered elsewhere. The description gives the essential workflow instruction (call upload_design_md first) and states the UI display side effect. It is complete for a create-like tool with moderate complexity, though it could mention error scenarios or what happens on conflict.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so parameters are already well-documented. The description mentions 'user uploaded DESIGN.md file' and links selectedScreenInstance to UploadDesignMd RPC, but this repeats schema information. With high coverage, the baseline is 3, and the description adds marginal meaning.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool 'Creates a design system for a project, with user uploaded DESIGN.md file, and displays the design system in the UI.' This specific verb+resource combination ('creates a design system') and the qualifier 'from DESIGN.md' distinguishes it from siblings like create_design_system. The instruction to call upload_design_md first further clarifies the workflow.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly instructs to call upload_design_md first, providing clear sequential context. However, it does not mention alternatives or when not to use this tool, so it lacks explicit exclusions. This fits 'clear context, no exclusions' at level 4.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
create_projectADestructive
Creates a new Stitch project. A project is a container for UI designs and frontend code.
| Name | Required | Description | Default |
|---|---|---|---|
| title | No | Optional. The title of the project. |
Output Schema
| Name | Required | Description |
|---|---|---|
| name | No | Identifier. The resource name of the project. Format: projects/{project} |
| title | No | Optional. The title of the project. |
| origin | No | Output only. The origin of the project. |
| metadata | No | Metadata of the project. |
| readTime | No | Output only. The time the project was last read. Populated only when listing recently viewed projects. |
| createTime | No | Output only. The time when the project was created. |
| deviceType | No | Optional. The device type of the project. |
| updateTime | No | Output only. The time when the project was last updated. |
| visibility | No | Optional. The visibility setting of the project. |
| designTheme | No | Output only. The theme used to generate the first design in the project. |
| projectType | No | Optional. The type of the project. If not specified, the project is a text to UI project. |
| backgroundTheme | No | Optional. The background theme of the project. |
| screenInstances | No | Output only. The screen instances on this project. |
| thumbnailScreenshot | No | Optional. The screenshot to be used as the thumbnail for the project. Same as normal design screenshots, this contains the FIFE serving_base_url which requires additional FIFE URL options to be set for sizing. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true, so the agent knows this is a mutating operation. The description adds the context that a project is a container for UI designs and frontend code, which helps set expectations. It does not disclose side effects, permissions, or what happens to existing data, but the annotation covers the destructive nature.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences with no wasted words. The core action is front-loaded, and the clarifying definition of a project is concise and useful.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple creation tool with one optional parameter and an output schema, the description is nearly complete. It explains what a project is, which is the main contextual gap. It does not mention return values, but the output schema exists, so that is not required.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents the 'title' parameter as optional. The description does not add any additional meaning about the parameter beyond what the schema provides, so the baseline of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Creates') and resource ('a new Stitch project'), and adds a clarifying definition of what a project is ('a container for UI designs and frontend code'). It is clear and distinguishes the tool from siblings like create_design_system, though it does not explicitly name any sibling.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies the tool is for creating a project, and the definition of a project gives some context for when to use it. However, it does not explicitly state when to use this tool versus alternatives like create_design_system or apply_design_system, nor does it mention any prerequisites or exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
delete_projectADestructive
Deletes a specific Stitch project using its project name.
Instructions for Tool Call:
This action cannot be undone. Please confirm with "yes" or "no" to proceed.
| Name | Required | Description | Default |
|---|---|---|---|
| name | Yes | Required. Identifier. The resource name of the project to delete. Format: `projects/{project}` Example: `projects/4044680601076201931` |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already provide destructiveHint=true, and the description adds critical context: 'This action cannot be undone' and requires confirmation before proceeding. This goes beyond the binary destructive flag, clarifying the irreversibility and the need for explicit consent, which is valuable for an agent.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is brief and front-loaded with the core purpose, followed by a clearly separated instruction block. Every word earns its place; there is no filler or repetition. The structure is clean and easy to parse.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple, single-parameter delete operation, the description covers the essential action and irreversibility, and the output schema supplies return-value details. It could mention side effects on associated resources, but given the low complexity and strong annotations/schema, the description is sufficiently complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema has 100% description coverage for the only parameter, including format and example. The description itself adds no additional parameter semantics, so the baseline of 3 applies per calibration. The slight mismatch between 'project name' and 'resource name' is a minor redundancy but does not reduce reliability.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action ('Deletes') and the target resource ('a specific Stitch project'), distinguishing it from siblings like get_project or create_project. The use of 'using its project name' aligns with the required parameter, though the schema more precisely calls it a resource name; this is a minor terminology nuance that does not obscure the purpose.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives no explicit guidance on when to use this tool versus alternatives (e.g., update_project, list_projects), nor does it mention prerequisites or scenarios where deletion should be avoided. The confirmation instruction is a procedural call-time guideline, not usage-context guidance, so it does not address this dimension.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
download_assetsC
Download screens and assets to a local directory
| Name | Required | Description | Default |
|---|---|---|---|
| outputDir | Yes | Output directory | |
| projectId | Yes | Project ID |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden of behavioral disclosure. It only states the basic action and does not disclose whether the output directory must already exist, whether existing files are overwritten, what file formats are produced, or whether authentication or special permissions are required.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single concise sentence with no filler words and the primary verb front-loaded. It earns its place, though it is minimal and does not add any structured context.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple two-parameter tool with full schema coverage, the definition is minimally viable. However, it lacks usage guidance, behavioral details, and clarification of what 'assets' includes, and there is no output schema to compensate for these omissions.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, with both projectId and outputDir documented in the schema, so the baseline is 3. The description adds minimal semantic value beyond the schema, only implying that outputDir is a local directory.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a clear verb ('Download'), a resource ('screens and assets'), and a destination ('to a local directory'), which goes beyond just repeating the tool name. It is distinguishable from sibling tools like list_screens and get_screen, though it does not state the exact scope of what is downloaded.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives no guidance on when to use download_assets versus alternatives such as get_screen, list_screens, or generate_screen_from_text. It does not mention any exclusions, prerequisites, or conditions that would help an agent choose this tool.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
edit_screensADestructive
Edits existing screens within a project using a text prompt.
Instructions for Tool Call:
This action can take a few minutes to complete. Please be patient. DO NOT RETRY.
If the tool call fails due to connection error, the process may still succeed.
| Name | Required | Description | Default |
|---|---|---|---|
| prompt | Yes | Required. The input text to generate the screen from. | |
| modelId | No | Optional. The model to use for generation. | |
| projectId | Yes | Required. The project ID of screens to edit, example: '4044680601076201931', without the `projects/` prefix. | |
| deviceType | No | Optional. The type of device that captured the screenshot, e.g., mobile or desktop. | |
| selectedScreenIds | Yes | Required. The screen IDs of screens to edit, example: ['98b50e2ddc9943efb387052637738f61', '98b50e2ddc9943efb387052637738f62'], without the `screens/` prefix. |
Output Schema
| Name | Required | Description |
|---|---|---|
| projectId | No | The project ID of the generated screen. This is the same as the input project ID. |
| sessionId | No | The session ID of the generated screen. |
| outputComponents | No | The generated output components. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description adds valuable behavioral context beyond annotations: it warns that the action can take minutes, instructs the agent not to retry, and notes that a connection error may still lead to success. This is exactly the kind of runtime behavior an agent needs to avoid harmful retries, and it does not contradict destructiveHint=true.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with a one-sentence purpose statement, then a compact bulleted list of crucial operational instructions. Every sentence earns its place and there is no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 5-parameter tool with a full schema and an output schema, the description plus annotations cover the safety profile and runtime behavior. The long-running/retry warning is the missing piece an agent cannot infer from schema, and it is included.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the input schema already documents all five parameters. The description adds no extra meaning about parameters beyond the phrase 'using a text prompt', making the baseline 3 appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('edits') and a precise resource ('existing screens within a project'), which clearly separates it from sibling tools like generate_screen_from_text. An agent can tell what resource is affected without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The general purpose 'edits existing screens' implies the intended use case, but there is no explicit guidance about when to choose this over generate_screen_from_text, generate_variants, or other siblings. It also gives no prerequisites or exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
generate_screen_from_textADestructive
Generates a new screen within a project from a text prompt.
Instructions for Tool Call:
This action can take a few minutes to complete. Please be patient. DO NOT RETRY.
If the tool fails with a timeout, don't retry. Instead, try to get the screen with
get_screenmethod every 30 seconds for up to 10 times before giving up.If the tool call fails due to connection error, the generation process may still succeed. Please try to get the screen with
get_screenmethod later.
Output:
output_components: Ifoutput_componentscontains text, return it to the user. Ifoutput_componentscontains suggestions (e.g. "Yes, make them all"), present these suggestions to the user. If the user accepts one of the suggestions, callgenerate_screen_from_textagain withpromptset to the accepted suggestion.
| Name | Required | Description | Default |
|---|---|---|---|
| prompt | Yes | Required. The input text to generate the screen from. | |
| modelId | No | Optional. The model to use for generation. | |
| projectId | Yes | Required. The project ID to generate the screen for, example: '4044680601076201931', without the `projects/` prefix. | |
| deviceType | No | The type of device that captured the screenshot, e.g., mobile or desktop. | |
| designSystem | No | Optional. The design system id to use for generating the new screen, should always be configured for design consistency, via `get_project` or `list_assets` methods. If not provided, a default design system will be used. Example: `assets/15996705518239280238`. |
Output Schema
| Name | Required | Description |
|---|---|---|
| projectId | No | The project ID of the generated screen. This is the same as the input project ID. |
| sessionId | No | The session ID of the generated screen. |
| outputComponents | No | The generated output components. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description goes beyond annotations by detailing the asynchronous behavior: it can take minutes, may timeout, and may still succeed on connection errors. It also specifies the fallback polling mechanism and the handling of output components and suggestions. The annotations only state readOnlyHint=false and destructiveHint=true, so the description adds valuable behavioral context without contradicting them.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is well-structured with sections for instructions and output. It is appropriately sized for the tool's complexity (async operation, error handling, output handling). Every sentence serves a purpose: the initial purpose statement, the timing instructions, the fallback strategy, and the output handling. No redundant or filler text; it is dense but clear.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity (long-running, error-prone, with suggestions), the description is complete. It covers what to do on timeout, connection errors, how to retrieve the result, and how to handle suggestions. It also references an output schema implicitly through 'output_components'. The agent has everything needed to invoke the tool correctly and react to its outcomes.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3. The description adds extra value by providing guidance on the 'designSystem' parameter ('should always be configured for design consistency, via `get_project` or `list_assets` methods') and clarifying the 'projectId' format (example given). These details are not present in the schema descriptions, making the description helpful for correct parameter usage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose: 'Generates a new screen within a project from a text prompt.' This is a specific verb ('generates'), resource ('screen'), and input method ('text prompt'). It distinguishes itself from sibling tools like 'edit_screens' (which edits existing screens) and 'generate_variants' (which generates variants of existing screens) by the text prompt basis. The purpose is unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides explicit guidance on when and how to use the tool: it warns about long execution time, instructs not to retry on timeout, and directs the agent to poll with `get_screen` every 30 seconds up to 10 times. It also explains how to handle connection errors and what to do with output suggestions (re-call the tool with the accepted suggestion). This is clear, actionable usage context that differentiates this tool from its alternatives.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
generate_variantsADestructive
Generates variants of existing screens within a project using a text prompt.
Instructions for Tool Call:
If the tool fails with a timeout, don't retry. Instead, try to get the screen with
get_screenmethod every 30 seconds for up to 10 times before giving up.
| Name | Required | Description | Default |
|---|---|---|---|
| prompt | Yes | Required. The input text used to generate the variants. | |
| modelId | No | Optional. The model to use for generation. | |
| projectId | Yes | Required. The project ID of screens to generate variants for, example: '4044680601076201931', without the `projects/` prefix. | |
| deviceType | No | Optional. The type of device that captured the screenshot, e.g., mobile or desktop. | |
| variantOptions | Yes | Required. The variant options for generation, including the number of variants, creative range, and aspects to focus on. | |
| selectedScreenIds | Yes | Required. The screen ids of screen to generate variants for, example: ['98b50e2ddc9943efb387052637738f61', '98b50e2ddc9943efb387052637738f62'], without the `screens/` prefix. |
Output Schema
| Name | Required | Description |
|---|---|---|
| projectId | No | The project ID of the generated screen. This is the same as the input project ID. |
| sessionId | No | The session ID of the generated screen. |
| outputComponents | No | The generated output components. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description adds a genuinely useful behavioral disclosure: 'If the tool fails with a timeout, don't retry. Instead, try to get the screen with get_screen method every 30 seconds for up to 10 times before giving up.' This tells the agent the operation is effectively asynchronous and how to recover. It does not explain the destructiveHint=true implication beyond annotations, so it is not perfect.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is short, front-loaded with the core purpose, and followed by a concise, actionable timeout instruction. Every sentence earns its place; there is no redundant filler or repetition of schema details.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the rich schema, output schema, and annotations, the description covers the main non-obvious operational detail (timeout recovery via get_screen). It lacks explicit guidance on when to prefer this tool over siblings, but that gap is already captured in usage_guidelines. Overall, an agent has enough to call it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents projectId, selectedScreenIds, prompt, variantOptions, modelId, and deviceType. The description only restates 'existing screens' and 'text prompt', adding little semantic value beyond the schema. Baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The first sentence states a specific verb and resource: 'Generates variants of existing screens within a project using a text prompt.' This clearly identifies what the tool does and distinguishes it from creating screens from scratch. It does not explicitly name or differentiate from sibling tools like generate_screen_from_text or edit_screens, so it misses the top score.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to use this tool versus generate_screen_from_text, edit_screens, or other siblings. The only procedural note is about timeout handling with get_screen, which is error-recovery guidance, not selection guidance. This is insufficient for choosing between alternatives.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_projectARead-only
Retrieves the details of a specific Stitch project using its project name.
| Name | Required | Description | Default |
|---|---|---|---|
| name | Yes | Required. Identifier. The resource name of the project to retrieve. Format: `projects/{project}` Example: `projects/4044680601076201931` |
Output Schema
| Name | Required | Description |
|---|---|---|
| name | No | Identifier. The resource name of the project. Format: projects/{project} |
| title | No | Optional. The title of the project. |
| origin | No | Output only. The origin of the project. |
| metadata | No | Metadata of the project. |
| readTime | No | Output only. The time the project was last read. Populated only when listing recently viewed projects. |
| createTime | No | Output only. The time when the project was created. |
| deviceType | No | Optional. The device type of the project. |
| updateTime | No | Output only. The time when the project was last updated. |
| visibility | No | Optional. The visibility setting of the project. |
| designTheme | No | Output only. The theme used to generate the first design in the project. |
| projectType | No | Optional. The type of the project. If not specified, the project is a text to UI project. |
| backgroundTheme | No | Optional. The background theme of the project. |
| screenInstances | No | Output only. The screen instances on this project. |
| thumbnailScreenshot | No | Optional. The screenshot to be used as the thumbnail for the project. Same as normal design screenshots, this contains the FIFE serving_base_url which requires additional FIFE URL options to be set for sizing. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already establish readOnlyHint=true and destructiveHint=false, and the description's 'Retrieves' is consistent with those. The description adds little behavioral context beyond that, such as failure behavior or scope, but the annotations lower the burden for this simple read operation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single sentence that states the action, target resource, and lookup key with no filler. The most important information is front-loaded and every word earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
This is a low-complexity tool with one fully documented required parameter, an output schema, and read-only annotations. Nothing essential for correct invocation is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, with the parameter's format and example fully documented as `projects/{project}`. The description only restates that the lookup uses the project name, adding no meaning beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('Retrieves') and a clear object ('details of a specific Stitch project'), so the core purpose is immediately understandable. It does not explicitly contrast with sibling tools, but 'specific project' distinguishes it from list_projects and other project mutations.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies the tool is used when retrieving one project by name, but it gives no explicit when-to-use guidance or references to alternatives such as list_projects. An agent can infer basic usage, but exclusion criteria are absent.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_screenARead-only
Retrieves the details of a specific screen within a project.
| Name | Required | Description | Default |
|---|---|---|---|
| name | Yes | Required. Identifier. The resource name of the screen to retrieve. Format: `projects/{project}/screens/{screen}` Example: `projects/4044680601076201931/screens/98b50e2ddc9943efb387052637738f61` |
Output Schema
| Name | Required | Description |
|---|---|---|
| name | No | Identifier. The name of the screen. Format: projects/{project}/screens/{screen} |
| title | No | The title of the screen. |
| width | No | The width of the screen. |
| height | No | The height of the screen. |
| htmlCode | No | The HTML code of the screen. |
| deviceType | No | The device type of the screen. |
| screenshot | No | The screenshot of the screen. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, so the tool is known safe and non-mutating. The description adds no extra behavioral context beyond that; it doesn't discuss errors, access requirements, or output format. Meets baseline with annotation coverage.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One concise sentence that fully captures the purpose. No wasted words, front-loaded with key action and resource.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the simple single-parameter tool, output schema present, and read-only annotation, the description is sufficient. No critical information is missing for an agent to invoke it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema covers 100% of parameters with detailed description, including format and example. The description adds no extra semantic meaning beyond that. Baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Clear verb and resource: retrieves details of a specific screen. Distinguishes from list_screens implicitly, but doesn't name the sibling explicitly. Adequate for selection.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No explicit guidance on when to use this over siblings like list_screens or edit_screens. However, the schema's required 'name' format implies a single-screen fetch, and the presence of output schema suggests it's for detailed inspection. Usage is implied but not explicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_design_systemsARead-only
Lists all design systems for a given project.
| Name | Required | Description | Default |
|---|---|---|---|
| projectId | No | Optional. The project ID to list design systems for, example: '4044680601076201931', without the `projects/` prefix. If empty, lists all global design systems. |
Output Schema
| Name | Required | Description |
|---|---|---|
| designSystems | No | The design systems for the given project. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, and the description is consistent. The description adds the project scope but does not disclose additional behavior such as falling back to global design systems when projectId is empty; this is covered in the parameter schema instead of the tool description.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single sentence, front-loaded with the action and resource, and contains no redundant or unnecessary information. It is appropriately concise for a simple list tool.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's simplicity (one optional parameter, output schema present, readOnly annotation), the description sufficiently captures the core purpose. Additional details like return format are covered by the output schema, and parameter behavior by the schema, so the description is complete for its context.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema description fully covers the single parameter projectId, including its optionality and behavior when empty. The tool description does not add extra parameter information, so the schema coverage baseline applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses the specific verb 'Lists' and the resource 'design systems' with a scope ('for a given project'), making it clear what the tool does. It differentiates from sibling list tools like list_projects and list_screens by naming the exact resource.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage when needing design systems for a project, but it does not explicitly state when to use this over alternatives or provide exclusions. No sibling tool comparisons are made, so guidance is only implied from the purpose.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_projectsARead-only
Lists all Stitch projects accessible to the user. By default, it lists projects owned by the user.
| Name | Required | Description | Default |
|---|---|---|---|
| filter | No | Optional. A filter to apply to the list of projects, following a subset of AIP-160. This service supports filtering on the `view` field. Supported filters: * `view=owned`: Lists only projects owned by the user. This is the default behavior if no filter is specified. * `view=shared`: Lists only projects shared with the user. Example: `view=owned` |
Output Schema
| Name | Required | Description |
|---|---|---|
| projects | No | The projects for the given user, sorted by update time from the most recent to the least recent. Adding a new session to an existing project is considered as an update to the project. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, so the safety profile is known. The description adds the default filter behavior (owned), which is also covered in the schema's parameter description. It does not disclose additional behavioral traits like pagination or ordering. Given the annotations cover the main trait, a 3 is appropriate.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two concise sentences with no fluff. The primary purpose is front-loaded, and the default behavior is stated in the second sentence. Every word earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool has an output schema, so return values are not the description's job. The tool is simple with one optional parameter, and the description covers purpose and default. It does not explicitly mention that filtering is possible, but that is in the schema. Given the context, the description is sufficient and well-aligned with the tool's simplicity.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, and the parameter description includes supported filters and an example. The tool description adds no new meaning about the parameter beyond what the schema already states (default behavior is repeated). Since the schema handles the parameter thoroughly, baseline 3 is correct.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action (lists), the resource (Stitch projects), and the scope (accessible to the user). It also specifies the default behavior (owned by user), which distinguishes it from siblings like get_project or list_screens. An agent can immediately understand what this tool does without needing to open the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage (when you need to list all projects) but does not explicitly mention alternatives or when not to use it. It states the default behavior but gives no exclusions or references to sibling tools. For a simple list tool, this is adequate but leaves room for clearer routing guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_screensARead-only
Lists all screens within a given Stitch project.
| Name | Required | Description | Default |
|---|---|---|---|
| projectId | Yes | Required. Identifier. The project ID to list screens for, example: '4044680601076201931', without the `projects/` prefix. |
Output Schema
| Name | Required | Description |
|---|---|---|
| screens | No | The screens for the given project. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true consequentially, covering the read-only nature of the operation. The description adds the project scoping detail but does not disclose pagination, result limits, or other behavioral traits that may matter for a list operation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, focused sentence with no filler. It front-loads the verb, resource, and scope, making it easy for an agent to parse quickly.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the low complexity, one fully documented parameter, readOnlyHint annotation, and presence of an output schema, the description provides enough context for an agent to invoke the tool correctly. Return value details are already covered by the output schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, and the parameter description already covers requiredness, type, an example, and the prefix convention. The tool description adds no additional meaning beyond saying the project is a scope for the list.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb and resource, 'Lists all screens within a given Stitch project', making the tool's basic function clear. It is distinguishable from siblings like get_screen by the word 'all', but it does not explicitly name or contrast any alternative tool.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies this tool should be used when an agent needs all screens for a project, but it provides no explicit guidance about when to prefer list_screens over get_screen or list_projects. No exclusions or alternatives are mentioned.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
update_design_systemADestructive
Updates a design system for a project. Use this tool when the user wants to change the overall visual theme, style, or branding of the application. This includes configuring:
Color Palette: Presets, custom primary colors, and saturation levels.
Typography: Font families (e.g., Inter, Roboto, etc.).
Shape: Corner roundness for UI elements.
Appearance: Light and dark mode background colors.
Design MD: Free-form design instructions in markdown. This tool establishes the foundational design tokens that apply across all screens in the project.
| Name | Required | Description | Default |
|---|---|---|---|
| name | Yes | Required. Identifier. The resource name of the design system to update. Format: `assets/{asset_id}` Example: `assets/15996705518239280238` | |
| projectId | Yes | Required. The project ID to update design system for, example: '4044680601076201931', without the `projects/` prefix. | |
| designSystem | Yes | Required. The design system to update. |
Output Schema
| Name | Required | Description |
|---|---|---|
| name | No | Identifier. The resource name of the asset. Format: assets/{asset} |
| version | No | Output only. The version of this asset. 0 indicates unversioned (legacy data). Incremented when the asset content changes. |
| copiedFrom | No | Optional. The resource name of the asset this was copied from, if any. Format: assets/{asset} Tracks the fork history of assets. |
| designSystem | No | Optional. The design system. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true, so the destructive nature is known. The description adds that the tool 'establishes the foundational design tokens that apply across all screens in the project,' explaining the broad scope of the change. It does not detail overwrite semantics, but the annotation covers the destructive aspect, and the added context about cross-screen impact is valuable.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is well-structured with a clear opening sentence, a 'Use this tool when...' statement, and a concise bullet list of features. Every sentence adds useful context without fluff or redundancy, making it efficient and easy to scan.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity (nested designSystem object) and the presence of an output schema, the description covers the core 'what', 'when', and 'scope' (applies across all screens). It includes the key configurable categories, and the schema fills in the detailed parameter definitions. The description is complete enough for an agent to decide when to invoke the tool and what to include.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so parameters are already well-documented. The description adds high-level conceptual categories (Color Palette, Typography, Shape, Appearance, Design MD) that help an agent map user intent to the relevant fields in the nested designSystem object, going beyond the raw schema by grouping related parameters into thematic areas.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states 'Updates a design system for a project' and specifies the exact use case: 'when the user wants to change the overall visual theme, style, or branding of the application.' It enumerates the configurable aspects (color, typography, shape, appearance, design MD), making the tool's purpose unambiguous and distinct from creation or other operations.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly says 'Use this tool when the user wants to change the overall visual theme, style, or branding,' which is clear when-to-use guidance. It does not explicitly mention when not to use or name alternative sibling tools like create_design_system, so it lacks explicit exclusions or alternatives, but the context is sufficient.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
upload_design_mdADestructive
Uploads DESIGN.md to a Stitch project. Use this tool when the user wants to create a design system from a DESIGN.md file.
Instructions for Tool Call:
Call
create_design_system_from_design_mdtool immediately after this tool to create the design system from the uploaded DESIGN.md, and display the design system in the UI.
| Name | Required | Description | Default |
|---|---|---|---|
| projectId | Yes | Required. The project ID to upload the DESIGN.md to, example: '4044680601076201931', without the `projects/` prefix. | |
| designMdBase64 | Yes | Required. The base64-encoded DESIGN.md content. The decoded content must be valid UTF-8; uploads with invalid UTF-8 bytes will be rejected. Run `base64 -w 0 ` to get the base64-encoded string. |
Output Schema
| Name | Required | Description |
|---|---|---|
| x | No | Optional. The x position of the screen. |
| y | No | Optional. The y position of the screen. |
| id | No | Optional. The screen instance id. |
| type | No | Optional. The type of screen instance. |
| label | No | Optional. The screen label. |
| width | No | Optional. The screen width. |
| height | No | Optional. The screen height. |
| hidden | No | Optional. User driven action hiding screen from the canvas. |
| groupId | No | Optional. Group identifier for Genie Agent output grouping. Screens with the same groupId are rendered and interact as a group. |
| groupName | No | Optional. Human-readable name for the group (e.g., "Warm Minimalism", "Ethereal Glow"). |
| isResized | No | Optional. Whether this screen has been resized by the user. |
| isFavourite | No | Optional. Whether this screen is marked as a favourite by the user. |
| needsLayout | No | Optional. Whether this screen instance needs frontend layout positioning. Set by the PostProcessor when auto-linking screens that the agent generated but the frontend has not yet positioned on the canvas. |
| sourceAsset | No | Optional. The resource name of the source asset. Format: assets/{asset}. Only one of the following fields should be set: this field, sourceScreen, textContent. |
| textContent | No | Optional. Text content for TEXT_INSTANCE nodes. Only one of the following fields should be set: this field, sourceScreen, sourceAsset. |
| sourceScreen | No | Optional. The resource name of the source screen. Format: projects/{project}/screens/{screen}. Only one of the following fields should be set: this field, sourceAsset, textContent. |
| variantScreenInstance | No | Optional. The variant Screen Instance. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already mark the tool as destructive, and the description does not contradict that. It adds the valuable workflow constraint that this tool must be followed by `create_design_system_from_design_md`, but it does not disclose what the destructive effect impacts (e.g., overwriting an existing DESIGN.md), so it only partially exceeds what annotations provide.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact, with a front-loaded purpose sentence and a clearly separated bolded instruction block. Every sentence contributes necessary guidance, with no padding.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a two-parameter upload tool with full schema coverage and an output schema present, the description supplies the purpose, the trigger condition, and the mandatory follow-up call. Nothing essential for correct invocation and orchestration is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The parameter schema fully documents both `projectId` and `designMdBase64`, including format and base64 requirements. The description itself only reinforces that the content is DESIGN.md, so it adds little beyond the schema, meriting the baseline score.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with 'Uploads DESIGN.md to a Stitch project,' naming a specific verb and resource. It also clarifies when to use it and references the sibling `create_design_system_from_design_md` as the follow-up, making it distinguishable from the related creation tool.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explicitly states 'Use this tool when the user wants to create a design system from a DESIGN.md file' and then mandates calling `create_design_system_from_design_md` immediately afterward. This gives the agent both a clear trigger condition and the required next step.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
16 tool updates
v1.0.0- First observed
apply_design_system - First observed
create_design_system - First observed
create_design_system_from_design_md - First observed
create_project - First observed
delete_project - First observed
download_assets - First observed
edit_screens - First observed
generate_screen_from_text - First observed
generate_variants - First observed
get_project - First observed
get_screen - First observed
list_design_systems - First observed
list_projects - First observed
list_screens - First observed
update_design_system - First observed
upload_design_md
TDQS
Scored across 16 tools
Most tools target distinct resources and actions (projects, screens, design systems, assets), making selection relatively straightforward. However, create_design_system and create_design_system_from_design_md overlap in purpose, and edit_screens vs generate_variants could be confused since both modify existing screens.
Tool names consistently follow a verb_noun pattern (create_project, get_screen, list_design_systems, apply_design_system, etc.). Even longer names like create_design_system_from_design_md and generate_screen_from_text follow the same predictable structure.
At 16 tools, the server is slightly large but still well-scoped across four coherent areas: projects, screens, design systems, and assets. A few tools could be consolidated (e.g., upload_design_md and create_design_system_from_design_md are tightly coupled), but the count is reasonable for the domain.
The core workflows are covered: project CRUD, screen listing/retrieval/generation/editing, and design system create/update/list/apply. Minor gaps include no delete for screens or design systems and no project update tool, but these are workable and not likely to cause major agent failures.
Maintenance
Related MCP Connectors
Nifty's MCP server — exposes tasks, projects, messages, and files as tools for AI agents.
- SupabaseOAuthcom.supabase
MCP server for interacting with the Supabase platform
MCP server for Statsig API - interact with Statsig's feature flags, experiments, and analytics
Remote MCP server for supportsheep: run AI interviews and manage support content for your blog.
Related MCP Servers
- AlicenseNot gradedqualityFmaintenanceAn automated MCP server for Google Stitch that enables AI-driven UI design generation, accessibility audits, and design system management. It streamlines workflows for creating responsive screens, extracting design tokens, and maintaining visual consistency across professional web projects.47 npm31Apache 2.0
- AlicenseNot gradedqualityDmaintenanceA universal MCP server for Google Stitch that enables AI-powered UI/UX design generation by extracting design context and metadata from existing screens. It allows users to fetch screen code and images to create consistent, styled UI components across multiple projects.500 npm121Apache 2.0
- FlicenseNot gradedqualityDmaintenanceAn MCP server that exposes Google Stitch's AI-powered UI design capabilities as local tools. It enables project management, screen generation/editing, and design system operations by forwarding requests to Google's Stitch API.2-
- AlicenseAqualityDmaintenanceAn intelligent MCP server for Google Stitch that generates production-ready UI from text prompts, with auto-orchestration of design systems, WCAG accessibility, responsive breakpoints, and framework conversion (React, Vue, Svelte).1719 npm9MIT