PromptingBox MCP Server
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@PromptingBox MCP Serversave this last response as 'Code Review Checklist' in my Work folder"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
@promptingbox/mcp
MCP (Model Context Protocol) server for PromptingBox — save, manage, and organize prompts directly from Claude.ai, Claude Desktop, Cursor, Windsurf, and other MCP-compatible AI tools.
Quick Start
There are two ways to connect:
Method | Best for | Install needed? |
Claude.ai conversations, Cowork | No | |
Claude Desktop, Cursor, Windsurf, Claude Code | Yes |
Both methods give you the same 20 tools and connect to the same account.
Related MCP server: Prompt Bookmarks
Claude.ai (Web / Cowork)
No install needed — works entirely in the browser. This connects PromptingBox to claude.ai for web conversations and Cowork (Claude's cloud agent mode).
Steps
Get your API key — Go to Settings → MCP Integration and create a key.
Add a custom connector — In claude.ai, go to Settings → Connectors → Add custom connector.
Enter the connector URL:
https://www.promptingbox.com/api/mcp-transportName it
pbox(recommended) — This lets you say "save this to pbox" naturally. You can use any name, but Claude will use whatever name you enter here.Click Connect — A window opens asking for your API key. Paste it and click Authorize.
Done! Start a new conversation and try it out.
Naming tip
We recommend pbox so you can say "save to pbox" naturally in conversation. If you choose a different name (e.g. "my-prompts"), use that name instead when talking to Claude (e.g. "save to my-prompts"). Claude uses the exact name you entered when adding the connector.
Examples in Claude.ai
"Save this conversation as a prompt to pbox"
"List all my pbox prompts"
"Search pbox for my email templates"
"Save this as 'Meeting Notes Template' in my Work folder on pbox"
"Get my prompt called 'Code Review' from pbox"
"Which pbox account am I connected to?"Local Setup (npm)
For Claude Desktop, Cursor, Windsurf, and Claude Code. Requires Node.js 18+.
1. Get your API key
Go to PromptingBox Settings → MCP Integration and create an API key.
2. Install the server (run once)
npm install -g @promptingbox/mcp3. Configure your AI tool
Tip: Name the server
pboxso you can naturally say things like "save this to pbox" or "list my pbox prompts".
Claude Desktop
Edit ~/Library/Application Support/Claude/claude_desktop_config.json (macOS) or %APPDATA%\Claude\claude_desktop_config.json (Windows):
{
"mcpServers": {
"pbox": {
"command": "promptingbox-mcp",
"env": {
"PROMPTINGBOX_API_KEY": "pb_your_key_here"
}
}
}
}Cursor
Add to .cursor/mcp.json in your project (or global config):
{
"mcpServers": {
"pbox": {
"command": "promptingbox-mcp",
"env": {
"PROMPTINGBOX_API_KEY": "pb_your_key_here"
}
}
}
}Windsurf
Add to your Windsurf MCP config:
{
"mcpServers": {
"pbox": {
"command": "promptingbox-mcp",
"env": {
"PROMPTINGBOX_API_KEY": "pb_your_key_here"
}
}
}
}Claude Code
Add to .claude/mcp.json in your project:
{
"mcpServers": {
"pbox": {
"command": "promptingbox-mcp",
"env": {
"PROMPTINGBOX_API_KEY": "pb_your_key_here"
}
}
}
}4. Restart your AI tool
Restart Claude Desktop, Cursor, or Windsurf for the MCP server to be detected.
Usage
Once configured, you can say things like:
Saving & retrieving prompts:
"Save this prompt to pbox"
"Save this as 'Code Review Checklist' in my Work folder on pbox"
"Save this to pbox with tags 'python' and 'debugging'"
"Get my prompt called 'Code Review'"
"Search my pbox prompts for API"
"List all my pbox prompts"
Editing & managing prompts:
"Update the content of 'Code Review'"
"Delete the prompt called 'Old Draft'"
"Duplicate 'Code Review'"
"Star my 'Code Review' prompt"
"Move 'Brainstorm Ideas' to my Marketing folder"
Folders & tags:
"Create a folder called 'Work'"
"Move 'Code Review' to my Work folder"
"Delete my 'Old' folder"
"Tag 'Code Review' with testing and automation"
"List my pbox folders"
"List my pbox tags"
Version history:
"Show version history for 'Code Review'"
"Restore version 2 of 'Code Review'"
Templates:
"Search pbox templates for email"
"Save that template to my collection"
Account:
"Which pbox account am I using?"
Available Tools
Prompt Management
Tool | Description |
| Save a prompt with title, content, optional folder and tags |
| Get the full content and metadata of a prompt |
| Search prompts by title, content, tag, folder, or favorites |
| Update a prompt's title and/or content (auto-versions) |
| Permanently delete a prompt and all its versions |
| Create a copy of an existing prompt |
| Star or unstar a prompt |
| List all prompts grouped by folder |
Folder Management
Tool | Description |
| List all folders in your account |
| Create a new folder (or return existing one) |
| Delete a folder (prompts move to root, not deleted) |
| Move a prompt to a different folder |
Tag Management
Tool | Description |
| List all tags in your account |
| Set tags on a prompt (replaces existing, auto-creates new tags) |
| Delete a tag from your account and all prompts |
Version History
Tool | Description |
| View all saved versions of a prompt |
| Restore a prompt to a previous version |
Templates
Tool | Description |
| Browse and search the public template library |
| Save a public template to your collection |
Account
Tool | Description |
| Show which PromptingBox account is connected |
Environment Variables
Variable | Required | Description |
| Yes | Your PromptingBox API key (starts with |
| No | Override the API base URL (default: |
FAQ
Can I use both the web connector and local MCP? Yes! The web connector (Claude.ai) and local MCP (Claude Desktop) can coexist. They share the same tools and account. The desktop version works in Claude Desktop app, while the web connector works in claude.ai conversations and Cowork.
What name should I use?
We recommend pbox — it's short and lets you say "save to pbox" naturally. You can also use promptingbox or any name you like. Just use your chosen name when talking to Claude.
Is my data secure? Yes. For local setups, the server runs on your machine over HTTPS. For the web connector, authentication uses OAuth 2.1 with PKCE — your API key is never exposed to Claude directly. We store only SHA-256 hashes of API keys, never raw keys.
Is the server open source? Yes. Source code is at github.com/promptingbox/mcp.
License
MIT
Available Tools
20 toolsadd_tagsA
Set tags on a prompt (by tag name). This replaces all existing tags on the prompt. Tags are auto-created if they don't exist.
| Name | Required | Description | Default |
|---|---|---|---|
| promptId | No | The prompt ID. Provide this or promptTitle. | |
| promptTitle | No | The prompt title to search for. Provide this or promptId. | |
| tagNames | Yes | Tag names to set on the prompt (replaces existing tags) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It effectively describes key behaviors: the destructive nature ('replaces all existing tags'), side effects ('Tags are auto-created if they don't exist'), and the operation's scope ('Set tags on a prompt'). It does not cover aspects like error handling, permissions, or rate limits, but provides sufficient context for basic use.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with the core purpose in the first sentence, followed by two concise sentences that add critical behavioral details. Every sentence earns its place by providing essential information without redundancy, making it efficient and well-structured.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's moderate complexity (3 parameters, no output schema, no annotations), the description is fairly complete. It covers the purpose, key behaviors, and parameter implications. However, it lacks details on return values or error cases, which would be beneficial since there is no output schema. It is adequate but has minor gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all parameters thoroughly. The description adds minimal value beyond the schema by reinforcing the replacement behavior for 'tagNames' and clarifying the prompt identification logic ('by tag name'), but does not provide additional syntax or format details. This meets the baseline for high schema coverage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the specific action ('Set tags on a prompt'), identifies the resource ('prompt'), and distinguishes it from siblings by specifying the replacement behavior ('replaces all existing tags') and auto-creation feature ('Tags are auto-created if they don't exist'). This goes beyond just restating the tool name.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage context by mentioning 'replaces all existing tags' and 'auto-created if they don't exist', which helps understand when to use it (e.g., for full tag replacement vs. incremental updates). However, it does not explicitly state when to use this tool versus alternatives like 'update_prompt' or 'list_tags', nor does it provide exclusions or prerequisites.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
create_folderA
Create a new folder in PromptingBox. If a folder with the same name exists, returns the existing one.
| Name | Required | Description | Default |
|---|---|---|---|
| name | Yes | The folder name to create |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries full burden. It discloses key behavior: creation with idempotent handling of duplicates. However, it lacks details on permissions, error conditions, or response format, which are important for a mutation tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences with zero waste: the first states the core purpose, the second adds crucial behavioral nuance. It is front-loaded and appropriately sized for a single-parameter tool.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a mutation tool with no annotations and no output schema, the description is minimal but covers the essential behavior. It lacks details on permissions, error handling, and return values, which would improve completeness given the context.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents the 'name' parameter fully. The description adds no additional semantic context beyond implying it's used for folder creation, meeting the baseline for high schema coverage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action ('Create a new folder') and the resource ('in PromptingBox'), with specific behavior about handling duplicates ('returns the existing one'). It distinguishes from siblings like 'list_folders' (read-only) and 'delete_folder' (destructive).
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage for creating folders, but does not explicitly state when to use this tool versus alternatives (e.g., 'list_folders' for checking existence, 'move_prompt_to_folder' for organizing content). No exclusions or prerequisites are mentioned.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
delete_folderA
Delete a folder from PromptingBox. Prompts in the folder are moved to the default folder (not deleted). The default folder cannot be deleted.
| Name | Required | Description | Default |
|---|---|---|---|
| folderId | No | The folder ID. Provide this or folderName. | |
| folderName | No | The folder name. Provide this or folderId. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It effectively describes key behaviors: it's a destructive operation (deletes folders), specifies side effects (prompts are moved, not deleted), and includes a constraint (default folder cannot be deleted). It doesn't mention authentication needs, rate limits, or error conditions, but covers the essential mutation behavior well.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is perfectly concise with three sentences that each earn their place: states the action, explains the side effect, and provides a critical constraint. It's front-loaded with the primary purpose and wastes no words.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given this is a destructive tool with no annotations and no output schema, the description does well by explaining the deletion behavior, what happens to contained prompts, and the default folder constraint. It could be more complete by mentioning error conditions or return values, but it covers the essential context for a folder deletion operation adequately.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, with both parameters clearly documented in the schema (folderId and folderName, with 'Provide this or folderName/folderId'). The description adds no additional parameter information beyond what the schema provides, so it meets the baseline of 3 for high schema coverage without compensating value.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action ('Delete a folder from PromptingBox') and specifies what happens to contained prompts ('Prompts in the folder are moved to the default folder (not deleted)'), distinguishing it from siblings like delete_prompt (which deletes prompts) and delete_tag (which deletes tags). It provides specific verb+resource+consequences.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides clear context for when to use this tool (to delete folders) and includes an important exclusion ('The default folder cannot be deleted'), which helps prevent misuse. However, it doesn't explicitly mention alternatives like moving prompts with move_prompt_to_folder or when to use folderId vs folderName, so it's not a perfect 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
delete_promptA
Permanently delete a prompt from PromptingBox. This also deletes all versions and tag associations.
| Name | Required | Description | Default |
|---|---|---|---|
| promptId | No | The prompt ID. Provide this or promptTitle. | |
| promptTitle | No | The prompt title to search for. Provide this or promptId. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It clearly indicates this is a destructive operation ('Permanently delete') and specifies the extent of deletion ('all versions and tag associations'), which is crucial context. However, it does not mention potential side effects like error handling, confirmation steps, or irreversible consequences beyond the stated scope.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with the core action and resource, followed by additional scope details in a second sentence. Both sentences earn their place by providing essential information without redundancy, making it efficient and well-structured.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given no annotations, no output schema, and a destructive operation, the description is adequate but has gaps. It covers the action and scope well, but lacks details on error conditions, return values, or confirmation requirements. For a deletion tool, more context on irreversibility or success indicators would enhance completeness.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents both parameters (promptId and promptTitle) and their optional nature. The description does not add any meaning beyond what the schema provides, such as explaining parameter interactions or usage examples. Baseline 3 is appropriate as the schema handles the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the specific action ('Permanently delete') and resource ('a prompt from PromptingBox'), and distinguishes it from siblings like delete_folder or delete_tag by specifying prompt deletion. It also mentions the scope of deletion ('all versions and tag associations'), which further clarifies its purpose.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage when permanently deleting a prompt, but does not explicitly state when to use this tool versus alternatives (e.g., delete_folder for folders, or restore_version for undoing deletions). It lacks guidance on prerequisites or exclusions, such as whether the prompt must exist or be accessible.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
delete_tagB
Delete a tag entirely from PromptingBox. Removes it from all prompts that use it.
| Name | Required | Description | Default |
|---|---|---|---|
| tagId | No | The tag ID. Provide this or tagName. | |
| tagName | No | The tag name. Provide this or tagId. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries full burden. It discloses the destructive behavior ('Delete... entirely', 'Removes it from all prompts'), which is critical. However, it lacks details on permissions needed, whether deletion is reversible, error conditions (e.g., if tag doesn't exist), or rate limits. For a destructive tool with zero annotation coverage, this is insufficient.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two concise sentences that are front-loaded with the core action. The first sentence states the purpose, and the second adds crucial scope information. No wasted words or redundant phrasing.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a destructive tool with no annotations and no output schema, the description is incomplete. It covers the basic action and scope but misses critical context: what happens on success/failure, return values, authentication needs, or error handling. Given the complexity of deletion and impact on prompts, more behavioral disclosure is needed.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, with both parameters clearly documented in the schema. The description adds no parameter-specific information beyond what the schema provides (e.g., no examples or edge cases). Baseline 3 is appropriate since the schema handles parameter documentation adequately.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the specific action ('Delete a tag entirely') and resource ('from PromptingBox'), with additional scope clarification ('Removes it from all prompts that use it'). It distinguishes from sibling tools like 'list_tags' or 'add_tags' by specifying destructive removal.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No explicit guidance on when to use this tool versus alternatives. While the description implies this is for permanent tag deletion, it doesn't mention prerequisites (e.g., tag must exist), exclusions, or alternatives like updating tags instead. The sibling list includes 'list_tags' which could be used first, but this isn't stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
duplicate_promptB
Create a copy of an existing prompt. The copy gets "(Copy)" appended to the title and inherits the same folder and tags.
| Name | Required | Description | Default |
|---|---|---|---|
| promptId | No | The prompt ID. Provide this or promptTitle. | |
| promptTitle | No | The prompt title to search for. Provide this or promptId. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries full burden but only covers basic behavior (copying with title modification and inheritance). It misses critical details like whether this requires specific permissions, if it's idempotent, what happens on errors, or the response format. For a mutation tool with zero annotation coverage, this is insufficient.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two concise sentences with zero waste—each sentence adds essential information about the action and the copy's properties. It's front-loaded with the core purpose and efficiently structured.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a mutation tool with no annotations and no output schema, the description is incomplete. It doesn't address permissions, error handling, or what the tool returns (e.g., the new prompt's ID). Given the complexity of duplicating resources, more behavioral context is needed.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so parameters are fully documented in the schema. The description adds no parameter-specific information beyond what the schema already states (e.g., it doesn't clarify if both promptId and promptTitle can be provided together or which takes precedence). Baseline 3 is appropriate when schema does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the specific action ('Create a copy') and resource ('an existing prompt'), distinguishing it from siblings like 'update_prompt' or 'save_prompt'. It explicitly mentions what gets copied and how the copy is modified, making the purpose unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage when duplicating a prompt, but provides no explicit guidance on when to use this versus alternatives like 'update_prompt' for modifications or 'save_prompt' for creating new prompts from scratch. It lacks context about prerequisites or exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_promptA
Get the full content of a single prompt from PromptingBox. Returns title, content, tags, folder, and metadata.
| Name | Required | Description | Default |
|---|---|---|---|
| promptId | No | The prompt ID. Provide this or promptTitle. | |
| promptTitle | No | The prompt title to search for. Provide this or promptId. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It discloses the return values (title, content, tags, folder, metadata), which is useful behavioral context. However, it lacks details on error handling, authentication needs, rate limits, or whether it's a read-only operation (implied by 'Get' but not explicit). The description adds some value but is incomplete for behavioral transparency.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, efficient sentence that front-loads the action and resource, followed by the return details. Every word earns its place with no redundancy or waste, making it highly concise and well-structured.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given no annotations and no output schema, the description provides basic completeness by stating the return values. However, for a tool with 2 parameters and no structured output, it should ideally include more context on usage scenarios, error cases, or behavioral traits. It's adequate but has clear gaps in contextual detail.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, with clear descriptions for promptId and promptTitle in the input schema. The description does not add any parameter semantics beyond what the schema provides (e.g., it doesn't explain format or constraints). Baseline is 3 since the schema does the heavy lifting, but no extra value is added.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the specific action ('Get the full content') and resource ('a single prompt from PromptingBox'), distinguishing it from siblings like list_prompts (which lists multiple) or search_prompts (which searches). It explicitly mentions what is returned (title, content, tags, folder, metadata), making the purpose unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage for retrieving a specific prompt's details, but does not explicitly state when to use it versus alternatives like search_prompts or list_prompts. However, the context is clear (getting full content of a single prompt), and the input schema hints at alternatives by allowing either promptId or promptTitle, though no explicit guidance is provided.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_foldersB
List all folders in the user's PromptingBox account. Useful to know where to save a prompt.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It states the tool lists folders but doesn't cover critical aspects like whether it's read-only (implied but not explicit), pagination behavior, error conditions, or authentication needs. For a tool with zero annotation coverage, this leaves significant gaps in understanding how it behaves beyond basic functionality.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is brief and front-loaded with the core purpose in the first sentence. The second sentence adds practical context without redundancy. Both sentences earn their place by clarifying usage, though it could be slightly more structured (e.g., explicitly stating it's a read operation).
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's low complexity (0 parameters, no output schema, no annotations), the description is minimally adequate. It covers what the tool does and hints at usage, but lacks details on behavior, output format, or error handling. For a simple list tool, this might suffice, but it doesn't fully compensate for the absence of annotations or output schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema has 0 parameters with 100% coverage, so no parameter documentation is needed. The description appropriately doesn't discuss parameters, focusing instead on the tool's purpose. This aligns with the baseline expectation for zero-parameter tools, where the description adds value by explaining the tool's role rather than repeating schema details.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action ('List all folders') and resource ('in the user's PromptingBox account'), making the purpose immediately understandable. It distinguishes from siblings like 'create_folder' or 'move_prompt_to_folder' by focusing on listing rather than modifying. However, it doesn't explicitly differentiate from other list tools like 'list_prompts' or 'list_tags' beyond the resource type.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides implied usage guidance with 'Useful to know where to save a prompt,' suggesting it's for discovery before saving operations. However, it lacks explicit when-to-use vs. alternatives (e.g., no mention of when to use this over 'search_prompts' for folder-related queries) or any exclusions, leaving some ambiguity for the agent.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_promptsB
List all prompts in the user's PromptingBox account grouped by folder. Use this to see what prompts exist and where they are organized.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It mentions grouping by folder and the general purpose, but doesn't disclose key behavioral traits such as whether it's read-only (implied by 'List'), pagination, rate limits, authentication needs, or what the output format looks like. For a tool with no annotation coverage, this leaves significant gaps in understanding its behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two concise sentences that are front-loaded with the core purpose ('List all prompts...') and followed by usage context. Every sentence adds value without redundancy, making it efficient and well-structured for quick understanding.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's low complexity (0 parameters, no output schema, no annotations), the description is minimally adequate. It covers the basic purpose and grouping aspect, but lacks details on output format, behavioral constraints, or explicit sibling differentiation. With no annotations or output schema, more context on behavior would improve completeness, but it meets the minimum for a simple list tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has 0 parameters, and schema description coverage is 100%, so there's no need for parameter details in the description. The baseline for 0 parameters is 4, as the description appropriately doesn't discuss parameters and focuses on the tool's function. No additional semantic value is required or provided.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose with a specific verb ('List all prompts') and resource ('in the user's PromptingBox account'), and mentions grouping by folder. It distinguishes from siblings like 'get_prompt' (single prompt) and 'search_prompts' (filtered search), though not explicitly. However, it doesn't fully differentiate from 'list_folders' (which lists folders, not prompts), keeping it at 4 rather than 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage for seeing existing prompts and their organization, which suggests it's for browsing rather than targeted retrieval. It doesn't explicitly state when to use this versus alternatives like 'search_prompts' for filtering or 'get_prompt' for a single prompt, nor does it mention prerequisites or exclusions. This provides some context but lacks explicit guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_tagsA
List all tags in the user's PromptingBox account. Useful to know what tags are available when saving a prompt.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It implies a read-only operation by using 'List' and mentions the scope ('user's PromptingBox account'), but does not cover aspects like authentication needs, rate limits, or return format. This leaves gaps in behavioral understanding for a tool with zero annotation coverage.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences, front-loaded with the core purpose and followed by a usage guideline. Every sentence earns its place by adding clarity and context without any wasted words, making it highly efficient and well-structured.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's simplicity (0 parameters, no output schema, no annotations), the description is adequate but has clear gaps. It explains the purpose and usage but lacks details on behavioral traits like authentication or return format, which are important for completeness even in simple tools. This meets the minimum viable standard with room for improvement.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema has 0 parameters with 100% coverage, so the schema fully documents the lack of inputs. The description adds value by clarifying the scope ('user's PromptingBox account') and purpose, which compensates for the minimal parameter information. A baseline of 4 is appropriate as it provides useful context beyond the empty schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the specific action ('List all tags') and resource ('in the user's PromptingBox account'), distinguishing it from sibling tools like 'add_tags' or 'delete_tag' which modify tags rather than list them. It provides a complete picture of what the tool does without being vague or tautological.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly states when to use this tool ('Useful to know what tags are available when saving a prompt'), providing clear context for its application. However, it does not specify when not to use it or name alternatives (e.g., using 'search_prompts' for tag-based filtering), which prevents a perfect score.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_versionsA
Get the version history for a prompt. Shows all saved versions with their version numbers, notes, and timestamps.
| Name | Required | Description | Default |
|---|---|---|---|
| promptId | No | The prompt ID. Provide this or promptTitle. | |
| promptTitle | No | The prompt title to search for. Provide this or promptId. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries full burden. It clearly indicates this is a read operation ('Get', 'Shows') and specifies the scope of returned data. However, it doesn't disclose behavioral aspects like pagination, sorting order, error conditions, or authentication requirements that would be helpful for an agent.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is perfectly concise with two sentences that each earn their place. The first sentence states the core action and resource, while the second elaborates on what information is returned. No wasted words, and the most important information is front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only tool with 2 parameters and 100% schema coverage but no output schema, the description provides adequate context about what the tool does and returns. However, without annotations or output schema, it lacks information about return format structure, error handling, and other behavioral details that would make it more complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents both parameters thoroughly. The description doesn't add any parameter-specific information beyond what's in the schema descriptions. It implies the tool needs either a promptId or promptTitle but doesn't elaborate on format, validation, or search behavior.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the specific action ('Get the version history') and resource ('for a prompt'), distinguishing it from siblings like get_prompt (which retrieves current version) or restore_version (which modifies state). It specifies what information is returned ('all saved versions with their version numbers, notes, and timestamps'), making the purpose unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage context by specifying 'for a prompt' and listing what information is returned, but doesn't explicitly state when to use this versus alternatives like get_prompt (for current version) or restore_version (to revert to a specific version). It provides clear context about the tool's function but lacks explicit comparison or exclusion guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
move_prompt_to_folderA
Move a prompt to a different folder. Provide either the prompt ID or the prompt title — if a title is given, it will be looked up automatically. Does not change or delete prompt content.
| Name | Required | Description | Default |
|---|---|---|---|
| promptId | No | The prompt ID (from list_prompts). Provide this or promptTitle. | |
| promptTitle | No | The prompt title to search for. Provide this or promptId. | |
| folder | Yes | The folder name to move the prompt into |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden. It discloses that the tool moves prompts without altering content, which is useful behavioral context. However, it lacks details on permissions needed, error handling (e.g., if folder doesn't exist), or whether the move is reversible, leaving gaps for a mutation tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with the core action in the first sentence, followed by parameter guidance and a key behavioral note. Every sentence earns its place by clarifying tool function and constraints without redundancy, making it efficient and well-structured.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a mutation tool with no annotations and no output schema, the description is adequate but incomplete. It covers the basic action and parameter interplay but misses details like return values, error cases, or side effects (e.g., impact on prompt references), which are important given the tool's complexity.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all parameters (promptId, promptTitle, folder) with their purposes and relationships. The description adds minimal value beyond this, mentioning the ID/title lookup but not elaborating on semantics like folder naming conventions or search behavior.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the specific action ('Move'), the resource ('a prompt'), and the target ('to a different folder'), which distinguishes it from siblings like duplicate_prompt (creates copy), delete_prompt (removes), or update_prompt (modifies content). It explicitly notes it 'Does not change or delete prompt content,' further differentiating from content-altering tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage when reorganizing prompts between folders, but it does not explicitly state when to use this tool versus alternatives like create_folder (for making new folders) or update_prompt (which might handle folder changes differently). No exclusions or prerequisites are mentioned, leaving some ambiguity in context.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
restore_versionB
Restore a prompt to a previous version. Creates a new version with the restored content.
| Name | Required | Description | Default |
|---|---|---|---|
| promptId | No | The prompt ID. Provide this or promptTitle. | |
| promptTitle | No | The prompt title to search for. Provide this or promptId. | |
| versionNumber | Yes | The version number to restore to |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It states the tool 'creates a new version with the restored content,' implying a mutation, but lacks details on permissions, whether the original version is preserved, response format, or error handling. This is a significant gap for a mutation tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded and concise, consisting of two clear sentences that directly explain the tool's purpose and outcome without any wasted words. Every sentence earns its place by adding value.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the complexity of a mutation tool with no annotations and no output schema, the description is incomplete. It doesn't cover behavioral aspects like side effects, error cases, or what the tool returns, leaving gaps that could hinder an AI agent's ability to use it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema description coverage is 100%, so the schema already documents all three parameters thoroughly. The description adds no additional meaning beyond what the schema provides, such as clarifying the relationship between 'promptId' and 'promptTitle' or the implications of 'versionNumber.' Baseline 3 is appropriate when the schema does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the specific action ('Restore a prompt to a previous version') and the resource ('prompt'), distinguishing it from siblings like 'update_prompt' or 'duplicate_prompt' by focusing on version restoration rather than general editing or copying.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is provided on when to use this tool versus alternatives. It doesn't mention prerequisites (e.g., needing existing versions from 'list_versions'), exclusions, or comparisons to tools like 'update_prompt' for making new changes instead of restoring old ones.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
save_promptA
Save a prompt to the user's PromptingBox account. Use this when the user wants to save, store, or bookmark a prompt. If no folder is specified, saves to the default folder.
| Name | Required | Description | Default |
|---|---|---|---|
| title | Yes | A short, descriptive title for the prompt | |
| content | Yes | The full prompt content to save | |
| folder | No | Folder name to save into (created if it doesn't exist) | |
| tagNames | No | Tag names to apply (created if they don't exist) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden of behavioral disclosure. It mentions the default folder behavior, which adds some context, but does not cover critical aspects like authentication needs, error handling (e.g., what happens if the title already exists), or response format. For a mutation tool with zero annotation coverage, this leaves significant gaps in understanding how the tool behaves beyond basic functionality.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is appropriately sized and front-loaded, consisting of two sentences that directly state the purpose and a key usage note. Every sentence earns its place by providing essential information without redundancy or fluff, making it efficient and easy to parse.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's moderate complexity (a mutation operation with 4 parameters) and no annotations or output schema, the description is partially complete. It covers the basic purpose and default behavior but lacks details on authentication, error handling, or return values. While concise, it does not fully address the gaps left by missing structured data, making it adequate but with clear room for improvement.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all parameters thoroughly. The description adds no additional meaning beyond what the schema provides (e.g., no extra details on title/content constraints or folder/tag creation logic). With high schema coverage, the baseline score of 3 is appropriate, as the description doesn't compensate but also doesn't need to given the schema's completeness.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action ('Save a prompt') and resource ('to the user's PromptingBox account'), distinguishing it from siblings like 'update_prompt' or 'duplicate_prompt' by focusing on initial storage rather than modification or copying. It specifies the purpose as saving, storing, or bookmarking a prompt, making the verb+resource combination specific and unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides clear context for when to use this tool: 'when the user wants to save, store, or bookmark a prompt.' It also mentions a default behavior ('If no folder is specified, saves to the default folder'), which helps guide usage. However, it does not explicitly state when not to use it or name alternatives (e.g., vs. 'update_prompt' for modifications), missing full sibling differentiation.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
search_promptsB
Search prompts in PromptingBox by title, content, tag, folder, or favorites. Returns matching prompts.
| Name | Required | Description | Default |
|---|---|---|---|
| query | No | Search text to match against title and content | |
| tag | No | Filter by tag name | |
| folder | No | Filter by folder name | |
| favorites | No | Set to true to only show favorited prompts |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It states the search functionality and return outcome, but lacks details on permissions, rate limits, pagination, error handling, or whether it's read-only/destructive. For a search tool with zero annotation coverage, this is a significant gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, efficient sentence that front-loads the purpose ('Search prompts...') and includes key details (searchable fields and outcome) without any wasted words. Every part earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given 4 parameters with full schema coverage and no output schema, the description is adequate for a basic search tool but lacks completeness. It doesn't cover behavioral aspects like permissions or pagination, and without annotations or output schema, the agent may struggle with implementation details. A 3 reflects a minimum viable description with clear gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all 4 parameters. The description adds value by listing the searchable fields (title, content, tag, folder, favorites), which aligns with the parameters, but doesn't provide additional syntax, format, or interaction details beyond what the schema specifies. Baseline 3 is appropriate when the schema does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action ('Search') and resource ('prompts in PromptingBox'), specifying searchable fields (title, content, tag, folder, favorites) and the outcome ('Returns matching prompts'). However, it doesn't explicitly differentiate from sibling tools like 'list_prompts' or 'search_templates', which would require a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is provided on when to use this tool versus alternatives such as 'list_prompts' (for unfiltered listing) or 'search_templates' (for a different resource). The description implies usage through the searchable fields but lacks explicit when/when-not instructions or named alternatives.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
search_templatesB
Browse and search the PromptingBox public template library. Find pre-built prompts you can save to your collection.
| Name | Required | Description | Default |
|---|---|---|---|
| query | No | Search text to match against template titles and descriptions | |
| category | No | Filter by category (e.g. "Business", "Writing", "Development") | |
| limit | No | Number of results to return (default 10) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries full burden. It mentions 'browse and search' and that results can be saved, but doesn't disclose behavioral traits like whether this is a read-only operation, if it requires authentication, rate limits, pagination behavior, or what the output format looks like. For a search tool with no annotations, this is a significant gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two concise sentences with zero waste. It front-loads the core purpose and follows with the outcome, making it easy to parse quickly.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's moderate complexity (search with filtering), 100% schema coverage but no annotations and no output schema, the description is minimally adequate. It covers the what and why but lacks details on behavior, output, or integration with siblings like 'use_template'. It meets basic needs but has clear gaps in completeness.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all three parameters (query, category, limit) with clear descriptions. The description adds no additional parameter semantics beyond what's in the schema, such as search ranking or category examples beyond those implied. Baseline 3 is appropriate when schema does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action ('Browse and search') and resource ('PromptingBox public template library'), and distinguishes it from sibling tools like 'search_prompts' by specifying it searches templates rather than prompts. However, it doesn't explicitly contrast with 'use_template', which might be a related action.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage for finding pre-built prompts to save, but doesn't explicitly state when to use this versus alternatives like 'search_prompts' or 'use_template'. It provides some context (public template library) but lacks clear exclusions or comparison to siblings.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
toggle_favoriteC
Star or unstar a prompt in PromptingBox.
| Name | Required | Description | Default |
|---|---|---|---|
| promptId | No | The prompt ID. Provide this or promptTitle. | |
| promptTitle | No | The prompt title to search for. Provide this or promptId. | |
| isFavorite | Yes | true to favorite, false to unfavorite |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It states the action ('star or unstar') but doesn't cover critical aspects like whether this requires authentication, if it's idempotent, what happens on errors (e.g., invalid promptId), or the expected response format. For a mutation tool with zero annotation coverage, this is a significant gap in transparency.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, efficient sentence that front-loads the core action ('star or unstar') and resource ('a prompt in PromptingBox'). There is no wasted language, repetition, or unnecessary elaboration, making it highly concise and well-structured for quick comprehension.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's mutation nature (toggling favorites), lack of annotations, and absence of an output schema, the description is incomplete. It doesn't address behavioral traits (e.g., side effects, error handling), usage context, or return values, leaving significant gaps for an agent to operate effectively in a real-world scenario.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema description coverage is 100%, with clear documentation for all three parameters (promptId, promptTitle, isFavorite). The description adds no additional parameter semantics beyond what the schema provides, such as explaining the priority between promptId and promptTitle or formatting requirements. Given the high schema coverage, a baseline score of 3 is appropriate as the schema does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action ('star or unstar') and the resource ('a prompt in PromptingBox'), making the purpose immediately understandable. However, it doesn't explicitly differentiate this tool from sibling tools like 'save_prompt' or 'update_prompt' which might also involve prompt modifications, missing the opportunity to clarify its unique role in managing favorites.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use this tool versus alternatives. It doesn't mention prerequisites (e.g., needing an existing prompt), exclusions (e.g., not for templates), or how it relates to sibling tools like 'list_prompts' for viewing favorites. This lack of context leaves the agent to infer usage scenarios independently.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
update_promptA
Update the title and/or content of an existing prompt. If content changes, a new version is automatically created.
| Name | Required | Description | Default |
|---|---|---|---|
| promptId | No | The prompt ID. Provide this or promptTitle. | |
| promptTitle | No | The prompt title to search for. Provide this or promptId. | |
| title | No | New title for the prompt | |
| content | No | New content for the prompt |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries full burden. It discloses key behavioral traits: that updates modify existing prompts, that content changes trigger automatic version creation, and that both title and content are optional updates. However, it doesn't mention permission requirements, error conditions, or what happens when neither promptId nor promptTitle is provided.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two concise sentences with zero waste. The first sentence states the core functionality, the second adds crucial behavioral context about versioning. Every word earns its place and the information is front-loaded appropriately.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a mutation tool with no annotations and no output schema, the description provides adequate but incomplete context. It covers the core update behavior and versioning implication, but lacks information about return values, error handling, authentication requirements, and how the promptId/promptTitle alternative selection works in practice.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all 4 parameters thoroughly. The description adds marginal value by clarifying that title and content updates are optional ('and/or') and that content changes trigger versioning, but doesn't provide additional parameter semantics beyond what's in the schema descriptions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the verb 'update' and the resource 'existing prompt' with specific fields 'title and/or content'. It distinguishes from create operations but doesn't explicitly differentiate from sibling tools like 'save_prompt' or 'restore_version' which might have overlapping functionality.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage for modifying existing prompts, but doesn't explicitly state when to use this versus alternatives like 'save_prompt' or 'restore_version'. It mentions automatic version creation for content changes, which provides some contextual guidance but lacks explicit when/when-not directives.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
use_templateA
Save a public template to your PromptingBox collection. Creates a copy you can edit and customize.
| Name | Required | Description | Default |
|---|---|---|---|
| templateId | Yes | The template ID (from search_templates) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It discloses that the tool creates an editable copy, implying a write operation, but does not mention permissions, rate limits, or error conditions. This leaves gaps in behavioral understanding, though the core action is clear.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences with zero waste: the first states the action and destination, the second clarifies the outcome. It is front-loaded and appropriately sized for the tool's complexity.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given no annotations and no output schema, the description adequately covers the purpose but lacks details on behavioral aspects like permissions or return values. It is complete enough for a simple tool but could be more informative for reliable agent use.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents the 'templateId' parameter. The description does not add any meaning beyond what the schema provides (e.g., it doesn't explain format or sourcing details), resulting in the baseline score for high schema coverage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the specific action ('Save a public template') and resource ('to your PromptingBox collection'), distinguishing it from siblings like 'search_templates' (which finds templates) or 'save_prompt' (which saves custom prompts). It explicitly mentions creating an editable copy, which clarifies the outcome beyond just saving.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides clear context for when to use this tool: to save and customize public templates. However, it does not explicitly state when not to use it or name alternatives (e.g., using 'save_prompt' for custom prompts instead of templates), which prevents a perfect score.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
whoamiA
Show which PromptingBox account is connected to this MCP server.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries full burden. It discloses the tool's behavior (shows connected account information), but doesn't mention whether this requires authentication, what format the output returns, or any rate limits. The description is accurate but lacks detailed behavioral context that would be helpful for an agent.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, efficient sentence that directly states the tool's purpose without any wasted words. It's appropriately sized for a simple, parameterless tool and is perfectly front-loaded with the essential information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's simplicity (no parameters, no annotations, no output schema), the description is adequate but minimal. It explains what the tool does, but doesn't provide information about the return format or any authentication requirements. For a connectivity verification tool, more detail about the response structure would be beneficial.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters, and schema description coverage is 100% (empty schema is fully described). The description appropriately doesn't discuss parameters since none exist. Baseline for zero parameters is 4, as there's no need for parameter explanation.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose with a specific verb ('Show') and resource ('which PromptingBox account is connected'), making it immediately understandable. It distinguishes itself from all sibling tools, which focus on prompt/folder/tag management rather than account identification.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage context (determining account connectivity to the MCP server), but does not explicitly state when to use it versus alternatives or provide exclusions. It's clear this tool is for authentication/connection verification, but lacks explicit guidance on when it's necessary versus other tools.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
20 tool updates
v0.4.1- First observed
add_tags - First observed
create_folder - First observed
delete_folder - First observed
delete_prompt - First observed
delete_tag - First observed
duplicate_prompt - First observed
get_prompt - First observed
list_folders - First observed
list_prompts - First observed
list_tags - First observed
list_versions - First observed
move_prompt_to_folder - First observed
restore_version - First observed
save_prompt - First observed
search_prompts - First observed
search_templates - First observed
toggle_favorite - First observed
update_prompt - First observed
use_template - First observed
whoami
TDQS
Scored across 20 tools
Each tool has a clearly distinct purpose targeting specific resources and actions, such as add_tags for tag management, create_folder for folder creation, and get_prompt for retrieving prompt details. There is no significant overlap, making it easy for an agent to differentiate between operations like delete_prompt (permanent deletion) and move_prompt_to_folder (relocation).
Tool names consistently follow a verb_noun pattern throughout, such as add_tags, create_folder, delete_prompt, and list_folders. This uniform naming convention enhances readability and predictability, with no deviations or mixed styles like camelCase or inconsistent verb usage.
With 20 tools, the count is slightly high but reasonable for a comprehensive prompt management system, covering operations for prompts, folders, tags, versions, and templates. It avoids being excessive (e.g., over 25) and ensures each tool serves a specific function without redundancy.
The toolset provides complete CRUD and lifecycle coverage for prompt management, including creation (save_prompt, use_template), retrieval (get_prompt, list_prompts), updating (update_prompt, add_tags), deletion (delete_prompt, delete_folder), and additional features like version control (list_versions, restore_version) and organization (move_prompt_to_folder, toggle_favorite). No obvious gaps exist for the domain.
Maintenance
Related MCP Connectors
Self-hosted AI prompt library: prompts, collections, tags, teams, chains. 29 MCP tools for agents.
- PromptOTOAuthcom.promptot
Manage, version, and publish LLM prompts with blocks, variables, and evaluations.
Persistent AI memory shared across Claude, ChatGPT, coding agents, and compatible MCP clients.
- mcpOAuthai.butlerbrain
Persistent memory for AI assistants. Save once; recall from Claude, ChatGPT, or any MCP client.
Related MCP Servers
- AlicenseNot gradedqualityCmaintenanceGit-driven MCP server that manages and provides prompt templates with Handlebars support, enabling teams to collaborate on and version-control reusable, dynamic prompts across AI editors like Cursor and Claude Desktop.ISC
- AlicenseNot gradedqualityCmaintenanceEnables users to organize, search, and manage a shared library of prompts across AI tools via the Model Context Protocol. It supports hierarchical folder organization, tagging, and template variable substitution for dynamic prompt generation.MIT
- AlicenseNot gradedqualityDmaintenanceA local-only MCP server that uses SQLite to store, search, and manage a personal library of AI prompts. It enables developers to organize and reuse prompts across multiple AI clients like Claude and Cursor while keeping all data on their local machine.3 npm1MIT
- AlicenseBqualityDmaintenanceEnables AI models like Claude to manage local prompt files with CRUD operations, fuzzy search, categorization, templates, version control, and favorites.325MIT