SuperSkill
Manage a project-scoped AI knowledge vault and coding-agent workflows: capture decisions/learnings/tasks/sessions, build context and graphs, and route specialized skills.
Vault IO: read, write/append/prepend, and search vault files (full-text and frontmatter filters), scoped to a project slug/path.
Project context: get or auto-detect project context, and generate a draft context.md from a git repo without writing.
Decisions & learnings: log ADRs/decisions, capture or extract insights into individual files, deprecate items, and list learnings with tags/confidence/source.
Task management: add/list/update/kanban-board tasks with status, priority, blockers, tags, and assignees.
Session coordination: register/heartbeat/complete agent sessions, list active sessions, persist outcomes, and get resume context with recent work and next steps.
Graph & linking: create wikilinks, find related notes (outgoing/backlinks, multi-hop), and search across all projects grouped by project.
Environment & repo state: snapshot git branch/dirty files/last commit, store non-secret env facts, and store pointers to credential documentation.
Rollback safety: add/list rollback checkpoints with commit hash, purpose, scope, and mark follow-up work.
Skills: install/list/remove skills from GitHub repos, and route a task to the best skill via
superskill.Init & maintenance: initialize the vault for the current repo, generate templates (ADR/PRD/spec/etc.), prune stale content with dry-run/archive/delete, and view stats.
Allows scanning a git repository to generate project context, snapshot repository state (branch, dirty files, last commit), and manage rollback checkpoints using commit hashes.
Integrates with Obsidian vaults by storing project knowledge as Markdown and generating a knowledge-graph.canvas file for visual exploration of the project graph.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@SuperSkillscore and load the best skill for testing my Go API"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
SuperSkill
One-entry orchestrator for coding agents. Curated packs, project-jailed vault, FTS knowledge graph. Not a dump of 90k remote skills.
Prompt normally. Call superskill with the task. It diagnoses, then delegates (Go → Go pack, QA → QA, security bug → review + security). Defaults: ADHD-shaped output, careful-minimal, algorithm-correct, systems thinking. Review walks 18 axes; open branches grill the human (HITL). Factory packs (plan, TDD, verify, SRE, grill) are ours.
Requires Node 22+ (node:sqlite).
Why
Marketplace skill dumps fill the window with advice that does not know this system. Review and security need:
Vertical — this project’s ADRs, learnings, sessions
Horizontal — callers, sibling paths, co-activations
Fresh — session complete +
learn
SuperSkill injects a system brief from the project graph, jails vault IO to projects/<this-slug>/, and keeps .superskill/ gitignored so each developer’s trajectory stays private.
Related MCP server: skill-curator-mcp
How it works
npx superskill-cli init— detect stack, index the in-repo catalog (not skills.sh)Describe the task
Inverted-index router picks packs (language, phase, specialists)
Content is budgeted. Real review/diff/audit (and security bugs) get the vault + caller protocol
Activations write
.superskill/graph.json(local only)
Packs
Pack | When |
| Always. Diagnose, then load a combination — not a 10-step ritual |
| Always. Invariant, O(…), HLD/LLD that pay rent |
| Always (tiny). This project’s vault only |
| Tiny always-on. Full |
| This repo’s stack, or the task names a language |
| Review / diff / security fix — 18 axes |
| Spec, TDD, evidence-before-done, debug, Chrome QA |
| Deploy / SLO / incident — the cloud this repo already uses |
skills.sh remains opt-in install, not the default catalog.
Knowledge graph
Markdown under projects/<slug>/ is source of truth. SQLite FTS5 + edges is a derived index (porter stems: authorize hits Authorization).
npx superskill-cli graph rebuild -p my-project
npx superskill-cli graph viz -p my-project
npx superskill-cli qa viz -p my-projectOpen one file: projects/<slug>/knowledge-graph.html
Tabs: Graph · HLA · LLA · ERD · Modules. Keys g h l e m. Vault / index panels list what is actually stored (titles + text), not empty blobs. Obsidian: vault root = VAULT_PATH, open knowledge-graph.canvas.
QA drives system Chrome via playwright-core (in-harness, not a plugin).
Isolation
Vault IO jailed to
projects/<slug>/. Sibling projects deniedSecret-like writes throw
SECRET_REJECTEDSearch stays in this slug
.superskill/is appended to.gitignoreoninit
Quick start
npm install -g superskill
# in the repo
npx superskill-cli initMCP (prefer the installed binary, not npx -y, so you get this version):
{
"mcpServers": {
"superskill": {
"command": "superskill",
"env": { "VAULT_PATH": "~/Vaults/ai" }
}
}
}Claude Code plugin: /plugin marketplace add permanu/superskill then /plugin install superskill.
Then prompt normally and call the superskill tool with the task.
Configuration
Variable | Default | Description |
|
| Knowledge vault (must be under |
|
| Max tokens for context injection |
|
| Session heartbeat TTL |
| macOS Chrome | Browser used by |
npm publish is manual (npm publish on a maintainer machine). CI does not ship tokens.
License
AGPL-3.0-or-later — LICENSE
Copyright 2026 Permanu (Atharva Pandey)
Available Tools
33 toolsbrainstormCDestructive
Start or continue a brainstorm document for a project.
| Name | Required | Description | Default |
|---|---|---|---|
| topic | Yes | Brainstorm topic (used as filename) | |
| content | Yes | Content to add | |
| project | No | Project slug (auto-detected if omitted) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations include destructiveHint=true, and the description adds only the vague 'start or continue' without explaining whether existing content is overwritten or appended. It does not disclose side effects, file naming behavior, or project context requirements, so the description adds minimal value beyond the annotation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, front-loaded sentence that is easy to parse. It is concise but perhaps too sparse to fully illuminate behavior, which is why it doesn't earn a 5.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool is simple with few parameters and no output schema, so the description covers the basic action. However, it lacks context about when to use vs. siblings, what 'continue' implies, and whether a project context is required, leaving gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so parameters (topic, content, project) are fully documented in the schema. The description does not add any additional meaning or clarify relationships between parameters, so baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses the specific verbs 'start or continue' with the resource 'brainstorm document,' clearly indicating the tool manages brainstorm documents. However, it does not distinguish itself from sibling tools like write or capture, which may also handle document creation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description offers no guidance on when to use this tool versus alternatives like write or capture, nor does it state any prerequisites (e.g., project must exist). No exclusions or alternative references are provided.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
captureADestructive
Batch-capture multiple insights from a conversation into individual vault items. Each item gets its own file with auto-numbering. Supports any content type (learning, decision, adr, prd, research, etc.).
| Name | Required | Description | Default |
|---|---|---|---|
| items | Yes | Items to capture from the conversation | |
| project | No | Project slug (auto-detected if omitted) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description adds meaningful behavioral context beyond the destructiveHint annotation: each item gets its own file with auto-numbering, and any content type is supported. It does not detail what might be destroyed, but the annotation covers the destructive nature, and the file-creation behavior is useful context.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, front-loaded with the core action, and every phrase earns its place. No redundant or vague wording.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a two-parameter tool with full schema descriptions and an output-free schema, the description covers the main purpose, behavior, and content type flexibility. Minor gaps like project parameter behavior or failure handling are not critical given the schema details, but the description is not exhaustive.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema already has 100% description coverage for both parameters. The description merely restates the type examples already present in the schema ('learning, decision, adr, prd, research') and adds no new parameter-specific meaning beyond what the schema provides.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('batch-capture') and resource ('insights from a conversation into individual vault items'), with clear scope (auto-numbering files, multiple content types). This distinguishes it from siblings like 'write' or 'learn' by emphasizing batch and multi-type support.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage when capturing multiple insights from a conversation, but does not explicitly state when to prefer this over alternatives like 'write' or 'learn', nor does it provide exclusions or when-not-to-use guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
cred_refsADestructive
Store pointers to where credentials are documented (not the credentials themselves). E.g., 'Django admin creds are in tests/live/test_all_endpoints.py'.
| Name | Required | Description | Default |
|---|---|---|---|
| name | No | Credential name (required for add) | |
| notes | No | Additional notes | |
| action | Yes | Action | |
| project | No | Project slug (auto-detected if omitted) | |
| location | No | Where the credential is documented (required for add) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true, so the bar is lower. The description adds important context that only references are stored, not the actual credentials. However, it does not disclose what destructive actions may occur (e.g., overwriting existing references), leaving some ambiguity.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One concise sentence plus an example. No filler, every word earns its place, and the structure is easy to scan.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple add/list tool with fully described schema parameters, the description is largely sufficient. It clearly explains the tool's purpose and gives a concrete example. Missing explicit usage guidelines and destructive behavior details keep it from being perfect.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so parameters are already well-documented in structured form. The description adds a concrete example but no additional parameter-level semantics beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose: 'Store pointers to where credentials are documented (not the credentials themselves).' It uses a specific verb ('store') and resource ('pointers to where credentials are documented'), and explicitly differentiates from storing the credentials themselves, distinguishing it from sibling write/read tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides an illustrative example ('Django admin creds are in tests/live/test_all_endpoints.py') that implies a use case, but it does not explicitly state when to use this tool versus alternatives, nor does it mention any exclusions or when-not-to-use scenarios.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
decideBDestructive
Log an architectural/design decision to the project's decisions directory.
| Name | Required | Description | Default |
|---|---|---|---|
| title | Yes | Decision title | |
| context | No | Why this decision was needed | |
| project | No | Project slug (auto-detected if omitted) | |
| decision | Yes | What was decided | |
| alternatives | No | Alternatives considered | |
| consequences | No | Known trade-offs |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The annotations include destructiveHint=true, but the description adds no behavioral context beyond 'log'. It does not disclose what could be destroyed, whether existing files are overwritten, or any side effects. Since the annotation already flags destructiveness, the description fails to add any extra transparency.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single concise sentence that front-loads the action and resource. Every word is meaningful, with no redundancy or wasted content.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple logging tool with complete schema and annotations, the description is minimally adequate. However, it lacks any mention of return values, file modifications, or consequences of the destructive hint, leaving some contextual gaps. It does not leverage the opportunity to explain expected outcomes beyond the action itself.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema has 100% coverage with descriptions for all six parameters, so the schema does the heavy lifting. The description itself adds no parameter-specific meaning, which aligns with the baseline score of 3 when schema coverage is high.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's specific action: "Log an architectural/design decision to the project's decisions directory." It names a specific verb (log), a resource (architectural/design decision), and a location (decisions directory), which distinguishes it from generic write tools among siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description does not provide any guidance on when to use this tool versus alternatives. It never mentions alternatives or excludes cases (e.g., 'use write instead if ...'). The purpose is implied by the name, but no explicit selection criteria are given.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
deprecateADestructive
Mark a vault item (ADR, learning, etc.) as deprecated with an optional reason.
| Name | Required | Description | Default |
|---|---|---|---|
| path | Yes | Relative path to the item in the vault | |
| reason | No | Why this item is being deprecated |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The annotation confirms destructiveHint: true, and the description does not contradict this. It adds minimal context beyond annotations—clarifying that the action is to 'mark' rather than delete, but it does not explain potential side effects, reversibility, or what 'deprecated' means behaviorally.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, complete sentence that is correctly front-loaded and contains no filler. Every word adds value, achieving high efficiency.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool has no output schema, and the description doesn't clarify return values or the practical consequences of deprecation (e.g., whether the item becomes hidden or read-only). While the schema covers parameters and annotations cover safety, the description could be more complete for a destructive tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, with both 'path' and 'reason' explicitly described. The description adds no new parameter semantics beyond what the schema already provides, so the baseline score of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action ('Mark a vault item... as deprecated') with a specific resource ('vault item') and provides an optional parameter ('with an optional reason'). It distinguishes itself from sibling tools like 'write' or 'prune' by specifying the deprecation intent.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies when to use the tool (when wanting to deprecate an item) but does not explicitly contrast with alternatives or state when not to use it. Given the sibling 'prune' exists, explicit guidance on choosing deprecate over prune would be valuable, but it is absent.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
env_factsADestructive
Store and query stable environment facts for a project (e.g., auth backend, env file locations, local URLs, required env vars). Not for secrets — use cred_refs for that.
| Name | Required | Description | Default |
|---|---|---|---|
| key | No | Fact key (required for add) | |
| value | No | Fact value (required for add) | |
| action | Yes | Action | |
| context | No | Additional context for the fact | |
| project | No | Project slug (auto-detected if omitted) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare destructiveHint=true, but the description does not explain any destructive behavior, such as whether adding an existing key overwrites it or if list operations have side effects. It adds example fact types but no behavioral details, leaving the agent without important operational nuance.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences with no filler. The first sentence states the core function, and the second provides a critical boundary with an alternative tool. Every word earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple add/list tool, the description plus schema is mostly sufficient. However, the lack of any output schema and the absence of note about overwrite or destructive effects means the agent may wonder what 'list' returns and what 'add' does to an existing key. This is a minor gap for an otherwise simple tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema has 100% description coverage for all 5 parameters, so the baseline is 3. The description's examples ('auth backend, env file locations, local URLs, required env vars') provide useful context for the value parameter, but it does not add param-specific syntax or clarify the relationship between key, value, and context beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses specific verbs 'Store and query' and identifies the resource as 'stable environment facts for a project,' making the tool's purpose immediately clear. It also distinguishes itself from the sibling 'cred_refs' tool by explicitly excluding secrets, which differentiates from a likely alternative.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly states when not to use the tool ('Not for secrets') and names the alternative ('use cred_refs for that'). This provides a clear exclusion and redirect, which is strong usage guidance. The primary use case is implied by the purpose statement.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
extractADestructive
Extract decisions, learnings, or other items from a source document into individual vault files. Each extracted item gets its own file with a backlink to the source.
| Name | Required | Description | Default |
|---|---|---|---|
| items | Yes | Items to extract from the source document | |
| source | Yes | Relative path to the source vault note | |
| project | No | Project slug (auto-detected if omitted) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description explains that each item gets its own file with a backlink, but does not clarify what happens to the source document or what destructive action is performed despite destructiveHint being true. It does not disclose whether the source is modified or items are removed, leaving significant ambiguity about the tool's effects.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two concise sentences, front-loaded with the primary action and followed by the key outcome. No filler or redundant information; every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description covers the core workflow and file creation but lacks clarity on destructive behavior and does not explain the impact on the source or the role of the 'project' parameter. Since there is no output schema and annotations only signal destructiveness, the description is not fully complete but is adequate for a basic understanding.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The description aligns with the schema's 'items' and 'source' parameters but does not add additional detail beyond what is already in the input schema, which has 100% coverage. It does not explain the 'project' parameter behavior or the 'type' field semantics beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool extracts items from a source document into individual vault files, specifying both the action and resource. It distinguishes from siblings like write or read by focusing on breaking out items from a source, which is a unique operation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage when wanting to break out decisions or learnings from a document, but it does not explicitly state when not to use it or compare with alternatives like write or learn. It provides no exclusions or alternative guidance, so usage is only implied.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
generate_contextARead-only
Scan a git repo and generate a draft context.md. Returns the draft — does NOT write to vault. Human reviews before committing.
| Name | Required | Description | Default |
|---|---|---|---|
| slug | No | Project slug (default: derived from directory name) | |
| project_path | Yes | Absolute path to the git repository to scan |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the readOnlyHint annotation, the description explains that it returns a draft and that human review is required before committing, clarifying the tool's non-persistent behavior. It does not contradict the annotation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single sentence of 19 words, front-loaded with the primary action ('Scan'), and every phrase contributes meaningful information with no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only draft generator with two parameters and a simple response, the description covers the core behavior, return value, and key constraint (no writing to vault). The absence of an output schema is mitigated by the clear statement that it returns a draft.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% for both parameters, so the schema already documents slug and project_path adequately. The description adds no additional parameter semantics, making a baseline score of 3 appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool scans a git repo and generates a draft context.md, with a specific verb and resource. It distinguishes itself from write tools by explicitly noting it does NOT write to the vault.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage for drafting context that requires human review and explicitly notes it does not write to the vault, but it does not name specific alternatives or provide explicit when-to-use/when-not-to-use guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
graph_cross_projectARead-only
Search the current project's vault only. Cross-project search is denied.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | Max results (default 20) | |
| query | Yes | Search query |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
readOnlyHint=true already signals a safe read operation, and the description adds a meaningful behavioral constraint: cross-project search is denied and only the current project's vault is searched. This goes beyond the annotations, though it doesn't detail error behavior or authentication requirements.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences with no filler. The core scope is front-loaded, and the denial of cross-project search is stated directly and economically.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a two-parameter read-only search tool, the description plus schema and annotations cover the essential information: scope, denial of cross-project access, and parameter definitions. No output schema exists, but search results are predictable enough that the description is sufficient.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the schema already documents 'query' and 'limit'. The description adds no parameter-specific detail beyond the general search action, so it meets the baseline but doesn't exceed it.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific action ('Search') and a bounded resource ('current project's vault only'), and explicitly denies cross-project scope. It doesn't name sibling tools, but the scope restriction clearly differentiates it from a cross-project search tool.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description clearly scopes usage to the current project's vault and states that cross-project search is denied, which provides an implicit when-not. However, it doesn't name alternatives like 'search' or 'graph_related' or explain when to prefer them.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
initADestructive
Initialize superskill for the current project. Detects stack, indexes the in-repo catalog, builds the knowledge graph. Does not scrape skills.sh.
| Name | Required | Description | Default |
|---|---|---|---|
| bridge | No | Enable native skill bridge (replaces native skill files with superskill redirects) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already flag destructiveHint=true; the description adds operational details (detects stack, indexes catalog, builds graph) and a useful negative boundary ('Does not scrape skills.sh'). It doesn't spell out that bridge replaces native skill files, though that appears in the parameter schema.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences with the main action front-loaded and the negative boundary expressed in a single clause. No filler or redundant restatement.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a one-optional-param tool with a destructive annotation and no output schema, the description covers the core workflow and an important non-behavior. It could add sibling routing or side-effect detail, but nothing essential is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the boolean bridge parameter is already fully documented. The tool description adds no parameter-level meaning beyond what the schema provides.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a clear verb ('Initialize'), resource ('superskill'), and scope ('current project'), then enumerates concrete steps: detects stack, indexes the in-repo catalog, builds the knowledge graph. It doesn't explicitly contrast with siblings like knowledge_rebuild, but the action is unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Implies first-time setup for the current project and explicitly rules out scraping skills.sh, but gives no when-to-use guidance versus alternatives such as knowledge_rebuild or generate_context, and no prerequisites.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
knowledge_rebuildA
Rebuild this project's SQLite FTS5 + edges index from markdown. Source of truth stays the files.
| Name | Required | Description | Default |
|---|---|---|---|
| project | No | Project slug (auto-detected if omitted) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations provide only destructiveHint=false, and the description adds that the source of truth stays the files, implying the operation does not modify source files. However, it does not disclose other behaviors such as potential runtime cost, whether the old index is replaced atomically, or any side effects. It is adequate but not rich.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two concise sentences with no filler. The action and core constraint (source of truth stays files) are front-loaded, making the intent immediately clear.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple rebuild operation with one optional parameter and no output schema, the description is sufficient for an agent to understand the purpose and the invariant that source files are unchanged. It does not detail post-conditions or prerequisites, but these are not critical for correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, with the 'project' parameter described as 'Project slug (auto-detected if omitted)'. The tool description adds no additional parameter meaning, meeting the baseline for fully covered schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action (rebuild), the resource (project's SQLite FTS5 + edges index), and the source (markdown). It also distinguishes itself from siblings like 'search' and 'graph_related' by focusing on index regeneration rather than querying.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives no explicit guidance on when to use this tool versus alternatives. It implies usage when the index needs rebuilding, but does not state exclusions or conditions. The phrase 'Source of truth stays the files' hints that files are authoritative, but this is not actionable guidance for selection.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
knowledge_vizB
Write knowledge-graph.html (browser) and knowledge-graph.canvas (open in Obsidian). Same edges as the FTS index.
| Name | Required | Description | Default |
|---|---|---|---|
| project | No | Project slug (auto-detected if omitted) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only provide destructiveHint: false, and the description adds that it writes two files and specifies their intended use (browser and Obsidian). This gives some context about output artifacts. However, it does not disclose potential side effects such as file overwriting, required directory, or whether it can be called repeatedly. The description relies on the annotation for safety but doesn't add much beyond what the annotation already implies (non-destructive). The note about 'same edges as FTS index' adds content context but not behavioral detail.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, efficient sentence that front-loads the primary action (writing two files) and includes the key detail about content equivalence to the FTS index. There is no redundant phrasing or filler. Every element earns its place, making it highly concise and well-structured.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple tool with one optional parameter and no output schema, the description covers the outputs and content source. However, it omits the output location (e.g., project root vs. current directory) and does not explain how the 'project' parameter influences the generated files. It also does not position the tool among its many siblings. These gaps mean an agent might not fully understand when and how to use it correctly, though the low complexity keeps the deficiency from being severe.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema has 100% coverage for the single parameter 'project', with a description stating it is auto-detected if omitted. The tool description does not add any further detail about how the parameter affects the output or its format. Since the schema already documents the parameter, the description adds no additional value, matching the baseline expectation for high coverage situations.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose: it writes two specific output files (knowledge-graph.html and knowledge-graph.canvas) and relates the content to the FTS index. It is not a tautology and uses specific verbs and resources. However, it does not explicitly differentiate from sibling visualization tools like graph_related or qa_viz, relying instead on the unique output filenames to signal distinction.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to use this tool versus the siblings. It does not mention exclusions, alternatives, or appropriate contexts. An agent is left to infer that this tool is for generating a specific visualization, but it lacks any 'instead of X' or 'use this when Y' information. This is a significant gap given the large sibling set.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
learnADestructive
Capture and query learnings. Learnings persist discoveries across sessions. Stored as individual files in projects//learnings/.
| Name | Required | Description | Default |
|---|---|---|---|
| tag | No | Filter learnings by tag (for list) | |
| tags | No | Tags | |
| title | No | Learning title (required for add) | |
| action | Yes | Learning action | |
| source | No | Source tool | |
| project | No | Project slug (auto-detected if omitted) | |
| discovery | No | What was discovered (required for add) | |
| confidence | No | Confidence level (default medium) | |
| session_id | No | Session ID that captured this learning |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The annotations declare destructiveHint=true, which the description does not contradict or explain. The description adds context about persistence and file storage, but it doesn't disclose any potentially destructive behavior (e.g., overwriting existing learnings) that might align with the annotation. With annotations present, the bar is lower, but the description could elaborate on side effects.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is three short, front-loaded sentences. The first sentence states the purpose, the second explains persistence, and the third provides storage location. Every sentence earns its place without wasted words.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description should explain what the tool returns or any side effects beyond storing files. It covers the core function and storage but doesn't describe list results, add confirmations, or how destructiveHint applies. Given the tool has 9 parameters and two actions, the description is somewhat incomplete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema covers 100% of parameters, so the description doesn't need to repeat parameter details. It doesn't add meaning beyond the schema, such as clarifying the relationship between 'add' and 'list' actions or how 'project' auto-detection works. Baseline 3 is appropriate given full schema coverage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb+resource pair ('Capture and query learnings') and clearly distinguishes the tool from siblings by focusing on learnings. It also provides storage details ('Stored as individual files in projects/<slug>/learnings/'), which adds concrete scope and context beyond the tool name.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear context for when to use the tool: to persist discoveries across sessions. It doesn't explicitly mention alternatives or exclusions, but the phrase 'persist discoveries across sessions' conveys the intended usage scenario. This is more than implied usage but lacks explicit comparisons to sibling tools like 'capture' or 'read'.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
linkADestructive
Create a forward link between two vault notes. Appends a [[wikilink]] to the source note, enabling graph_related to discover the connection.
| Name | Required | Description | Default |
|---|---|---|---|
| source | Yes | Relative path to the source vault note | |
| target | Yes | Relative path (or note name) to link to | |
| project | No | Project slug (auto-detected if omitted) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true, and the description adds that it appends a [[wikilink]] to the source note, which is useful context. However, it does not disclose edge cases such as what happens if the source note does not exist, whether the link is appended at the end, or whether repeated links are allowed, so the added transparency is modest.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is exactly two sentences, uses active voice, and includes no filler. The key action and effect are front-loaded, making it easy for an agent to parse quickly.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's simplicity (3 parameters, no output schema, annotations present), the description covers the core behavior and its relationship to graph_related. It omits some potential side effects, but the schema and annotations fill in the rest, so it is sufficiently complete for an agent to use it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%: each parameter has a description (source, target, project). The description's mention of 'source note' and 'target note' maps to the schema but does not add new meaning beyond what the schema already provides. Thus a baseline score of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action ('Create a forward link') and the resource ('between two vault notes'), and elaborates by explaining it appends a [[wikilink]] to the source note. This distinguishes it from sibling tools like graph_related (which discovers connections) and write (which writes general content).
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage by noting that it 'enables graph_related to discover the connection', but it does not explicitly say when to use this tool versus alternatives like write or search. No exclusions or alternative guidance are provided, so the context is clear but not fully prescriptive.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
project_contextARead-only
Get the context document for a project. Auto-detects project from CWD if not specified.
| Name | Required | Description | Default |
|---|---|---|---|
| project | No | Project slug. If omitted, auto-detected from CWD. | |
| detail_level | No | Detail level (default summary) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description adds the behavioral trait of auto-detecting the project from CWD, which goes beyond the readOnlyHint annotation. It is consistent with the read-only annotation and clearly conveys that it is a read operation. No side effects are implied, which aligns with the annotation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, front-loaded sentence that conveys the essential information without fluff. Every word earns its place, making it highly concise and well-structured.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the simplicity of the tool, the description is complete enough. It covers the primary action, the auto-detection behavior, and the schema fully documents parameters. However, without an output schema, it does not describe the return format, though the detail_level parameter implies summary vs. full document.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3. The description adds no new parameter semantics beyond what the schema already provides; it only restates the auto-detection behavior already in the project parameter description.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action ('Get') and the resource ('the context document for a project'). It is specific and distinguishes itself from siblings like 'generate_context' (which would create) and 'read' (which is more general) by focusing on project context documents.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage: use it to obtain a project's context document. It provides a useful detail (auto-detection from CWD) but does not explicitly state when to use this tool over alternatives or any exclusions. It lacks a clear 'when not to use' or mention of sibling tools.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
pruneADestructive
Archive or delete stale vault content based on retention policies. Use mode='dry-run' first to preview.
| Name | Required | Description | Default |
|---|---|---|---|
| all | No | Prune all projects | |
| mode | No | Prune mode (default dry-run) | |
| project | No | Project slug (or omit for auto-detect) | |
| sessions_days | No | Session retention in days (default 30) | |
| done_tasks_days | No | Done task retention in days (default 30) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description discloses that the tool archives or deletes content, aligning with the 'destructiveHint' annotation. It adds value by recommending a dry-run first to preview, which is a behavioral warning beyond the annotation. It doesn't detail reversibility or permissions, but the guidance is helpful.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, front-loaded sentence that conveys purpose and essential guidance without extra words. It is clear, direct, and earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's moderate complexity (five optional parameters, no output schema), the description covers the core behavior and safety protocol. It doesn't explain archive semantics or return values, but the schema fills in parameter meanings, so the description is reasonably complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The description does not add any parameter-specific meaning beyond the input schema, which already has 100% coverage. Per the rubric, high schema coverage gives a baseline of 3, and the reference to 'mode=\'dry-run\'' is a usage tip rather than new parameter semantics.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's function: 'Archive or delete stale vault content based on retention policies.' This provides a specific verb and resource, and distinguishes it from siblings like 'deprecate' or 'write' by focusing on pruning stale content.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear usage context ('stale vault content based on retention policies') and a specific safety instruction ('Use mode=\'dry-run\' first to preview'). It does not explicitly discuss alternatives or when not to use, but the context is clear enough for a 4.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
qa_vizCRead-only
Regenerate the knowledge graph HTML and drive system Chrome to click nodes and read the panel. In-harness QA, not a plugin.
| Name | Required | Description | Default |
|---|---|---|---|
| project | No | Project slug (auto-detected if omitted) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description says 'Regenerate the knowledge graph HTML,' which implies a write or state-changing operation, while annotations declare readOnlyHint=true. This is a direct contradiction. It also does not disclose side effects of driving Chrome or what 'read the panel' returns.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and front-loaded with the primary action. The second sentence adds useful context about being in-harness QA rather than a plugin, though it is not strictly necessary.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
This is a relatively complex tool involving HTML regeneration and browser automation, yet the description omits what the tool returns, whether it launches Chrome, and any side-effect or safety details. The contradictory read-only annotation makes the picture incomplete for an agent deciding whether to invoke it.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% and the only parameter, project, already has a clear description ('Project slug (auto-detected if omitted)'). The tool description adds no parameter-level meaning beyond what the schema provides, so the baseline of 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific action and resource: regenerating the knowledge graph HTML and driving Chrome to click nodes and read the panel. It is clear enough to understand the tool's job, though it does not explicitly differentiate it from sibling tools like knowledge_viz or graph_related.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to use this tool versus alternatives such as knowledge_viz or graph_related. The phrase 'In-harness QA, not a plugin' hints at context but does not explain selection criteria or exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
readARead-only
Read a file or directory listing from the AI knowledge vault.
| Name | Required | Description | Default |
|---|---|---|---|
| path | Yes | Relative path within vault. Use '.' for root. | |
| depth | No | Directory listing depth (default 1) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The readOnlyHint annotation already declares the tool's read-only nature. The description adds useful context about operating within the AI knowledge vault and reading files or directory listings, but it does not disclose error handling, return format, or limits, which are left unspecified.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single concise sentence that immediately conveys the tool's purpose. It contains no filler or redundant information, making it highly efficient for an agent to parse.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's simplicity, 100% schema coverage, and readOnlyHint annotation, the description covers the core functionality adequately. It does not have an output schema, but the phrase 'read a file or directory listing' implies the return content. Some detail on directory listing formatting or depth behavior is missing, but it is sufficient for a basic read tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema fully describes both parameters (path and depth) with clear descriptions, including the use of '.' for root and the default depth of 1. The description adds no additional parameter semantics beyond what the schema already provides, so the baseline of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states 'Read a file or directory listing from the AI knowledge vault,' providing a specific verb and resource. It distinguishes itself from sibling tools like write and search by focusing on direct reading of vault content.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is provided on when to use this tool versus alternatives such as search or project_context. There are no explicit exclusions or scenarios mentioned, leaving the agent to infer usage solely from the tool name and description.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
resumeARead-only
Get resume context for continuing work — recent sessions, interrupted work, in-progress tasks, suggested next steps. Call this at session start to understand what happened before.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | Number of recent sessions to show (default 5) | |
| format | No | Output format (default markdown) | |
| project | No | Project slug (auto-detected if omitted) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With readOnlyHint=true already covering non-destructiveness, the description adds behavioral context by enumerating the types of information returned (recent sessions, interrupted work, in-progress tasks, next steps). It does not disclose potential staleness or data limitations, but this is a minor gap given the simple read-only nature.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two concise sentences: the first states the purpose and content, the second gives the usage context. No redundant words, front-loaded with the most important information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description sufficiently covers what the tool returns and when to call it. Given the simple 3-optional-parameter schema and readOnly annotation, it does not need to explain return values in detail. However, it could have briefly mentioned output format options, though these are in the schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% with each parameter (limit, format, project) clearly described. The tool description adds no additional meaning about parameters, so baseline 3 is appropriate since the schema does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Get') and resource ('resume context'), then elaborates with concrete content: recent sessions, interrupted work, in-progress tasks, and suggested next steps. This clearly distinguishes it from sibling tools like 'project_context' or 'status' by framing it as a session-start continuity tool.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly says 'Call this at session start to understand what happened before', providing a clear when-to-use instruction. It does not mention when not to use it or name alternative tools, but the context is strong enough to guide an agent.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
rollbackADestructive
Manage rollback checkpoints. Store commit hashes with purpose/scope so rollback is safer. Mark when follow-up work starts after a checkpoint.
| Name | Required | Description | Default |
|---|---|---|---|
| scope | No | Scope of changes in the checkpoint | |
| action | Yes | Action | |
| project | No | Project slug (auto-detected if omitted) | |
| purpose | No | Purpose of the checkpoint (required for add) | |
| commit_hash | No | Commit hash (required for add) | |
| checkpoint_id | No | Checkpoint ID to mark follow-up for |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The destructiveHint annotation already signals destructive potential. The description adds context about storing hashes and marking follow-ups but doesn't disclose specific destructive behaviors (e.g., deleting checkpoints). No contradiction with annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two concise sentences that front-load the core purpose ('Manage rollback checkpoints') and provide enough detail on sub-actions without unnecessary fluff.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description is adequate for a checkpoint manager given the schema and annotations, but it lacks detail on action-specific behavior (e.g., what 'list' returns) or side effects. No output schema exists, so slight gaps remain.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% with all parameters described. The description repeats purpose/scope/commit hash but doesn't add new semantic meaning beyond the schema. Baseline of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose: 'Manage rollback checkpoints' with specific sub-tasks like storing commit hashes and marking follow-ups. This distinguishes it from sibling tools like snapshot_repo_state by focusing on rollback safety management.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied ('so rollback is safer') but there's no explicit guidance on when to use this tool vs alternatives like snapshot_repo_state. No exclusions or alternative recommendations are provided.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
searchARead-only
Full-text search (SQLite FTS5 porter stemming) scoped to this project. Structured mode still matches frontmatter.
| Name | Required | Description | Default |
|---|---|---|---|
| mode | No | Search mode (default text) | |
| limit | No | Max results (default 10) | |
| query | Yes | Search query or structured filter (e.g. 'type:adr project:permanu') | |
| project | No | Project slug to scope results | |
| path_filter | No | Glob to restrict search scope |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare readOnlyHint=true, so the read-only nature is covered. The description adds valuable behavioral detail: porter stemming and that structured mode still matches frontmatter. These go beyond the annotation and help the agent anticipate matching behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences with zero filler. Core purpose is front-loaded, and the technical detail (porter stemming, frontmatter) is packed efficiently. No redundancy with schema or annotations.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only search tool with 5 parameters and no output schema, the description covers the essential functionality and the key mode nuance. It doesn't describe result format or pagination, but the schema already documents limit and query format. Slightly more context on what 'structured' returns would improve completeness, but it's adequate.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% with all five parameters described. The description adds extra meaning by clarifying that 'structured' mode handles frontmatter queries, which enriches the mode parameter conceptually. It does not replicate schema descriptions but supplements them.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Description clearly states it performs full-text search scoped to the current project, specifying the underlying mechanism (SQLite FTS5 porter stemming) and distinguishing from other tools. No sibling tool offers search, so it stands out unambiguously.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance on when to use this tool versus alternatives. It mentions 'scoped to this project' but doesn't mention any exclusion criteria or alternative tools. The distinction between 'text' and 'structured' modes is also unexplained in terms of when to pick each.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
sessionADestructiveIdempotent
Register, update, or query active agent sessions for multi-agent coordination. On complete, persists a session note to the vault.
| Name | Required | Description | Default |
|---|---|---|---|
| tool | No | Tool name: claude-code|opencode|codex | |
| action | Yes | Session action | |
| blocked | No | Blocking issues (for complete) | |
| outcome | No | Session outcome (for complete) | |
| project | No | Project being worked on | |
| completed | No | Completed work items (for complete) | |
| session_id | No | Session ID (for heartbeat/complete) | |
| task_summary | No | What this session is doing | |
| files_touched | No | Files this session modifies | |
| tasks_completed | No | Task IDs completed (for complete) | |
| verification_run | No | Verification steps run and results (for complete) | |
| commands_to_resume | No | Commands to resume work (for complete) | |
| partially_completed | No | Partially completed work items (for complete) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations provide idempotentHint and destructiveHint, but the description adds a concrete side effect: 'On complete, persists a session note to the vault.' This goes beyond the annotations and helps the agent anticipate a persistent effect. It does not contradict any annotations and provides useful behavioral context.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences, front-loaded with the primary purpose, and every word contributes. It is concise without being under-specified, effectively summarizing the tool's core behavior and a key side effect.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool has 13 parameters but only 1 required, and the schema provides detailed descriptions for all. The description adds the important side effect of persisting a note to the vault, which is not in the schema. Given the high schema coverage and clear summary of actions, the description is sufficiently complete for an agent to understand the tool's role and main behavior.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the description does not need to compensate for undocumented parameters. The description mentions 'register, update, or query' which aligns with the 'action' enum, but it does not add meaning beyond what the schema already provides for individual parameters. Baseline of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's function: 'Register, update, or query active agent sessions for multi-agent coordination.' This includes a specific verb ('register, update, query') and resource ('active agent sessions'), making it unambiguous and distinct from sibling tools, none of which handle sessions.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives context ('for multi-agent coordination') indicating when it might be used, but does not explicitly state when to use it versus alternative tools, nor when not to use it. There is no mention of exclusions or preferred alternatives, leaving the agent to infer usage from the purpose.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
skill_installADestructive
Install skills from a GitHub repo (e.g. owner/repo or full URL). Clones the repo, discovers SKILL.md files, and copies them to the local skill directory.
| Name | Required | Description | Default |
|---|---|---|---|
| source | Yes | GitHub source: owner/repo, github:owner/repo, or full URL | |
| select_skills | No | Optional list of skill names to install (installs all if omitted) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations include destructiveHint: true, and the description adds context by detailing the clone/discover/copy steps. It does not contradict the annotation and provides a concrete picture of what the tool does, though it doesn't mention overwrite behavior or permissions.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences, no fluff. The first sentence states the core purpose and source options, the second explains the process. Every word earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool is simple (2 params, no output schema, annotation provided). The description covers the process, source format, and selection behavior. Minor gaps like error handling or overwrite semantics are not critical for an install tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, with both parameters already well-described. The description adds a bit of context (source format example) but does not significantly augment the schema meaning.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states a specific verb ('Install') and resource ('skills from a GitHub repo'), and also explains the mechanism (clones repo, discovers SKILL.md, copies to local directory). This distinguishes it from sibling tools like skill_remove or skill_list_installed.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides clear context for when to use this tool (installing skills from a GitHub repo) and gives source format examples. It does not explicitly mention exclusions or alternatives, but the context is unambiguous.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
skill_list_installedARead-only
List skills installed locally from GitHub repos.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The readOnlyHint annotation already establishes the tool is read-only. The description adds that it lists skills installed locally from GitHub repos, providing scoping but no additional behavioral details such as return format or network behavior. This is adequate but not rich.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, front-loaded sentence that states the action and resource without wasted words. It is perfectly concise for the tool's simplicity.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple list tool with no parameters, no output schema, and a read-only annotation, the description provides enough context to understand its purpose. There is no missing critical information.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
There are no parameters, and schema coverage is 100% by default. The description doesn't need to explain parameters. This baseline of 4 is appropriate for a zero-parameter tool.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's action (list) and resource (skills installed locally from GitHub repos). It distinguishes from sibling tools like skill_install and skill_remove, which perform different operations.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The context is clear: use this tool when you want to see which skills are installed locally from GitHub repos. There are no explicit exclusions or alternatives mentioned, but given the simplicity of the list operation, the context is sufficient.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
skill_removeADestructive
Remove an installed skill by name.
| Name | Required | Description | Default |
|---|---|---|---|
| name | Yes | Skill name to remove |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The destructiveHint annotation already flags the destructive nature. The description adds no further behavioral details, such as permanence, dependency effects, or potential side effects. It simply restates the action without enriching the agent's understanding of what happens when the tool is invoked.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, clear sentence with no redundancy. It is front-loaded and immediately conveys the intended action.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's simplicity (one parameter, no output schema), the description is adequate to explain the action. The destructiveHint annotation covers safety context. It lacks only minor details about edge cases or prerequisites, but for a straightforward removal operation, this is sufficient.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, and the parameter 'name' is described as 'Skill name to remove'. The tool description reiterates 'by name' without adding new meaning beyond the schema. Baseline 3 is appropriate since the schema does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description "Remove an installed skill by name" uses a specific verb (remove) and a clear resource (installed skill). It directly distinguishes from sibling tools like skill_install (adds a skill) and skill_list_installed (lists skills).
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies when to use the tool: when you want to remove a skill. However, it does not mention any alternatives (e.g., deprecate) or situations to avoid using it. No explicit exclusions or usage context beyond the obvious action.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
snapshot_repo_stateADestructive
Snapshot current git state (branch, dirty files, last commit) into the vault. Helps avoid repeating repo discovery across sessions.
| Name | Required | Description | Default |
|---|---|---|---|
| branch | No | Current git branch | |
| project | No | Project slug (auto-detected if omitted) | |
| dirty_files | No | List of dirty/uncommitted files | |
| last_commit | No | Last commit hash or message |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already flag the tool as destructive (destructiveHint: true), and the description does not add context about what destructive action occurs (e.g., overwriting existing snapshots). It describes the content being snapshotted, but no additional safety-relevant details beyond the annotation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two crisp sentences: the first states the action and inputs, the second explains the purpose. No filler or repetition, making it highly efficient and well-structured.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with no output schema, the description covers the core behavior and rationale well. However, it does not mention what the snapshot does to existing vault content (relevant given destructiveHint) or what the tool returns, leaving minor gaps in completeness.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so each parameter is already documented in the input schema. The description names the parameters (branch, dirty files, last commit) but adds no semantic meaning beyond what the schema provides, thus meeting the baseline of 3.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's function with a specific verb ('Snapshot') and resource ('current git state (branch, dirty files, last commit) into the vault'). It distinguishes itself from sibling tools by focusing on persisting state for later reuse rather than generic read/write/search operations.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides clear context for when to use the tool ('Helps avoid repeating repo discovery across sessions') but does not explicitly mention when not to use it or name alternative tools. This gives a strong sense of purpose without full exclusionary guidance, matching a 4.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
statsARead-only
Show content statistics for a project — file counts, task breakdown, growth monitoring.
| Name | Required | Description | Default |
|---|---|---|---|
| project | No | Project slug (auto-detected if omitted) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, so the safety profile is covered. The description adds useful context about the types of statistics returned, but does not disclose details such as whether the project must exist or how auto-detection works (that resides in the schema). With annotations present, this is adequate but not rich.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single sentence, immediately states the purpose, and is free of redundant phrasing. It is concise and well-structured.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple read-only tool with one optional parameter, the description lists the key output categories (file counts, task breakdown, growth monitoring) without needing an output schema. While it could mention the return format or more detail on growth monitoring, the current level is sufficient for selection and invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The only parameter 'project' is fully documented in the schema with a description and auto-detection note. The tool description adds no additional parameter-level semantics, but the schema carries the full burden, so a baseline score is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly identifies the tool as a statistics/analytics operation for projects, with a specific verb 'Show' and resource 'content statistics for a project'. The examples of file counts, task breakdown, and growth monitoring distinguish it from sibling tools like read, write, or status.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies this tool is for retrieving statistical summaries, but it does not explicitly state when to prefer this over sibling tools like status or project_context, nor does it mention any exclusions. Usage context is clear but alternative guidance is missing.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
statusARead-only
Show superskill knowledge graph state: loaded skills, weights, recent sessions.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, and the description aligns with this by using 'Show'. It adds context about what aspects of the state are visible (loaded skills, weights, recent sessions), which is valuable beyond the annotation. No contradictions.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single concise sentence with no filler. It front-loads the main purpose and immediately lists the key components, earning the maximum score for efficiency.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple zero-parameter read-only status tool, the description is sufficiently complete. It names the core content areas (loaded skills, weights, recent sessions). While it doesn't specify output format or provide deeper details, the simplicity of the tool makes this acceptable. Sibling tools like 'stats' exist, but the description distinguishes the knowledge graph focus.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
There are 0 parameters, so the baseline is 4. The description doesn't need to explain parameters but effectively communicates the scope of the operation by listing what will be shown.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool shows superskill knowledge graph state with specific components (loaded skills, weights, recent sessions). It uses a specific verb 'Show' and a resource, differentiating it from sibling tools like 'stats' or 'skill_list_installed' which likely focus on different aspects.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description clearly implies when to use this tool: when an agent needs to inspect the current state of the superskill knowledge graph. It does not explicitly mention alternatives or exclusions, but the context is clear enough for basic selection.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
superskillARead-only
Route a task to curated in-repo packs (code/review/security/ops/devops). Lazy: only matching language + phase. Does not scrape skills.sh.
| Name | Required | Description | Default |
|---|---|---|---|
| task | No | Describe what you're doing — superskill finds the right methodology and loads it. | |
| skill_id | No | Load a specific skill by ID (e.g. 'vercel-labs/agent-skills@react-best-practices') |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The readOnly annotation already signals safety; the description adds meaningful behavioral context by saying it is lazy and does not scrape skills.sh. It does not describe precedence when both task and skill_id are supplied, but this is minor.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two concise sentences front-load the core action and scope, then add two useful behavioral qualifiers. No filler or repetition of schema content.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple read-only routing tool with two optional params, the description plus schema covers purpose, behavior, and boundaries. It could be more explicit about expected output or selection precedence, but nothing essential is missing for calling it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents task and skill_id fully. The description's 'task' wording is consistent but adds no parameter detail beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific action ('Route a task') and a concrete resource ('curated in-repo packs') with domain categories. It also adds a distinguishing boundary ('Does not scrape skills.sh'), though it doesn't name a sibling tool to differentiate against.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage for tasks needing methodology and clarifies lazy matching ('only matching language + phase'), but it never states when to prefer this over siblings like search, decide, or skill_list_installed. No explicit when-not-to-use or alternatives.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
taskADestructive
Manage project tasks. Supports add, list, update, and board (kanban) views. Tasks are stored as individual files in projects//tasks/.
| Name | Required | Description | Default |
|---|---|---|---|
| tags | No | Tags | |
| title | No | Task title (required for add) | |
| action | Yes | Task action | |
| status | No | Task status | |
| project | No | Project slug (auto-detected if omitted) | |
| task_id | No | Task ID e.g. task-001 (required for update) | |
| priority | No | Task priority (default p1) | |
| blocked_by | No | Task IDs that block this task | |
| assigned_to | No | Assignee: claude-code|opencode|codex|human |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The annotation destructiveHint=true already signals potential side effects. The description adds that tasks are stored as individual files in projects/<slug>/tasks/, providing a persistence context. However, it doesn't detail what specific actions are destructive (e.g., update overwrites, delete removes) or any prerequisites.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is three short sentences, each informative: purpose, actions, and storage. There is no fluff or redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with 9 parameters and no output schema, the description is relatively sparse. While it covers the general purpose and storage, it doesn't explain the semantics of the board view, the output format, or how the action parameter maps to file operations. The schema and annotations fill some gaps, but the description alone would leave an agent uncertain about behaviors beyond the obvious.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
All nine parameters have descriptions in the schema, so the baseline for parameter semantics is a 3. The description adds minimal parameter-related context—it only mentions the action types and the file storage location, which indirectly relates to the 'project' parameter. No additional parameter meanings are provided.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's role with a specific verb ('Manage') and resource ('project tasks'), and enumerates the supported actions (add, list, update, board). This distinguishes it from the sibling toolset, which contains no other task-management tool.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies when to use the tool ('Supports add, list, update, and board views') by listing the primary operations. However, it doesn't explicitly discuss alternatives or exclusion criteria. Given the absence of sibling task tools, the usage context is nonetheless clear.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
templateARead-only
Get pre-filled templates for common vault item types (adr, prd, decision, learning, spec, rfc, roadmap, competitive-analysis, incident, research, vision, strategy). Use to scaffold new documents.
| Name | Required | Description | Default |
|---|---|---|---|
| type | No | Template type to retrieve | |
| action | No | Action (default: list) | |
| variables | No | Variables to substitute in the template (e.g. { title: 'My ADR' }) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The annotations already declare readOnlyHint: true, so the read-only nature is covered. The description adds the list of template types and the scaffolding context, which is useful, but it does not disclose behavioral details like variable substitution behavior or return format. With annotations providing the safety profile, a 3 is appropriate.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences, front-loaded with the action and resource, and lists the template types compactly. Every word earns its place; no verbose or redundant content.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a relatively simple tool with full schema coverage and no output schema, the description provides the essential context—the template types and the scaffolding use case. It does not need to explain action defaults or variable syntax since the schema covers them, making it sufficiently complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%—all three parameters (type, action, variables) are described in the schema. The description does not add any parameter semantics beyond what the schema already provides, so the baseline of 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb 'Get' and identifies the resource as 'pre-filled templates' for common vault item types, listing them explicitly. This clearly distinguishes it from sibling tools like read or write, and the phrase 'scaffold new documents' reinforces its purpose.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description includes 'Use to scaffold new documents', which provides a clear context for when to use this tool. It does not explicitly state when not to use it or mention alternatives, but the use case is well implied.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
writeBDestructive
Write or append to a file in the AI knowledge vault.
| Name | Required | Description | Default |
|---|---|---|---|
| mode | No | Write mode (default append) | |
| path | Yes | Relative path within vault | |
| content | Yes | Markdown content to write | |
| frontmatter | No | YAML frontmatter key-value pairs |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The annotation destructiveHint: true indicates the tool can destructively modify files, and the description does not contradict this. However, the description adds no behavioral traits beyond the annotation—it does not mention overwrite behavior, file creation, directory handling, or any side effects. The phrase 'Write or append' merely restates the core function.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, efficient sentence that is front-loaded with the action and target. There is no redundant wording or filler. It earns its place by clearly stating the tool's purpose without wasting tokens.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the rich schema and destructiveHint annotation, the description covers the basic purpose but leaves out behavioral nuances such as the effect of different modes, frontmatter handling, or return values. However, since the schema fills the parameter details and the annotation warns about destructiveness, the description is minimally adequate but not thorough.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema provides 100% coverage with descriptions for all four parameters (path, content, mode, frontmatter), including the enum and default for mode. The description adds no extra parameter semantics, so the baseline of 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's function with a specific verb ('Write or append') and resource ('file in the AI knowledge vault'). It distinguishes itself from the sibling 'read' tool and other non-file tools, making its purpose immediately clear.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use this tool versus alternatives. It does not mention any exclusions, prerequisites, or comparison to other write-related operations. The only context is the vault location, which is insufficient for decision-making.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
3 tool updates
v0.7.1- Added
knowledge_rebuild - Added
knowledge_viz - Added
qa_viz
30 tool updates
v0.6.1- First observed
brainstorm - First observed
capture - First observed
cred_refs - First observed
decide - First observed
deprecate - First observed
env_facts - First observed
extract - First observed
generate_context - First observed
graph_cross_project - First observed
graph_related - First observed
init - First observed
learn - First observed
link - First observed
project_context - First observed
prune - First observed
read - First observed
resume - First observed
rollback - First observed
search - First observed
session - First observed
skill_install - First observed
skill_list_installed - First observed
skill_remove - First observed
snapshot_repo_state - First observed
stats - First observed
status - First observed
superskill - First observed
task - First observed
template - First observed
write
TDQS
Scored across 33 tools
Most tools have distinct purposes, but several pairs blur together: qa_viz redoes knowledge_viz's HTML generation, graph_cross_project reads like a search tool rather than a cross-project graph tool, and capture/learn/extract all create vault items. The descriptions are detailed enough to disambiguate many cases, but a few boundaries are genuinely unclear.
All names use snake_case and are readable, but there is no consistent verb_noun or noun_verb convention: generate_context/search are verb-led, skill_install is noun+verb, skill_list_installed reverses the pattern, and status/task/template are bare nouns. qa_viz and superskill add product-specific abbreviations/names that don't fit the rest.
33 tools is too many for a single server's tool surface; the set spans knowledge vault CRUD, graph visualization, skill management, sessions, env facts, and rollback checkpoints. Several tools could be consolidated (knowledge_viz/qa_viz) or split into separate servers, and the volume makes selection harder.
The surface covers creation, reading, updating (via write), archival/deletion (prune/deprecate), plus specialized lifecycle tools for skills, sessions, tasks, and learnings. Notable minor gaps include no dedicated skill update, no unlink, and no explicit list/query for decisions, though search/read can work around them.
Maintenance
Related MCP Connectors
The governed runtime for agent skills. Search the catalog and inspect a skill before running it.
Intent execution engine for autonomous agent task routing
- SkilderOAuthai.skilder
One place to build, share, and govern the skills and tools your AI agents use at work.
Governed AI agent skills — one library, distributed to devs and exposed to remote agents over MCP.
Related MCP Servers
- AlicenseNot gradedqualityFmaintenanceProvides dynamic, context-aware code assistant skills through hybrid RAG (vector + knowledge graph), enabling runtime skill discovery, automatic toolchain-based recommendations, and on-demand loading from multiple git repositories.20MIT
- AlicenseNot gradedqualityAmaintenanceEnables AI agents to intelligently match tasks to skills through semantic embeddings, track skill effectiveness, detect skill gaps, and discover new skills from external sources.Apache 2.0
- AlicenseNot gradedqualityCmaintenanceProvides context-aware skill selection for AI agents, reducing token usage by 85-98% and improving accuracy through semantic retrieval, session memory, and feedback learning.1MIT
- AlicenseAqualityDmaintenanceRoutes SKILL.md libraries to any MCP client, enabling task matching and skill loading with embedding-based scoring, keyword fallback, and context-window discipline.517 npmMIT