ZeroWidth Compass
Server Details
Read and update your Compass map: pages, changes, opportunities and interviews.
- Status
- Healthy
- Last Tested
- Transport
- Streamable HTTP · MCP 2025-11-25
- URL
TDQS
Scored across 44 tools
Most tools target clearly distinct resources and actions (pages, links, gaps, interviews, comments, tags, docs), and descriptions explicitly name sibling tools to contrast (e.g. resolve vs update for gaps, create vs propose vs accept for opportunities). The main blur is that `compass_changes_upsert` operates on the same 'changes' that the `compass_opportunities_*` family manages, giving the same concept two names and two entry points. A few adjacent pairs (opportunities_create/propose/accept, pages_create vs opportunities_create) require reading carefully but are mostly disambiguated in text.
The compass/entity families use a consistent prefix_noun_verb pattern (compass_pages_create, comments_resolve, entity_tags_set), but the docs/search tools invert it to verb_noun (get_doc, list_docs, search_docs, search_workspace). There is also singular/plural drift (compass_interview_invite and compass_interview_targets vs the compass_interviews_* family) and the changes/opportunities dual prefix. Readable overall, but conventions are mixed.
44 tools is heavy well beyond the 16-25 range, and the surface covers numerous distinct domains (comments, opportunities/changes, gaps, inbox, interviews, links, map, pages, statuses, tags, docs, search). Some of the weight is genuine breadth, but the parallel changes-vs-opportunities families add redundant surface for the same resource. Scoping or consolidating the change/opportunity tools would bring it into range.
Coverage is strong: pages have full create/get/list/search/update/delete/restore, links and opportunities have full lifecycles, and gaps/interviews/comments each have the expected operations plus registry and review verbs. Minor gaps remain (comments have no update/delete) and several descriptions reference Ledger tools (ledger_entries_settle, ledger_metrics_list) that do not appear in this tool list, which could strand an agent mid-workflow.
Available Tools
44 toolscomments_createComment on an entityAInspect
Posts a comment on a workspace entity — a new thread, or a reply when rootId is given. Use it to leave findings where the discussion already lives (an eval result on the flow being debated, a summary on a long thread). Mention people via mentionedUserIds (from workspace member ids) to ring their notification bell; never mention someone who didn't ask to be pulled in.
| Name | Required | Description | Default |
|---|---|---|---|
| body | Yes | ||
| rootId | No | Reply into this thread; omit to start a new one. | |
| entityId | Yes | ||
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. | |
| approvalId | No | ||
| entityKind | Yes | What the thread hangs on. | |
| mentionedUserIds | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=false, destructiveHint=false and openWorldHint=false, so the write/safety profile is covered. The description adds value beyond that by disclosing a side effect annotations cannot express: mentioning users rings their notification bell, with an accompanying social caution. It omits permission/auth requirements for posting.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the core action and its two modes, then layers usage and mention etiquette compactly. Two sentences, minimal waste, though the trailing 'never mention someone who didn't ask' is advisory padding rather than invocation-critical information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 7-parameter mutation with 43% schema coverage and no output schema, the description covers the social/mention dimension well but leaves the entity-targeting parameters and the undocumented approvalId unexplained, which an agent would need to call this reliably in all cases.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is only 43%, so the description must compensate, and it partially does: rootId (reply target), mentionedUserIds (workspace member ids, notification behavior), and body are implied. However entityKind's 17-value enum, entityId, and especially approvalId are unexplained in both schema and description, leaving real gaps.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a precise verb and resource ('Posts a comment on a workspace entity') and distinguishes the two modes of operation: new thread vs. reply when rootId is given. It never names its closest siblings (comments_list, comments_resolve), so it stops just short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives concrete usage context ('leave findings where the discussion already lives') with two illustrative scenarios (eval result on a debated flow, summary on a long thread). It does not state when not to use it or point to a sibling alternative, so it lacks the explicit routing of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
comments_listRead an entity's comment threadsARead-onlyInspect
Lists the comment threads on one workspace entity (open first, then resolved) with authors and timestamps. Read this before weighing in on contested work — the threads are where disagreement lives before it becomes a decision.
| Name | Required | Description | Default |
|---|---|---|---|
| entityId | Yes | ||
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. | |
| entityKind | Yes | What the thread hangs on. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, destructiveHint=false, so the safety profile is covered. The description adds genuinely new behavior: threads are returned open-before-resolved and include authors and timestamps, which shapes how an agent interprets output. It stops short of pagination or volume limits.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences with the key facts front-loaded and no redundancy. The second sentence is motivational framing that carries mild value but is slightly softer than a hard routing rule.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only list tool with no output schema and a mostly documented schema, the description covers ordering, content, and the entity scoping needed to call it. Pagination/result-size behavior is the only notable gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 67%, and the schema itself handles the workspace slug nuances and the enum list. The description only says 'one workspace entity', adding little beyond the structured fields, so the baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (lists) and resource (comment threads on one workspace entity), plus the ordering rule (open first, then resolved) and payload (authors, timestamps). This clearly separates it from comments_create and comments_resolve, which mutate rather than read.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The second sentence gives a soft situational cue ('read this before weighing in on contested work'), which implies when the tool is useful. However it names no alternatives and gives no explicit when-not or prerequisite guidance, so usage remains inferred rather than stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
comments_resolveResolve or reopen a threadADestructiveInspect
Sets a comment thread's resolved state (rootId = the thread's root comment id). Resolve ONLY when the human asked or the thread's question is demonstrably settled — and say what settled it in a reply first. Reopening is for new evidence.
| Name | Required | Description | Default |
|---|---|---|---|
| rootId | Yes | ||
| resolved | Yes | ||
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. | |
| approvalId | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true, readOnlyHint=false and openWorldHint=false, so the mutation/safety profile is covered. The description adds real behavioral context beyond that: the precondition (a reply explaining what settled the thread) and the reopen semantics. It doesn't clarify reversibility or how the state change affects existing replies, keeping it below 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences with no filler; the action and the primary precondition are front-loaded before the reopening clause. Every clause carries decision-relevant information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a small toggle tool with no output schema, the description covers the action, the target, and the conditions well. It leaves approvalId and workspace behavior unexplained, which matters for a destructive write, but overall an agent has enough to act correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is only 25%, and the description compensates for rootId only (root comment id). 'resolved' is implied by resolve/reopen framing, but 'workspace' is documented only in the schema and 'approvalId' is explained nowhere — a notable gap for a mutation tool with an approval parameter.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (sets resolved state) plus resource (comment thread), and covers both directions — resolve and reopen. The parenthetical 'rootId = the thread's root comment id' disambiguates the target, making it clearly distinct from comments_create/comments_list.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly gates the action: resolve ONLY when the human asked or the question is demonstrably settled, and reply first explaining what settled it; reopen is for new evidence. This is genuine when/when-not guidance with a required prior step.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
compass_changes_upsertBring changes in from another trackerADestructiveInspect
Create or update up to 100 changes in one call, each matched by url — the GitHub issue, Linear or Jira ticket, PR or doc it mirrors. A url already linked to a change updates that change (label, description, status, tags, experiment, workflow — omitted fields are left alone); any other url creates a change with that link. Run it again with the same urls to keep Compass in step; nothing duplicates. Statuses are by name (compass_statuses_list). Returns one result per item — created / updated with the change's key, or failed with why; a failed item doesn't stop the rest. Summarize what you're about to bring in before calling. May return needs_confirmation — one approval covers the whole batch.
| Name | Required | Description | Default |
|---|---|---|---|
| items | Yes | ||
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. | |
| approvalId | No | Approval id from a prior needs_confirmation response, after the user has approved. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Goes well beyond the annotations (which only flag the write + destructive + closed-world profile): it discloses partial-update semantics ('omitted fields are left alone'), idempotency ('nothing duplicates'), per-item isolation ('a failed item doesn't stop the rest'), the 100-item cap, and the needs_confirmation approval flow where one approval covers the batch. This is unusually rich behavioral disclosure.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Dense but front-loaded — the core upsert action leads, followed by matching semantics, side effects, and confirmation behavior. Parenthetical and em-dash clauses pack multiple facts per sentence, which is efficient though slightly heavy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description still explains the return shape (one result per item: created/updated with key, or failed with reason) and the confirmation path. For a batched mutation tool this covers everything an agent needs to call and interpret it.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 67% and the description closes gaps: `url` is the match key mirroring a GitHub/Linear/Jira ticket, PR or doc; omitted fields are left untouched; statuses are referenced by name via compass_statuses_list. It does not explain `workspace` or `approvalId`, but those already carry schema descriptions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb pair (create or update) and resource (changes), plus scope: up to 100 per call, each matched by `url`. The upsert-by-url framing is unambiguous and no sibling tool operates on changes, so an agent can select this without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear operational context: run again with the same urls to keep Compass in step, summarize before calling, and use compass_statuses_list for status names. It does not name a when-not-to-use case or an alternative tool, but no overlapping sibling exists, so the routing guidance is effectively complete.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
compass_gaps_createRaise a Compass open question (gap)AInspect
Records something you don't yet know — the gap registry is your working memory, so use it liberally while mapping or interviewing. Give one clear question, WHY it matters (what answering it unblocks), and how you'd resolve it (ask the user / interview a specific person / connect a source). No approval needed — noting your own uncertainty isn't acting on the user's behalf. Attach it to a subject (subjectType/subjectId) when it's ABOUT a specific page or person.
| Name | Required | Description | Default |
|---|---|---|---|
| question | Yes | The open question, plainly ('Who approves refunds over $500?'). | |
| rationale | No | Why it matters / what answering it unblocks. | |
| subjectId | No | Id of the subject. | |
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. | |
| subjectType | No | What it's about, e.g. compass_page. | |
| suggestedResolution | No | How to resolve it: 'ask the user', 'interview Dana', 'connect Drive'. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already flag a non-destructive write (readOnlyHint=false, destructiveHint=false), and the description adds meaningful behavioral context beyond them: no approval is required and why ('noting your own uncertainty isn't acting on the user's behalf'). It doesn't describe the returned record or any limits, but the auth/approval disclosure is genuinely additive.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three front-loaded sentences that move from what to record, to why no gating is needed, to when to attach a subject. Dense but each sentence carries guidance; no filler or repetition of the schema.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a single-record write with no output schema and full schema coverage, the description supplies the missing 'why' and 'approval' context plus subject-attachment guidance. It leaves workspace selection entirely to the schema but otherwise covers what an agent needs to invoke it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so baseline is 3, and the description earns an increment by tying the fields to intent — the question, the WHY (rationale), and the resolution path (suggestedResolution) — plus the conditional semantics of subjectType/subjectId ('when it's ABOUT a specific page or person').
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource — recording an open question into the gap registry — and immediately distinguishes itself from siblings by describing the registry's role as working memory. An agent can separate it from compass_gaps_list/resolve/update without opening a schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear context for when to use it ('while mapping or interviewing', 'use it liberally') and specifies the condition for attaching a subject ('when it's ABOUT a specific page or person'). It does not explicitly name the sibling tools it supersedes (gaps_update/resolve), so it stops short of full alternative routing.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
compass_gaps_listList Compass open questions (gaps)ARead-onlyInspect
Lists the workspace's open questions — the gap registry: what the map doesn't know yet. Defaults to OPEN gaps. This is where you keep your head — check it before asking the user something you may already have flagged, and compose interview briefs from a subject's open questions.
| Name | Required | Description | Default |
|---|---|---|---|
| status | No | Filter by lifecycle stage. Default OPEN. | |
| subjectId | No | Only gaps ABOUT this subject id. | |
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. | |
| subjectType | No | Only gaps ABOUT this subject type (e.g. compass_page). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, destructiveHint=false and openWorldHint=false, so safety is covered. The description adds the default-OPEN behavior and the pre-ask check habit, but says nothing about result volume, pagination, or ordering — modest added value beyond structured fields.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core statement of what is listed and the default, followed by usage. The colloquial aside ('This is where you keep your head') costs a few words but does carry usage guidance, so little is wasted overall.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only list tool with four fully documented, all-optional parameters and no output schema, the description covers purpose, default filtering, and workflow integration adequately. Only pagination/return-shape behavior is left unstated, which is a minor gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all four parameters (status enum, subjectId, subjectType, workspace) are already documented in the schema. The description only restates the OPEN default and the 'subject's open questions' notion, adding no syntax or semantics beyond what the schema provides. Baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Lists the workspace's open questions — the gap registry') and characterizes what a gap is ('what the map doesn't know yet'). This clearly distinguishes it from the sibling mutations compass_gaps_create/resolve/update without needing to open any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives concrete workflow triggers: check this before asking the user something already flagged, and use it to compose interview briefs from a subject's open questions. That is strong when-to-use guidance, though it names no explicit alternative or exclusion (e.g. vs. compass_inbox_list).
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
compass_gaps_resolveResolve a Compass open question (gap)ADestructiveInspect
Closes a gap once you've learned the answer (ANSWERED, with the resolution) or decided it doesn't matter (DISMISSED). Keep the registry honest — resolve gaps as their answers land (from an interview, a doc, or the user) so it always reflects what's still unknown. Gaps you raised close without a card; a question a person wrote needs their approval.
| Name | Required | Description | Default |
|---|---|---|---|
| gapId | Yes | Id of the gap (from compass_gaps_list). | |
| status | Yes | ANSWERED (you learned it) or DISMISSED (irrelevant). | |
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. | |
| approvalId | No | Re-call with the approvalId from a needs_confirmation response after the user approves. | |
| resolution | No | The answer / note, when marking ANSWERED. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true and readOnlyHint=false, so the terminal-state safety profile is covered. The description adds genuinely new behavioral context: gaps you raised close without a card, while a person-written question requires their approval — an authorization workflow not encoded in the annotations or schema.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, front-loaded with the core action and its two outcomes. The middle 'Keep the registry honest' sentence is mildly exhortative but carries the when-to-use guidance, so it earns its place; nothing dangles.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a mutation tool with no output schema, the description covers purpose, trigger conditions, status semantics, and the approval prerequisite. It does not describe the success response or reversibility, but the annotations plus 100% schema coverage leave the agent adequately equipped.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so gapId, status, workspace, approvalId, and resolution are all documented in the schema. The description reinforces the ANSWERED-with-resolution relationship and the approvalId flow but adds little beyond what the schema already states, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb (Closes/resolve) and resource (a gap = open question), and enumerates the two terminal outcomes ANSWERED and DISMISSED with their meanings. It clearly conveys the terminal-state nature, but never names the adjacent sibling compass_gaps_update, so an agent must infer the resolve-vs-update boundary.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It says when to act — resolve gaps as their answers land, explicitly listing sources (interview, doc, or the user) and the DISMISSED case when it proves irrelevant. It lacks explicit 'instead of' routing to siblings, but the trigger conditions are clear.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
compass_gaps_updateReword a Compass open question (gap)ADestructiveInspect
Edits an open question's wording, rationale, or suggested resolution — for sharpening a vague question or fixing one you phrased badly. Closing a gap is compass_gaps_resolve, not this. Gap ids come from compass_gaps_list. May return needs_confirmation.
| Name | Required | Description | Default |
|---|---|---|---|
| gapId | Yes | Id of the gap (from compass_gaps_list). | |
| question | No | New wording of the question. | |
| rationale | No | Why it matters / what answering it unblocks. | |
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. | |
| approvalId | No | Approval id from a prior needs_confirmation response, after the user has approved. Omit on the first call. | |
| suggestedResolution | No | How to resolve it: 'ask the user', 'interview Dana', 'connect Drive'. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=false and destructiveHint=true, so the safety profile is carried structurally. The description adds genuine context by disclosing the 'needs_confirmation' return path, but it never explains why an edit is destructive (e.g. what prior wording is overwritten) nor the approval requirement behind the destructive hint, which is the notable gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, front-loaded with purpose, then exclusions, then id provenance and the confirmation behavior. Every clause earns its place with no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 6-param mutation with full schema coverage and annotations covering safety, the description covers purpose, alternative, id source, and the confirmation flow. It stops short of explaining the destructive nature of the edit or the workspace/token rules that the schema mentions, and with no output schema it only names 'needs_confirmation' rather than describing returns, but it is nearly sufficient.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so gapId, question, rationale, workspace, approvalId, and suggestedResolution are all documented in the schema. The description only reinforces the field set and the provenance of gapId, adding nothing on syntax or format beyond what the schema already provides, which is the expected baseline.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Edits) and resource (an open question's wording, rationale, or suggested resolution), enumerating exactly which fields change. It names the sibling it is not (compass_gaps_resolve) so an agent can distinguish editing from closing without opening either schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives an explicit when-to-use ('for sharpening a vague question or fixing one you phrased badly') and an explicit when-not, routing closing to compass_gaps_resolve. It also tells the agent where the required gapId comes from (compass_gaps_list), leaving nothing to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
compass_inbox_listWhat's waiting for a decision in CompassARead-onlyInspect
The Compass Inbox in one read: finished interviews awaiting review (transcript in, not yet turned into pages), AI-proposed opportunities awaiting accept / dismiss, and the caller's pending approval cards for Compass writes. THE place to answer 'what needs me?' for the map. Next moves: compass_interviews_get to read a transcript, then propose pages from it with compass_pages_create; compass_opportunities_accept or compass_opportunities_delete for a proposal.
| Name | Required | Description | Default |
|---|---|---|---|
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, destructiveHint=false, and openWorldHint=false, so the safety profile is covered. The description adds genuine context beyond that: the aggregate nature of the result, the state filter ('transcript in, not yet turned into pages'), and that approval cards are scoped to the caller. Pagination/limits are not mentioned, keeping it below a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the three inbox categories before the routing sentence, so the agent gets the core scope first. The colon-list and semicolon-chained next-moves sentence are dense but every clause carries information; no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema, so the description must describe what comes back, and it does: the three item categories and their pending states. Combined with annotations covering safety and the schema covering the sole parameter, an agent has enough to call it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and there is a single optional workspace parameter, so the schema fully documents it. The description adds nothing about the workspace parameter (e.g., how it affects which inbox items appear), so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (list) and a precisely scoped resource (the Compass Inbox) and enumerates the three categories it aggregates: finished interviews awaiting review, AI-proposed opportunities awaiting accept/dismiss, and the caller's pending approval cards. This distinguishes it from the narrower siblings compass_interviews_list and compass_opportunities_list without the agent needing to open either schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly positions itself as 'THE place to answer what needs me?' and gives concrete next moves that route to compass_interviews_get/compass_pages_create and compass_opportunities_accept/delete. It lacks an explicit when-not (e.g., pointing to compass_opportunities_list for a full unfiltered list), so it stops short of full alternative coverage.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
compass_interview_inviteEmail an interview inviteAInspect
Emails the guest link for an existing interview, framed as coming from the REQUESTING USER (their name signs it; replies go to them). Author message in their voice — short, human, says why THEIR knowledge matters and that it takes ~15 minutes, no account needed. The approval card shows the exact subject + message before anything sends. Limit: the invite plus ONE reminder; a third ask is the user's conversation to have. Completion arrives as a notification with draft-page counts — don't poll.
| Name | Required | Description | Default |
|---|---|---|---|
| Yes | The interviewee's email address. | ||
| message | Yes | The body, written in the requesting user's voice. Greeting + link + signature are added automatically — write only the middle. | |
| subject | Yes | Email subject, e.g. '15 minutes on how refunds actually work?' | |
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. | |
| approvalId | No | Approval id from a prior needs_confirmation response, after the user has approved. Omit on the first call. | |
| interviewId | Yes | Interview id from compass_interviews_create. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Goes well beyond the annotations: discloses the approval-card preview of subject+message before sending, the email is sent as the requesting user (replies route to them), a hard limit of one reminder, and that completion arrives as a notification rather than a polled response. This is exactly the behavioral context annotations cannot convey.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core action, then packs voice, approval, limit, and completion into tight clauses with no filler. It is dense with several distinct concerns in one paragraph, but each sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, and the description compensates by explaining how the result surfaces (notification with draft-page counts, don't poll) plus the confirmation flow and send limits. An agent has everything needed to invoke and follow through correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the field descriptions already carry the load (baseline 3). The description adds authoring guidance beyond the schema — 'short, human, says why THEIR knowledge matters and that it takes ~15 minutes' — which shapes how `message` and `subject` should be written.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: 'Emails the guest link for an existing interview.' The word 'existing' distinguishes it from compass_interviews_create, and the framing detail (from the requesting user) further narrows its purpose. An agent can place it without opening any sibling schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Says to use it for existing interviews and gives clear operational rules — invite plus ONE reminder, the third ask is the user's conversation, and don't poll. It implies but never explicitly names the sibling alternatives (create vs invite vs revoke), so it stops short of full routing guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
compass_interviews_createInvite someone to a Compass interviewAInspect
Mints a stakeholder-interview invite: a no-account guest link where the person talks to an interviewer agent briefed by your focusPrompt, and the transcript flows back into Compass as reviewable draft pages. THE KNOWLEDGE-GAP MOVE: when caliper_flow_performance shows failures clustered on missing company facts (the judge says the flow invented a policy, missed a rule, didn't know who owns something), the fix is usually not a prompt edit — it's asking the human who actually knows. Write a focusPrompt that names the SPECIFIC gaps (cite the eval run id), pick the owner of the relevant workflow as interviewee when the map knows one, and hand the user the invite link to forward. May return needs_confirmation — tell the user who you want to interview and why, then wait.
| Name | Required | Description | Default |
|---|---|---|---|
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. | |
| approvalId | No | Approval id from a prior needs_confirmation response, after the user has approved. Omit on the first call. | |
| focusPrompt | Yes | What the interviewer should dig into — specific, grounded in the gap you found (≤2000 chars). Never shown verbatim to the guest. | |
| contextPageIds | No | Compass page ids the interviewer gets as briefing context (e.g. the workflow page whose flow underperformed). | |
| intervieweeName | No | Who this invite is for, when known. | |
| intervieweeEmail | No | Their email, when known — enables compass_interview_invite to send the link directly (with approval). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare openWorldHint=true and readOnlyHint=false; the description adds real context on top of that — the invite is a no-account guest link, the transcript returns as reviewable draft pages, and the call may return needs_confirmation requiring a user pause. It does not cover auth or rate-limit behavior, but the confirmation and side-effect profile is well disclosed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core definition, then the usage rule, then the closing procedure. Dense but nearly every clause carries information; the final sentence is a slightly compressed run-on covering needs_confirmation and the user hand-off, but it is not padding.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 6-param tool with no output schema, the description covers the important return behavior (needs_confirmation pause) and the downstream outcome (draft pages), and pairs naturally with approvalId in the schema. Enough for an agent to invoke and handle the response correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Even though schema coverage is 100% (baseline 3), the description adds guidance the schema lacks: focusPrompt should name SPECIFIC gaps and cite the eval run id, and interviewee should be the workflow owner when the map knows one. This meaningfully shapes how an agent fills the parameters rather than restating them.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('mints a stakeholder-interview invite') and describes the concrete artifact produced: a no-account guest link, an interviewer agent briefed by focusPrompt, and a transcript flowing back as draft pages. It is clearly distinguishable from sibling compass_interview_invite, which is referenced as the send step rather than the mint step.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly names when to use this over the obvious alternative: when caliper_flow_performance shows failures clustered on missing company facts, 'the fix is usually not a prompt edit — it's asking the human who actually knows.' It also prescribes the operational sequence (name gaps with eval run id, pick the workflow owner, hand over the link) and the needs_confirmation handling.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
compass_interviews_getRead one Compass interview, transcript includedARead-onlyInspect
One interview in full — metadata plus the transcript of what the guest actually said. Read this before summarizing an interview or drafting pages from it: the transcript is the source you work from, and your compass skill has the rules for typing and splitting what's in it. Ids come from compass_interviews_list or compass_inbox_list.
| Name | Required | Description | Default |
|---|---|---|---|
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. | |
| interviewId | Yes | Interview id (from compass_interviews_list). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, destructiveHint=false and closed-world scope, so the safety profile is covered. The description adds genuinely useful behavioral context with no output schema: what the payload contains (metadata plus the actual transcript) and that downstream typing/splitting rules live in the compass skill.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, front-loaded with what the tool returns before the usage advice. The clause about the compass skill's typing/splitting rules is slightly tangential but still earns its place by explaining why the transcript matters.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a single-resource read with full schema coverage, covered annotations and no output schema, the description supplies what an agent needs: the return content, the id source, and the workflow context. It does not address failure modes (unknown id, missing workspace) but little more is required here.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline would be 3. The description goes slightly beyond the schema by naming compass_inbox_list as an additional source of interviewId, which the schema's own description does not mention. It says nothing about the workspace parameter, so this is a modest lift, not a full one.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Opens with a specific verb and resource: fetch one interview in full, including metadata and transcript. It is immediately distinguishable from compass_interviews_list (plural enumeration) and from the create/revoke siblings that mutate interview state.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
States the trigger condition explicitly ('read this before summarizing an interview or drafting pages from it') and tells the agent where the required id comes from (compass_interviews_list or compass_inbox_list). That is concrete routing guidance rather than implied usage.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
compass_interviews_listList Compass interviewsARead-onlyInspect
Every stakeholder interview in the workspace, newest first: who was invited, status (INVITED / IN_PROGRESS / COMPLETED), focus, message count, whether the transcript was already reviewed (documentPageId set), and the invite URL while the link is still live. Check this before inviting someone again — interview fatigue is real — and to answer 'who have we already asked?'. Read one in full with compass_interviews_get.
| Name | Required | Description | Default |
|---|---|---|---|
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, destructiveHint=false and openWorldHint=false, so the safety profile is covered. The description adds real behavioral context beyond that: ordering (newest first), the status lifecycle values, and the fact that the invite URL is only present while the link is live. It does not mention pagination or result-size limits, which is the one remaining gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, front-loaded with scope and ordering, then the returned-field inventory, then the usage trigger and the sibling pointer. The informal aside ('interview fatigue is real') is short and justifies the routing guidance rather than padding.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema, and the description compensates by enumerating the returned fields and their semantics. Combined with annotations covering the read-only/non-destructive nature and the explicit pointer to the detail tool, an agent has everything needed to call this correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and the single workspace parameter is fully documented in the schema, so the baseline of 3 applies. The description adds nothing about the workspace scoping or how default-workspace tokens behave, leaving all parameter meaning to the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Every stakeholder interview in the workspace, newest first') and enumerates exactly what each record contains — invitee, status enum, focus, message count, documentPageId, invite URL. It is immediately distinguishable from compass_interviews_get, which it explicitly routes to for full detail.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives an explicit when-to-use ('Check this before inviting someone again') plus a concrete user question it answers ('who have we already asked?'), and names the alternative tool (compass_interviews_get) with the condition that selects it. Nothing is left to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
compass_interviews_revokeRevoke an interview invite linkADestructiveInspect
Kills an interview's invite link — the guest's next visit sees that the link is no longer active. Use when an invite went to the wrong person, the user changed their mind, or the link leaked. Idempotent. Ids come from compass_interviews_list. May return needs_confirmation.
| Name | Required | Description | Default |
|---|---|---|---|
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. | |
| approvalId | No | Approval id from a prior needs_confirmation response, after the user has approved. Omit on the first call. | |
| interviewId | Yes | Interview id (from compass_interviews_list). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the destructiveHint=true/readOnlyHint=false annotations, the description discloses idempotency, the exact user-visible consequence of revocation, and the `needs_confirmation` approval flow. These are meaningful behavioral traits an agent could not infer from the annotations alone.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, front-loaded with the action and effect before usage triggers and edge behaviors. No filler; every clause adds information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a destructive single-object mutation with no output schema, the description covers effect, idempotency, id provenance, and the confirmation follow-up path. Nothing essential for correct invocation is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so interviewId, workspace, and approvalId are already documented. The description only restates the interviewId source already present in the schema, so with the schema carrying the load a baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Kills an interview's invite link') plus the concrete effect ('the guest's next visit sees that the link is no longer active'). This clearly separates it from the sibling compass_interview_invite, which creates invites rather than revoking them.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit trigger conditions: wrong recipient, changed mind, or leaked link. It covers when to use but does not name a contrasting alternative or state when-not to use (e.g., a different tool for regenerating vs. fully revoking).
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
compass_interview_targetsWho should be interviewed about a pageARead-onlyInspect
Ranks the PEOPLE the map says know about a page (workflow, system, pain point): owns links first, then involved_in, then weaker edges. Each target carries contact email + role from their PERSON page and their latest interview (skip someone who just gave one — interview fatigue is real). Empty result = the map doesn't know an owner: ASK THE USER who runs this, create the PERSON page + owns link from the answer, and the map gets smarter. Use before compass_interviews_create to pick the interviewee.
| Name | Required | Description | Default |
|---|---|---|---|
| pageId | Yes | The page the knowledge gap is about. | |
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare a safe read (readOnlyHint=true, destructiveHint=false), and the description adds substantial extra behavior: the ranking order (owns > involved_in > weaker edges), fatigue-avoidance skip logic for recent interviewees, and the empty-result recovery path. This is well beyond what the annotations convey.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core action, then ranking order, then empty-result handling, then the sibling pointer. Dense but each sentence carries actionable information; minor length from parentheticals keeps it just short of ideal.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description still explains what each target carries (contact email, role, latest interview), so an agent knows what it gets back. It also covers the failure mode, making it complete for a 2-param read tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% and the schema already documents pageId and workspace semantics in detail. The description adds no syntax or format detail beyond 'a page', so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (ranks) and resource (PEOPLE who know about a page), with explicit scope and an ordered edge taxonomy. It is clearly distinguishable from siblings like compass_pages_get or compass_interviews_create.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly says 'Use before compass_interviews_create to pick the interviewee', naming the sibling and the sequencing. It also prescribes what to do on an empty result (ask the user, create PERSON + owns link), covering the when and the when-not.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
compass_links_createConnect two Compass pagesAInspect
Creates a typed, directed edge between two pages — the knowledge graph's connective tissue. Canonical directions: PERSON owns WORKFLOW/SYSTEM, PERSON involved_in WORKFLOW, WORKFLOW uses SYSTEM, SYSTEM uses SYSTEM, PAIN_POINT affects WORKFLOW/SYSTEM/PERSON, DOCUMENT documents anything, relates_to as fallback. Idempotent on (from, to, kind) — re-creating an existing edge returns it. May return needs_confirmation; tell the user what you're proposing, wait for their approval, then re-call with the approvalId.
| Name | Required | Description | Default |
|---|---|---|---|
| kind | Yes | Edge type, in canonical direction. | |
| note | No | ≤500-char qualifier when the kind alone undersells it. | |
| source | No | Provenance label. Defaults AI_ACCEPTED (agent writing under live human direction); pass USER for human-driven scripts. | |
| toPageId | Yes | Edge target page id. | |
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. | |
| approvalId | No | Approval id from a prior needs_confirmation response, after the user has approved. Omit on the first call. | |
| fromPageId | Yes | Edge source page id. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only declare the basic safety profile (not read-only, not destructive, closed-world). The description adds real behavioral traits beyond that: idempotency keyed on (from, to, kind), the returned-existing-edge behavior, and the needs_confirmation → user approval → re-call with approvalId loop.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Purpose is front-loaded in sentence one, followed by direction semantics, then idempotency, then the confirmation protocol — a deliberate and dense ordering with no filler sentences.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 7-parameter mutation with no output schema, the description covers the create semantics, idempotency, and the approval round-trip. It omits error/failure behavior and permission requirements, but the schema already carries workspace/auth and provenance details.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so baseline is 3, but the description goes further by enumerating canonical direction pairs for the `kind` parameter — semantics the schema's circular 'Edge type, in canonical direction' text does not supply. The approvalId and workspace behavior are already well documented in the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
First sentence states a specific verb and resource — creates a typed, directed edge between two pages — plus a concrete metaphor for its role in the knowledge graph. It is clearly distinguishable from siblings like compass_links_update and compass_links_delete.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides substantive when-to-use guidance via the canonical direction list (PERSON owns WORKFLOW, relates_to as fallback), which tells the agent how to pick the right invocation. It does not explicitly name when to prefer links_update over create, but the idempotency note covers the main overlap case.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
compass_links_deleteRemove a connection between two pagesADestructiveInspect
Deletes one edge from the graph (the pages stay). Use when a connection is simply wrong — the system isn't used by that workflow, the person left the team. Link ids come from compass_page_links_list. Re-creating the same edge later revives it. May return needs_confirmation.
| Name | Required | Description | Default |
|---|---|---|---|
| linkId | Yes | Id of the link (from compass_page_links_list). | |
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. | |
| approvalId | No | Approval id from a prior needs_confirmation response, after the user has approved. Omit on the first call. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true, but the description adds genuinely non-obvious behavior: deleting is reversible ('Re-creating the same edge later revives it') and the call may return needs_confirmation, implying a gated approval flow. That is real context beyond the structured fields; it stops short of stating permission requirements or exactly what the confirmation response contains.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences: the destructive scope first, the usage trigger second, the reversibility/confirmation caveat last. No filler, and the most decision-relevant fact (edges only, pages stay) is front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description carries the return-value burden and does flag the needs_confirmation outcome and its approval follow-up. Coverage is good for a single-edge delete; only details like idempotency on a stale linkId are left implicit.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% — linkId sourcing, workspace rules and the approvalId flow are all documented in the schema. The description only echoes the linkId provenance already stated in the schema, so this is the baseline case where the schema does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Deletes one edge from the graph') and immediately scopes it against the obvious confusion — the pages themselves survive, which separates it from compass_pages_delete. An agent can distinguish it from compass_links_create/update and the page tools without opening a schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives an explicit when-to-use rationale with concrete triggers ('the system isn't used by that workflow, the person left the team') and points to compass_page_links_list as the source of ids. It does not name the sibling alternative (e.g., compass_links_update) or state when not to delete, so it falls short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
compass_links_updateChange a connection's reading or noteADestructiveInspect
Edits an existing edge between two pages: its kind (how the two relate) and/or its note. Use to CORRECT a connection typed wrongly — 'Dana doesn't own billing, she's involved in it'. Link ids come from compass_page_links_list. Fails with conflict when the new reading already exists between the same two pages (delete this one instead). May return needs_confirmation.
| Name | Required | Description | Default |
|---|---|---|---|
| kind | No | New edge type, in canonical direction. | |
| note | No | New qualifier note (≤500 chars); empty string clears it. | |
| linkId | Yes | Id of the link (from compass_page_links_list). | |
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. | |
| approvalId | No | Approval id from a prior needs_confirmation response, after the user has approved. Omit on the first call. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only declare the mutation/destructive profile; the description adds real behavioral content beyond that — the conflict failure mode when an identical reading already exists, the routing advice to delete instead, and the possibility of a 'needs_confirmation' response. It stops short of stating reversibility or permission requirements for a destructive edit.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads purpose, then usage, then failure behavior in roughly four tight sentences with no filler. The ALL-CAPS emphasis and quoted example are slightly chatty but earn their space by disambiguating the correction use case.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description carries the burden of describing returns and does mention the 'needs_confirmation' response and conflict failure, which are the two non-obvious outcomes. It is complete enough to call correctly, though it omits the ordinary success return shape and any permission prerequisites for this destructive edit.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all five parameters including the kind enum and the 500-char note limit are already documented in the schema. The description's gloss of `kind` as 'how the two relate' adds only marginal meaning beyond the schema's 'New edge type, in canonical direction.'
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Edits an existing edge between two pages') and enumerates the two mutable fields, kind and note. This cleanly separates it from compass_links_create and compass_links_delete without the agent needing to open either schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives an explicit triggering scenario ('Use to CORRECT a connection typed wrongly') with a concrete example, names the prerequisite source for link ids (compass_page_links_list), and routes the agent to an alternative ('delete this one instead') when the new reading already exists.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
compass_map_viewView the Compass map (rendered image)ARead-onlyInspect
Renders the workspace's visual map to an image and returns it so you can SEE it the way the user does: an isometric drawing where every page is a structure whose shape is its type — people are figures, audiences are crowds, workflows are gears lying on the ground, systems are database drums, offerings are price tags, pain points are warning signs, documents are standing sheets of paper, values are shields, priorities are flags — sized by how many other pages connect to them, placed near what they link to, with pages connected to nothing parked to one side, and routes drawn between linked pages. A coloured ring on the ground around a page shows the changes touching it by stage (blue undecided, amber not started, green in progress, violet done; thicker = more), with a key in the top-left corner. Use this when the user asks about the shape of their map, where the problems are, what connects to what, or anything spatial. The text part counts the pages of each type, the links, and how many pages are unconnected, plus mapUrl. The image is for you to look at; do NOT put it in your reply as a markdown image (the chat can't show it). When the user should see the map, link mapUrl — e.g. [Open your map](mapUrl) — and describe what you saw.
| Name | Required | Description | Default |
|---|---|---|---|
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only establish that this is a safe read; the description goes well beyond them by disclosing what the non-image payload contains (page-type counts, link count, unconnected-page count, mapUrl) and, critically, an anti-pattern instruction: do NOT embed the image in a reply because the chat cannot render it. That is exactly the kind of behavioral guidance structured fields cannot carry.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The usage trigger is front-loaded and the closing instructions are tight, but the long enumeration of every page type and its symbol, plus the full colour-ring legend, is bulky and would be needed only when interpreting a specific image. It earns most of its length but not all of it.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema and a single trivial parameter, the description carries the full burden and does so: it explains what the image depicts, what the text block returns, and how the agent should present the result. Nothing needed to call or act on this tool correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% with a single optional 'workspace' parameter, so the schema already documents the token/default-workspace semantics. The description adds nothing about the parameter, so the baseline of 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb+resource ('renders the workspace's visual map to an image and returns it') and the vivid type-to-shape mapping makes clear exactly what the artifact contains. No sibling tool renders a map, so it is trivially distinguishable from compass_pages_list, compass_page_links_list, and the rest.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives an explicit when-to-use trigger ('when the user asks about the shape of their map, where the problems are, what connects to what, or anything spatial') and pairs it with presentation guidance ('link mapUrl... and describe what you saw'). The routing condition is unambiguous rather than inferred.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
compass_opportunities_acceptAccept an AI-proposed opportunityADestructiveInspect
The review verb for the Inbox: promotes an AI-PROPOSED opportunity (from compass_opportunities_propose or the opportunity scout) into the workspace's own pipeline. compass_opportunities_set_status does NOT do this — a proposal stays in the Inbox until accepted. Optional status decides its lane in the same step (BACKLOG to park it, EXPERIMENTING to start it); omitted, it lands in NEW. Fails with not_proposed when the row isn't an AI proposal. Ids come from compass_inbox_list / compass_opportunities_list. May return needs_confirmation — name the proposal and wait.
| Name | Required | Description | Default |
|---|---|---|---|
| status | No | Lane to place it in on accept (default NEW). | |
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. | |
| approvalId | No | Approval id from a prior needs_confirmation response, after the user has approved. Omit on the first call. | |
| opportunityId | Yes | The AI-proposed opportunity's id. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true and readOnly=false, so mutation semantics are covered. The description adds genuinely useful behavior beyond that: the not_proposed failure condition, the needs_confirmation round-trip with approvalId, and the status defaulting to NEW. It stops short of explaining what 'promote' does to the source proposal or any irreversibility, so not a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Dense but front-loaded: the core action and sibling distinction lead, followed by parameter behavior, failure mode, and the confirmation caveat. Every sentence carries information, though the several clauses make it slightly heavy for a single-paragraph description.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, yet the description covers the return-signal territory an agent needs (not_proposed failure, needs_confirmation handling, where ids originate). For a mutation tool with full annotation coverage, nothing essential is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3. The description nonetheless adds real meaning: status chooses the lane in the same step with a stated default, and approvalId is tied to the two-step needs_confirmation flow (omit on first call). This goes beyond the schema's field-level text.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource (promotes an AI-proposed opportunity into the workspace pipeline) and explicitly differentiates from compass_opportunities_set_status, which 'does NOT do this.' An agent can distinguish it from all siblings without opening a schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly names the source of proposals (compass_opportunities_propose / opportunity scout), the alternative that won't work (set_status), and where ids come from (compass_inbox_list / compass_opportunities_list). When-to-use is fully covered with the counter-case named.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
compass_opportunities_createCapture an opportunityAInspect
Record a change — a fix, a chore, a new step, or an experiment worth trying. Anchor it to a workflow page when one fits (compass_pages_list); leave unanchored otherwise. It lands in the board's first status unless you name another (compass_statuses_list). Set experiment only when the user wants it measured against the Ledger. May return needs_confirmation — summarize and wait for approval.
| Name | Required | Description | Default |
|---|---|---|---|
| tags | No | Tags to file it under. | |
| label | Yes | ≤200 chars, the change in one line. | |
| links | No | Where the work also lives — issue, ticket, PR or doc URLs. | |
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. | |
| approvalId | No | Approval id from a prior needs_confirmation response, after the user has approved. Omit on the first call. | |
| experiment | No | Measure it against the Ledger. Default false. | |
| statusName | No | Start in this status (a name from compass_statuses_list). | |
| description | No | What the change is, and why it's worth doing. | |
| workflowPageId | No | Workflow page to anchor to, when one fits. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With annotations already signaling a non-destructive, closed-world write, the description adds genuinely new behavior: the default that it 'lands in the board's first status' and the two-phase approval flow ('May return `needs_confirmation` — summarize and wait for approval'). These are non-obvious traits an agent must handle, not restatements of annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core action, then each sentence covers one decision (anchoring, default status, experiment, approval). No filler; every clause carries an actionable instruction.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 9-param create tool with no output schema, the description covers anchoring, status defaulting, experiment gating, and the needs_confirmation exceptional return. It does not describe the success return shape (e.g., the created opportunity's id/link), which is the only meaningful gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% with rich per-parameter descriptions, so the baseline would be 3, but the prose adds decision logic the schema lacks: when to anchor workflowPageId, when to set statusName, when experiment applies, and how approvalId follows a needs_confirmation response. It reinforces and contextualizes rather than introducing new syntax.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Record a change') and enumerates concrete examples (fix, chore, new step, experiment), so the agent knows what lands in the board. It does not, however, differentiate itself from close siblings like compass_opportunities_propose or compass_changes_upsert. Clear on what it does, silent on how it differs from neighbors.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit conditions for the key decisions: anchor to a workflow page 'when one fits' (with compass_pages_list), leave unanchored otherwise, and set `experiment` 'only when the user wants it measured against the Ledger.' It also names compass_statuses_list for status selection. No exclusion against the sibling `propose` tool, so it stops short of full routing.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
compass_opportunities_deleteDelete an opportunityADestructiveInspect
Removes an opportunity from the pipeline (soft delete). For an AI proposal the user doesn't want, this is the dismiss verb; for a captured experiment that was a duplicate or a mistake, the remove verb. To conclude a real experiment without evidence, prefer compass_opportunities_set_status REJECTED — that keeps the record. Ids come from compass_opportunities_list. May return needs_confirmation.
| Name | Required | Description | Default |
|---|---|---|---|
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. | |
| approvalId | No | Approval id from a prior needs_confirmation response, after the user has approved. Omit on the first call. | |
| opportunityId | Yes | The change's key (e.g. ACME-12) or id, from compass_opportunities_list. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare destructiveHint=true, and the description adds crucial nuance beyond them: this is a soft delete, and the call may return `needs_confirmation` requiring an approvalId on a follow-up call. That disclosure of the confirmation flow and the soft-delete semantics is exactly the behavioral context annotations cannot convey.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the action and its soft-delete nature, followed by the routing guidance and the confirmation caveat. Every sentence earns its place; no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a destructive-flagged mutation with no output schema, the description covers the operation semantics, when to prefer an alternative, id provenance, and the needs_confirmation return path. Nothing an agent needs to invoke it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the schema already documents all three parameters, making 3 the baseline. The description earns a bump by tying the `needs_confirmation` return to the approvalId parameter and stating that ids come from compass_opportunities_list, adding source-of-value context the schema lacks.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Removes an opportunity from the pipeline') and immediately qualifies the operation as a soft delete. It further disambiguates from sibling operations by naming the two distinct intents — dismiss for AI proposals, remove for duplicates/mistakes — so an agent can tell it apart from compass_opportunities_accept or set_status.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly routes the agent: names the alternative tool (compass_opportunities_set_status REJECTED) and the condition that selects it ('to conclude a real experiment without evidence ... that keeps the record'). It also states where ids originate, leaving nothing to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
compass_opportunities_listList Compass opportunitiesBRead-onlyInspect
The workspace's changes — fixes, chores, new steps, experiments. Each carries statusName (the workspace's own word for where it is) and statusCategory (TRIAGE undecided, BACKLOG not started, ACTIVE in progress, DONE, CANCELED dropped), plus the older lane key in status. experiment: true marks a change measured against the Ledger: ledgerEntryId null means no expectation registered; ledger.verdict carries confirmed/missed after settlement. Scores (value/feasibility/risk) and a workflow anchor are optional. Filter by lane key in status or by workflowPageId; compass_statuses_list has the status names.
| Name | Required | Description | Default |
|---|---|---|---|
| status | No | Filter to one lane. | |
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. | |
| workflowPageId | No | Only opportunities anchored to this workflow page. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover the read-only, non-destructive, closed-world safety profile, so the bar is lower. The description adds real context beyond that: it decodes statusCategory values (TRIAGE/BACKLOG/ACTIVE/DONE/CANCELED), explains the experiment/ledger semantics (ledgerEntryId null = no expectation, verdict = confirmed/missed after settlement), and notes scores/anchor are optional. It stops short of listing behavior such as pagination or ordering.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single dense paragraph that front-loads a data-model definition rather than the tool's action, and is written in heavy shorthand ('the older lane key in `status`'). Almost every clause carries information, but the structure assumes prior domain knowledge and is not scannable.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description usefully enumerates the returned fields (statusName, statusCategory, status, ledger fields, scores, workflow anchor). But for a list tool with 0 required params it omits the default scope, ordering, result limits, and what filtering by lane key actually returns, leaving gaps in how to call it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the schema already documents all three parameters, making 3 the baseline. The description adds only marginal meaning, restating that filtering is by lane key in `status` or workflowPageId and pointing to compass_statuses_list; it does not clarify the enum values (NEW/QUALIFYING/...) mapped to the `status` param.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description identifies the resource ('the workspace's changes — fixes, chores, new steps, experiments') but never states the verb; the listing action is only implied by the trailing 'Filter by lane key...'. It does not distinguish this tool from siblings like compass_changes_upsert or the other compass_*_list tools, so an agent must infer the read/list role from the title alone.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives one concrete usage hint — filter by lane key in `status` or by `workflowPageId` — and routes the agent to compass_statuses_list for status names. However, it never states when to use this list versus create/update/set_status siblings, nor what happens when no filter is supplied (all opportunities? default scope?).
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
compass_opportunities_mark_implementedRecord that the change landedAInspect
The measurement window's boundary: when the user says the experiment's change actually shipped / went live / rolled out, record the landing date. Readings before it are baseline; after it, evidence of effect. Recorded once — it cannot move afterward, so confirm the date. Attaches to the linked Ledger entry as evidence. May return needs_confirmation.
| Name | Required | Description | Default |
|---|---|---|---|
| at | No | ISO date the change landed — omit for today. | |
| note | No | What shipped, if worth recording. | |
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. | |
| approvalId | No | Approval id from a prior needs_confirmation response, after the user has approved. Omit on the first call. | |
| opportunityId | Yes | The change's key (e.g. ACME-12) or id, from compass_opportunities_list. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Goes well beyond the annotations (readOnlyHint=false, destructiveHint=false, openWorldHint=false) by disclosing the irreversible, write-once nature ('cannot move afterward'), the side effect of attaching to the linked Ledger entry as evidence, and the possible `needs_confirmation` response with its approvalId loop.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Five dense sentences that are front-loaded with the trigger and the boundary meaning. The opening label 'The measurement window's boundary:' is slightly abstract, but every sentence carries distinct information and nothing is wasted.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description still covers the mutation's effect, its permanence, the confirmation/approval flow, and the evidence attachment, which is everything an agent needs to invoke it correctly against a 5-param, 100%-covered schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the description adds conceptual meaning to `at` (baseline before, effect evidence after) and ties `approvalId` to the needs_confirmation flow that is not spelled out in the annotated description alone.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource: recording the date the experiment's change shipped, framed as the measurement window boundary. This is clearly distinguishable from sibling mutations like compass_opportunities_set_status, update, or register_expectation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives an explicit trigger ('when the user says the experiment's change actually shipped / went live / rolled out') and a confirmation precondition ('Recorded once — it cannot move afterward, so confirm the date'). It does not name or exclude specific sibling tools, but the when-to-use condition is unmistakable.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
compass_opportunities_proposePropose an AI opportunityAInspect
File an AI-PROPOSED opportunity into the workspace's review queue — the scouting verb (the opportunity-scout routine's main move). Unlike compass_opportunities_create this needs NO approval: the proposal itself is the human gate — it lands in the Compass Inbox and the cockpit's Needs-you for accept/dismiss. Check compass_opportunities_list first so you never duplicate an idea. Anchor to a workflow page when one fits; score value/feasibility/risk 1–5.
| Name | Required | Description | Default |
|---|---|---|---|
| label | Yes | ≤200 chars, the opportunity in one line. | |
| reasoning | No | One or two sentences on why you're proposing this. | |
| riskScore | No | ||
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. | |
| valueScore | No | ||
| description | Yes | What it is and why it's worth trying, grounded in the map. | |
| workflowPageId | No | Workflow page to anchor to, when one fits. | |
| feasibilityScore | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations cover the safety profile (readOnlyHint=false, destructiveHint=false, openWorldHint=false), so the bar is lower. The description still adds real workflow context beyond them: no approval gate, the item lands in the Compass Inbox and the cockpit's Needs-you for accept/dismiss, and the inherent duplicate risk. It omits auth/permission specifics, which is the only real gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Dense and front-loaded: the core action and its routing distinction come first, followed by the guard-rail and the scoring hint. Slightly packed with parentheticals but every clause carries routing or behavioral value.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema and 8 parameters at 63% coverage, the description fills the important gaps: what happens after the call (review queue, accept/dismiss) and the duplicate-check prerequisite. It is complete enough for correct invocation, though workspace/auth edge cases are left to the schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 63%, so the schema does much of the work. The description adds the scoring scale for value/feasibility/risk (1-5) and the anchoring intent for workflowPageId, but says nothing about reasoning or the workspace parameter's token semantics beyond what the schema already documents.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('File an AI-PROPOSED opportunity into the workspace's review queue') and frames it as 'the scouting verb'. It explicitly distinguishes itself from the sibling compass_opportunities_create, so an agent can route correctly without opening schemas.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit when-to-use routing ('Unlike compass_opportunities_create this needs NO approval'), a prerequisite ('Check compass_opportunities_list first so you never duplicate an idea'), and a conditional recommendation ('Anchor to a workflow page when one fits'). This is exactly the when/alternatives guidance expected.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
compass_opportunities_register_expectationPre-register an experiment's expectationAInspect
The honesty mechanism: write what the experiment is expected to change BEFORE evidence exists. Creates a Ledger decision entry and links it to the opportunity — never backfill an expectation to match an outcome. Bind a metric (ledger_metrics_list) + comparator + target when the expectation is measurable; readings then land on the entry as evidence automatically. One expectation per opportunity — revise by superseding in Ledger. May return needs_confirmation.
| Name | Required | Description | Default |
|---|---|---|---|
| target | No | Target value, in the metric's unit. | |
| deadline | No | ISO date the expectation is due by. | |
| metricId | No | Ledger metric to bind (requires target). From ledger_metrics_list. | |
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. | |
| approvalId | No | Approval id from a prior needs_confirmation response, after the user has approved. Omit on the first call. | |
| comparator | No | ||
| expectation | Yes | One sentence — what we expect this to change. | |
| opportunityId | Yes | The change's key (e.g. ACME-12) or id, from compass_opportunities_list. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already mark this as a non-destructive write (readOnlyHint=false, destructiveHint=false), and the description adds substantial context beyond them: it creates a Ledger entry linked to the opportunity, forbids backfilling, auto-attaches readings as evidence, enforces a single expectation, and discloses the possible `needs_confirmation` return plus the supersede-to-revise workflow.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three dense sentences, front-loaded with the core purpose and the honesty rationale, then workflow and constraints. Every clause carries information, though the density leaves little breathing room and some cross-references (ledger_metrics_list) are embedded mid-sentence.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For an 8-parameter mutation tool with 88% schema coverage and no output schema, the description covers the return signal (needs_confirmation), the evidence-linking behavior, and the one-per-opportunity rule. It leaves workspace-scoping and the comparator enum semantics to the schema, which is reasonable given the high coverage.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 88% schema coverage the schema already documents parameters, but the description adds genuine meaning: metricId must be bound with comparator + target for the expectation to be measurable, and readings then flow in as evidence. It also implicitly ties the approvalId flow to the needs_confirmation response.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource (register/pre-register an experiment's expectation) and frames it as 'the honesty mechanism' with the concrete action 'Creates a Ledger decision entry and links it to the opportunity.' It distinguishes itself from siblings by the one-per-opportunity constraint and the superseding revision path.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear when-to-use ('write what the experiment is expected to change BEFORE evidence exists', 'never backfill'), a hard constraint ('One expectation per opportunity — revise by superseding in Ledger'), and a conditional path for measurable expectations. It references ledger_metrics_list but does not explicitly contrast with siblings like compass_opportunities_update for non-expectation edits, so it falls just short of explicit alternatives.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
compass_opportunities_set_statusMove a change to another statusADestructiveInspect
Move a change to one of the workspace's statuses, by name (compass_statuses_list has them — e.g. "In review"). The old lane keys (NEW / QUALIFYING / BACKLOG / EXPERIMENTING / SETTLED / REJECTED) still work and land on that lane's default status. Only changes marked as experiments follow the Ledger: when an experiment enters an In progress status with no registered expectation, offer compass_opportunities_register_expectation; when it enters Done with its Ledger entry still open, offer ledger_entries_settle. Other changes just move. May return needs_confirmation.
| Name | Required | Description | Default |
|---|---|---|---|
| status | No | A lane key. Use this or `statusName`. | |
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. | |
| approvalId | No | Approval id from a prior needs_confirmation response, after the user has approved. Omit on the first call. | |
| statusName | No | A status name from compass_statuses_list. Use this or `status`. | |
| opportunityId | Yes | The change's key (e.g. ACME-12) or id, from compass_opportunities_list. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare destructiveHint=true, and the description complements this by disclosing the legacy lane-key fallback behavior, the experiment-only Ledger coupling, and that the call "May return needs_confirmation" (tying to the approvalId param). It doesn't explain reversibility or what a status change does to dependent data, so it stops short of a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the core action and the name-based addressing before the conditional Ledger guidance, and every sentence carries operative information. It is dense-to-long, but nothing is padding.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a mutation tool with no output schema, the description covers the addressing options, the experiment-conditional side effects, and the confirmation return shape. It leaves the mutation's side effects on the change itself unstated, but the essential call-time context is present.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 100% schema description coverage, the schema already documents every parameter, including the enum values and the status/statusName alternative. The description adds only marginal value by naming example lane keys and default-status landing behavior, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ("Move") and resource ("a change") with the target (one of the workspace's statuses), and clarifies the two accepted addressing modes. An agent can distinguish it from compass_opportunities_update and compass_opportunities_mark_implemented without opening a schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly routes the agent: use compass_statuses_list for names, offer compass_opportunities_register_expectation when an experiment enters In progress without an expectation, and ledger_entries_settle when it enters Done with an open Ledger entry. It also states the negative case ("Other changes just move").
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
compass_opportunities_updateEdit an opportunityADestructiveInspect
Corrects an opportunity's content: label, description, priority, what done means, the 1–5 value / feasibility / risk scores, the owner (a PERSON page), notes, tags, workflow anchor, external link, visibility. Fields you omit are untouched; the lane is NOT here — move it with compass_opportunities_set_status. Opportunity ids come from compass_opportunities_list. May return needs_confirmation — say what you're changing and wait.
| Name | Required | Description | Default |
|---|---|---|---|
| tags | No | New tag set — replaces existing tags. | |
| label | No | New one-line label (2–200 chars). | |
| assignee | No | The workspace member doing the work, by email or user id; null unassigns. They're notified. | |
| doneWhen | No | What's true when it's done, that someone else could check (≤1,000 chars). | |
| priority | No | How soon it matters; null clears it, and the map's suggestion shows again. | |
| riskScore | No | ||
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. | |
| approvalId | No | Approval id from a prior needs_confirmation response, after the user has approved. Omit on the first call. | |
| experiment | No | Measure it against the Ledger (expectation when it starts, verdict when it's done). | |
| valueScore | No | ||
| visibility | No | Who can see it: PRIVATE (only the user), WORKSPACE (every member, the default), or SHARED (specific people, granted afterwards). Say 'make it private' → PRIVATE. | |
| description | No | New description (≤5,000 chars). | |
| externalUrl | No | Where the doing happens off-platform (URL); null clears. | |
| ownerPageId | No | PERSON page id who owns the work; null clears. | |
| opportunityId | Yes | The change's key (e.g. ACME-12) or id, from compass_opportunities_list. | |
| workflowPageId | No | Workflow page to anchor to; null un-anchors. | |
| feasibilityScore | No | ||
| qualitativeNotes | No | Free-form notes (≤20,000 chars) — replaces the current notes. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations supply destructiveHint=true and readOnlyHint=false, so the safety profile is covered; the description adds genuinely non-obvious behavior: partial-update semantics, the needs_confirmation/approvalId loop, and that an assignee triggers a notification. It does not explicitly warn that tags/notes replace rather than merge, which the destructive annotation implies but the prose only hints at.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two dense sentences front-loaded with the core action and the mutable-field list, followed by the constraint/alternative and the confirmation flow. No filler; every clause carries operational information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For an 18-parameter mutation with no output schema, the description covers id provenance, the confirmation protocol, and sibling routing, which is most of what an agent needs. It stops short of documenting return values or the exact consequences of null-clearing fields, but the essentials are present.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 83%, so most parameters are already self-documented. The description restates a handful of fields (owner is a PERSON page, 1–5 scores) that the schema largely covers, adding little syntax or format detail beyond it. Baseline 3 is appropriate when the schema carries the load.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Starts with a specific verb+resource ('Corrects an opportunity's content') and enumerates exactly which fields are editable, then explicitly carves out what is NOT editable ('the lane is NOT here — move it with compass_opportunities_set_status'). An agent can distinguish this from sibling updaters without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives the partial-update contract ('Fields you omit are untouched'), routes the lane change to the correct sibling, tells where opportunity ids come from, and explains the needs_confirmation handshake including the instruction to pause and confirm. Multiple when/when-not conditions, nothing left implicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
compass_page_links_listList a page's connectionsARead-onlyInspect
All typed edges touching one page, both directions, each hydrated with the other endpoint's page summary. Use this to understand a subject's neighborhood before adding to it.
| Name | Required | Description | Default |
|---|---|---|---|
| pageId | Yes | The page id. | |
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare the safe read-only profile, so the description's real contribution is disclosing the result shape: bidirectional edges each hydrated with the other endpoint's page summary. That return-format detail is genuinely useful and absent from annotations, though nothing is said about volume limits or ordering.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, zero filler, with the scope of what is returned front-loaded before the usage hint.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a two-parameter read tool with fully documented schema and no output schema, the description covers what it returns and when to reach for it. The only gap is any mention of result size or pagination for a page with many edges.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and both parameters (pageId, workspace, including the workspace default/override semantics) are documented there. The description adds no syntax or format meaning beyond the schema, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (list) plus the exact resource and scope: 'all typed edges touching one page, both directions.' Clearly separates it from the write siblings (compass_links_create/update/delete) without needing to name them.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
'Use this to understand a subject's neighborhood before adding to it' supplies a clear context and implicitly routes the agent to the link-mutation siblings for the 'adding' step, but it names no alternative tool explicitly and gives no when-not-to-use condition.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
compass_pages_createCreate a Compass pageAInspect
Creates a typed page in the workspace's Compass compendium. Check compass_pages_list first — don't create a page for a subject the map already has; link to it instead. May return needs_confirmation — if so, tell the user what you're proposing, wait for their approval, then re-call with the approvalId. Only record what the human has actually told you.
| Name | Required | Description | Default |
|---|---|---|---|
| body | No | Markdown body shaped to the type (steps for WORKFLOW, role/contact sections for PERSON, etc.). | |
| tags | No | Optional tags. | |
| type | Yes | Page type for this subject. VALUE = something the business won't sacrifice; PRIORITY = an outcome it's pushing toward; AUDIENCE = who the work is for; OFFERING = what the business delivers. | |
| title | Yes | ≤80 chars, sentence case, no filler verbs. | |
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. | |
| approvalId | No | Approval id from a prior needs_confirmation response, after the user has approved. Omit on the first call. | |
| visibility | No | Who can see it: PRIVATE (only the user), WORKSPACE (every member, the default), or SHARED (specific people, granted afterwards). Say 'make it private' → PRIVATE. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only declare it is a non-destructive write. The description adds real behavioral context beyond them: a possible `needs_confirmation` return, the requirement to surface the proposal to the user and re-call with an `approvalId`, and a grounding constraint ('Only record what the human has actually told you').
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four tightly packed sentences, front-loaded with purpose then routing, confirmation handling, and a grounding rule. No filler and every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a create tool with no output schema, the description covers the creation semantics, the dedup prerequisite, the confirmation handshake, and the data-grounding constraint. Nothing an agent needs to call it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all 7 parameters including type values, workspace rules, visibility enum, and approvalId. The description reinforces the approvalId round-trip but adds little parameter detail beyond what the schema provides; baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Creates a typed page in the workspace's Compass compendium') and distinguishes itself from siblings by naming `compass_pages_list` and the alternative action (link instead of create). An agent can identify the tool without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicit when-not guidance ('don't create a page for a subject the map already has; link to it instead') plus a named prerequisite (`compass_pages_list`). It also spells out the confirmation flow and what to do when `needs_confirmation` is returned.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
compass_pages_deleteDelete a Compass pageADestructiveInspect
Moves a page to the trash (soft delete — its connections and anchored opportunities go with it, and compass_pages_restore brings the whole set back). Use this to REMOVE a page you or an extraction created wrongly, or one the user says no longer belongs on the map; to fix a wrong title or body, use compass_pages_update instead. Page ids come from compass_pages_list / compass_pages_search. May return needs_confirmation — name the page and wait for approval.
| Name | Required | Description | Default |
|---|---|---|---|
| pageId | Yes | Id of the page to delete. | |
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. | |
| approvalId | No | Approval id from a prior needs_confirmation response, after the user has approved. Omit on the first call. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare destructiveHint=true but say nothing about scope or recovery; the description fills both gaps by stating the cascade (connections and anchored opportunities go with it) and the reversal path via compass_pages_restore. It also discloses the interactive approval behavior (may return `needs_confirmation`, name the page and wait), which is non-obvious runtime behavior not captured anywhere in structured fields.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the operation and its soft-delete semantics before the usage routing. Every sentence does distinct work: mechanism, cascade, reversal, use-case, counter-case, id provenance, confirmation flow.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a destructive, confirmation-gated tool with no output schema, the description covers mutation scope, reversibility, the required interaction loop, and prerequisite id lookup. An agent has everything needed to call it correctly, including knowing it may have to pause for user approval.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3 and the schema already carries per-parameter text. The description still adds meaning the schema does not: where pageId values originate, and the approval protocol that gives approvalId its purpose.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Moves a page to the trash') and immediately qualifies the mechanism as a soft delete, so the agent knows this is recoverable rather than permanent. It explicitly distinguishes itself from compass_pages_update and points at compass_pages_restore, so no sibling ambiguity remains.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives an explicit when-to-use ('a page you or an extraction created wrongly, or one the user says no longer belongs') and an explicit when-not ('to fix a wrong title or body, use compass_pages_update instead'). It also routes the agent to the id source (compass_pages_list / compass_pages_search) before the call is made.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
compass_pages_getGet a Compass pageARead-onlyInspect
Fetches one Compass page by id, including its full Markdown body, header fields, and any attached Napkin sketches — view an attached sketch's actual drawing with napkin_boards_view (its boardId) before discussing it.
| Name | Required | Description | Default |
|---|---|---|---|
| pageId | Yes | The page id (from compass_pages_list). | |
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, destructiveHint=false, and openWorldHint=false, so safety is covered. The description adds value beyond that by disclosing the return contents (Markdown body, header fields, sketches) and a downstream workflow dependency, though it says nothing about pagination, size limits, or missing-page behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One front-loaded sentence that names the resource and then the payload, with the napkin hint appended as a useful aside. Slightly dense but no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description usefully compensates by listing the returned contents, and annotations cover the safety profile. Remaining gaps (error/not-found behavior, pagination of attached sketches) are minor for a single-resource fetch.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, and the schema itself documents both pageId (with its source, compass_pages_list) and the workspace slug rules. The description only restates 'by id', so it adds essentially nothing over the structured fields — baseline 3.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Fetches one Compass page by id') and enumerates the payload ('full Markdown body, header fields, and any attached Napkin sketches'), clearly distinguishing it from the sibling list/search/update tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit cross-tool routing: view an attached sketch's actual drawing with napkin_boards_view (its boardId) before discussing it. However, it never states when to prefer this over compass_pages_list or compass_pages_search, so guidance is contextual rather than complete.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
compass_pages_listList Compass pagesARead-onlyInspect
Lists pages in the active workspace's Compass compendium. Optional type filter narrows to one node type (WORKFLOW, PERSON, SYSTEM, PAIN_POINT, DOCUMENT, VALUE, PRIORITY, AUDIENCE, OFFERING). Returns summaries, 50 at a time (limit / offset, total and nextOffset in the result) — fetch one with compass_pages_get for the full body. When you know what you're looking for, compass_pages_search is the better first call.
| Name | Required | Description | Default |
|---|---|---|---|
| type | No | Narrow to a specific page type (WORKFLOW, PERSON, SYSTEM, PAIN_POINT, DOCUMENT, VALUE, PRIORITY, AUDIENCE, OFFERING). | |
| limit | No | Page size, default 50. Prefer compass_pages_search when you know what you're looking for. | |
| offset | No | Skip this many, for the next page. | |
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, destructiveHint=false, and openWorldHint=false, so the safety profile is covered. The description adds real behavioral context beyond that: page size of 50, `limit`/`offset` paging, and the presence of `total` and `nextOffset` in the result. It stops short of describing auth or rate-limit behavior, which the schema carries for `workspace`.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads what it lists and the workspace scope, then layers filtering, pagination, and sibling routing in three tight sentences. Nothing is padding, and the most decision-relevant routing advice is easy to find.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, but the description discloses the return shape (summaries, total, nextOffset) and points to `compass_pages_get` for full bodies, which is everything an agent needs to page and follow up correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so every parameter is already documented in the schema, establishing a baseline of 3. The description reinforces the `type` enum and pagination parameters but adds no syntax or format meaning beyond what the schema provides.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource (lists pages in the active workspace's Compass compendium) plus scope. It explicitly routes to sibling tools (`compass_pages_get` for full bodies, `compass_pages_search` as a better first call), so an agent can differentiate it without opening other schemas.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit alternatives with conditions: use `compass_pages_get` when you have a page and want the full body, and use `compass_pages_search` when you already know what you're looking for. This is concrete when-to-use guidance, not inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
compass_pages_restoreRestore a deleted Compass pageAInspect
Puts a trashed page back on the map, together with exactly the connections and opportunities its delete took. Use when the user wants a deleted page back (the pageId from the earlier compass_pages_delete, or from the Compass trash). Fails with not_found when the page is already live or never existed. May return needs_confirmation.
| Name | Required | Description | Default |
|---|---|---|---|
| pageId | Yes | Id of the deleted page. | |
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. | |
| approvalId | No | Approval id from a prior needs_confirmation response, after the user has approved. Omit on the first call. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only cover the safety profile (readOnlyHint=false, destructiveHint=false, openWorldHint=false); the description goes well beyond them by disclosing that the restore is not just the page but also its connections and opportunities, that a not_found failure mode exists for already-live/nonexistent pages, and that a needs_confirmation response is possible. That last point is critical behavioral context an agent must plan for.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight sentences, front-loaded with the core action and scope before the trigger, parameter sourcing, and failure modes. No filler; every clause carries information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description carries the return-value burden and covers the failure (not_found) and confirmation (needs_confirmation) paths, which is the important part. It does not describe the success payload beyond the restored-page implication, a minor gap for a restore operation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so baseline is 3. The description exceeds it by explaining the workflow semantics behind approvalId ('from a prior needs_confirmation response, after the user has approved') and pageId's origin, tying the parameters into the confirmation loop rather than just restating them.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb+resource ('Puts a trashed page back on the map') and adds scope detail that distinguishes it from compass_pages_delete: it restores the page *plus* the connections and opportunities the delete removed. An agent can tell this apart from compass_pages_get, compass_pages_list, or compass_pages_delete without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicit trigger ('Use when the user wants a deleted page back') and it even tells the agent where the pageId comes from (a prior compass_pages_delete, or the Compass trash). The failure condition ('not_found when the page is already live or never existed') effectively defines the when-not-to-call case, though no alternative tool is named for that situation.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
compass_pages_searchSearch Compass pagesARead-onlyInspect
Keyword search over page titles, bodies, and tags (case-insensitive substring match), paginated. The fast way to check whether a subject already has a page before creating or linking. Returns summaries with a 300-char body snippet — fetch full text with compass_pages_get.
| Name | Required | Description | Default |
|---|---|---|---|
| q | Yes | Search term. Use the most distinctive word or phrase for the subject ('Zendesk', 'invoice approval') — not full sentences. | |
| type | No | Narrow to one page type. | |
| limit | No | Results per page, 1-25. Default 10. | |
| offset | No | Pagination offset. Default 0. | |
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly/non-destructive/openWorld=false, so the safety profile is covered; the description adds real value by disclosing match semantics (case-insensitive substring), that results are paginated summaries with a 300-char body snippet, and the handoff to compass_pages_get. It does not state result ordering or how snippet truncation affects matching, which is the only remaining gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, all load-bearing: match semantics first, then the routing use case, then the return shape and follow-up tool. No filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description compensates by specifying the return shape (summaries with 300-char snippets), the pagination contract, and the fetch-full-text alternative. For a 5-parameter read tool with full schema coverage and clean annotations, nothing an agent needs is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, including the useful 'use the most distinctive word, not full sentences' guidance for q, the limit range, and the workspace-auth caveat, so the schema does the heavy lifting. The description only restates the keyword/pagination concept, adding nothing beyond it; baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (keyword search) and resource (Compass pages), plus the exact match scope: titles, bodies, tags, case-insensitive substring. An agent can distinguish it from compass_pages_list (unfiltered listing) and compass_pages_get (full text retrieval) without opening a schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear context — 'the fast way to check whether a subject already has a page before creating or linking' — and names compass_pages_get as the follow-up for full text. It does not explicitly say when to prefer it over compass_pages_list or search_docs/search_workspace, so it stops short of full when/when-not routing.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
compass_pages_updateUpdate a Compass pageADestructiveInspect
Edits an existing page's title, description, body, tags, header fields, or visibility — use this to FIX what you (or an extraction) got wrong instead of creating a duplicate. Fetch the current page with compass_pages_get first and preserve what the human wrote; title / body / tags replace the field wholesale, while headerFields MERGES over the current ones (send only the keys you're changing — 'set the owner' leaves status alone). Page ids come from compass_pages_list. May return needs_confirmation — tell the user what you're changing, wait for approval, then re-call with the approvalId.
| Name | Required | Description | Default |
|---|---|---|---|
| body | No | New Markdown body — replaces the whole body. | |
| tags | No | New tag set — replaces existing tags. | |
| title | No | New title (≤200 chars). | |
| pageId | Yes | Id of the page to edit. | |
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. | |
| approvalId | No | Approval id from a prior needs_confirmation response, after the user has approved. Omit on the first call. | |
| visibility | No | Who can see it: PRIVATE (only the user), WORKSPACE (every member, the default), or SHARED (specific people, granted afterwards). Say 'make it private' → PRIVATE. | |
| description | No | New one-line gallery summary (≤2,000 chars). | |
| headerFields | No | Partial header fields for the page's type (PERSON: role, email, …; WORKFLOW: owner, status, …) — merged over the current values; the server validates the merged result against the type. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare destructiveHint=true / readOnly=false, and the description meaningfully adds to that: it explains the replace-wholesale vs headerFields-merge semantics, the need to preserve human-authored content, and the needs_confirmation → approvalId round-trip. Some of the replace language overlaps the schema, but the merge/confirmation behavior is extra context the annotations don't carry.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the verb and target, then packed with routing, prerequisite, and merge/confirmation guidance in which every sentence earns its place — no filler or restated schema text.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 9-parameter destructive edit with a nested object and no output schema, the description covers the risky behaviors an agent needs (field replacement vs merge, approval gating, id provenance, preservation of human content), leaving nothing critical to inference.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so baseline is 3; the description goes further by contrasting replace semantics (title/body/tags) against merge semantics for headerFields with a concrete example ('set the owner' leaves status alone) and by explaining where pageId values originate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Edits) and resource (an existing page), enumerates the editable fields, and explicitly distinguishes itself from the create path ('instead of creating a duplicate'). An agent can separate it from compass_pages_create and compass_pages_get without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit when-to-use (fixing mistakes/extra value from extractions rather than duplicating), prerequisites (fetch with compass_pages_get first, page ids from compass_pages_list), and the confirmation workflow to follow when needs_confirmation is returned.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
compass_statuses_listList the workspace's change statusesARead-onlyInspect
The statuses this workspace's changes move through, in board order — the names people use. Each sits in a group: Undecided (TRIAGE), Not started (BACKLOG), In progress (ACTIVE), Done (DONE), Dropped (CANCELED). Groups are stable across workspaces; names aren't, so use the names when talking to people and the groups when reasoning about progress. Move a change with compass_opportunities_set_status.
| Name | Required | Description | Default |
|---|---|---|---|
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and destructiveHint=false, so safety is covered. The description adds genuine domain context beyond structured data: the five group taxonomy, the fact that groups are stable while names are not, and the ordering guarantee ('in board order'). It does not discuss return format or pagination, which is acceptable for a small list tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the identity of the resource, then layers in the taxonomy and the stable-vs-unstable distinction. Four sentences, all informative, though the group enumeration and cross-reference sentence make it slightly denser than strictly necessary for a one-parameter list tool.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema, so the description carries the burden of conveying what comes back; it does this well by explaining names, groups, and ordering. Lacking only explicit return-shape/pagination detail, which is minor for a simple status lookup.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and the single workspace parameter is fully documented in the schema (default-workspace and API-key behavior). The description adds nothing about the parameter, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource — listing the workspace's change statuses — and uniquely characterizes them as the values changes 'move through, in board order.' It also names the sibling that mutates them (compass_opportunities_set_status), so an agent can place it precisely among the many compass_* tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives practical guidance for how to consume the result (use names with people, groups for reasoning) and points to compass_opportunities_set_status for the corresponding mutation. It stops short of an explicit 'call this when you need X' trigger or any when-not, but the context is clear.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
entity_tags_browseBrowse the workspace's tagsARead-onlyInspect
Without a tag: every tag in use across the workspace with how many entities carry it, most-used first — the vocabulary the team already organizes by. With a tag: everything filed under it across every tool, each with its kind, id, title, and path. Use it to reuse existing labels instead of inventing near-duplicates, and to answer 'show me everything about X' when X is a label.
| Name | Required | Description | Default |
|---|---|---|---|
| tag | No | A tag to expand into its items. Omit to list tags. | |
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly/destructive/openWorld, so safety is covered, and the description adds real behavioral detail: result ordering (most-used first), entity counts per tag, and the per-item fields returned in tag mode (kind, id, title, path). No pagination or size-limit behavior is mentioned, keeping this from a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded and cleanly parallel: 'Without a tag:' then 'With a tag:' then the usage clause. Every sentence carries distinct information with no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description correctly carries the return-value burden for both modes, and annotations cover the safety profile while the schema covers both parameters. Nothing an agent needs to select or call this tool is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the description adds meaning beyond the schema by characterizing what each mode of the tag parameter actually returns, turning a bare 'omit to list tags' into the two distinct result shapes. The workspace parameter semantics remain entirely schema-borne.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific dual operation (list every tag in the workspace, or expand one tag into all entities filed under it) with the exact resource and scope. It is immediately distinguishable from the write-oriented siblings entity_tags_get and entity_tags_set because the browse mode and cross-tool aggregation are spelled out.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives concrete when-to-use guidance: reuse existing labels rather than inventing near-duplicates, and answer 'show me everything about X' when X is a label. It stops short of naming sibling alternatives or stating when-not-to-use (e.g. versus search_workspace), so it is clear context without explicit exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
entity_tags_getRead the tags on entitiesARead-onlyInspect
Returns the tags on a batch of entities of one kind — the labels galleries organize by. Ids come from the kind's list/get tool or from search_workspace. Use it before entity_tags_set so you replace the full set knowingly, and to answer 'what is this filed under'. Entities the user can't see are omitted.
| Name | Required | Description | Default |
|---|---|---|---|
| ids | Yes | ||
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. | |
| entityKind | Yes | Which kind the ids belong to. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and destructiveHint=false, so safety is covered; the description adds the non-obvious behavioral fact that 'Entities the user can't see are omitted,' i.e. results are permission-filtered rather than erroring. It does not mention limits (e.g. the 100-id cap or ordering), but the value-add beyond annotations is genuine.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, all front-loaded: what it returns first, then usage, then the permission caveat. No filler and each sentence carries distinct information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description carries the return-value burden and does state what comes back (tags) and the omission behavior. It leaves minor gaps — tag value format, ordering, and whether unknown ids are silently dropped or error — but nothing that blocks correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 67%, so the schema documents most params, but the description adds real provenance for the hardest parameter: ids 'come from the kind's list/get tool or from search_workspace.' That tells an agent where to obtain valid ids, which the bare array schema does not.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource with scope: 'Returns the tags on a batch of entities of one kind.' The parenthetical 'the labels galleries organize by' distinguishes tags from other entity metadata, and the named siblings (entity_tags_set, search_workspace) make the boundary clear.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit usage contexts: 'Use it before entity_tags_set so you replace the full set knowingly, and to answer what is this filed under.' That is a real when-to-use plus a stated alternative. It does not mention entity_tags_browse, the other obvious read-side sibling, so it falls short of fully disambiguating the read alternatives.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
entity_tags_setSet an entity's tagsADestructiveInspect
Replaces the FULL tag set on one entity (an empty list clears it). Read the current tags with entity_tags_get first and pass the merged list — this is not additive. Tags are lowercase letters, numbers, spaces, and hyphens; prefer labels already in use (entity_tags_browse) so the workspace's vocabulary stays small. The id comes from the kind's list/get tool or search_workspace.
| Name | Required | Description | Default |
|---|---|---|---|
| tags | Yes | ||
| entityId | Yes | ||
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. | |
| approvalId | No | Approval id from a prior needs_confirmation response. Omit on the first call. | |
| entityKind | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true and readOnlyHint=false, and the description reinforces this with the replace-not-merge semantics and the empty-list-clears behavior. It does not mention the needs_confirmation/approvalId retry flow that the schema implies, which is the one notable behavioral gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three dense sentences, all front-loaded with the destructive replace semantics first, followed by workflow and vocabulary guidance. Every sentence earns its place; the parentheticals make it slightly busy but not wasteful.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, and the description covers destructive semantics, prerequisites, id sourcing, and tag format well enough to invoke the tool correctly. The confirmation/approval retry path is left to the schema field description rather than the tool description.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is only 40%, and the description compensates by documenting the tag character set (lowercase letters, numbers, spaces, hyphens), the merge requirement for the tags array, and the origin of entityId. The workspace and approvalId parameters are only explained in the schema, so coverage is good but not complete.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Replaces the FULL tag set on one entity') and immediately clarifies the destructive scope with '(an empty list clears it)'. This cleanly separates it from the sibling readers entity_tags_get and entity_tags_browse.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit workflow: read current tags with entity_tags_get first, pass the merged list, and warns 'this is not additive'. Also routes to entity_tags_browse for vocabulary reuse and names where the entityId comes from (kind's list/get tool or search_workspace), so the agent knows both prerequisites and alternatives.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_docFetch ZeroWidth doc by slugARead-onlyInspect
Fetch the full Markdown body of a specific docs page by its slug. Use this after search_docs when the user needs the complete content of a page. No authentication required.
| Name | Required | Description | Default |
|---|---|---|---|
| slug | Yes | Page slug. Accepts 'compass/api', '/compass/api', or 'docs/compass/api'. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, openWorldHint=false, and destructiveHint=false, covering the safety profile. The description adds valuable context by stating 'No authentication required' and specifying the return body as full Markdown, which agents need to know beyond the structured hints.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, front-loaded with the core purpose ('Fetch the full Markdown body...'), followed by usage routing and a salient behavioral note. Every sentence earns its place with no redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple one-parameter read tool with full schema coverage and annotations covering safety, the description is complete: it states the return format, the required input type, usage context, and authentication status. Nothing an agent needs to invoke it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, and the single `slug` parameter is fully documented in the schema with accepted formats. The description adds no syntax or format details beyond what the schema provides, so it meets the baseline rather than exceeding it.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Fetch'), resource ('full Markdown body of a specific docs page'), and retrieval key ('by its slug'). It distinguishes itself from the search-oriented sibling by positioning as the step after `search_docs` for complete page content.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly says to use this after `search_docs` when the user needs the complete content of a page, which names the alternative and the condition that selects it. It does not state when not to use it (e.g., for listing or metadata), but the context is clear enough for a simple read tool.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_docsList ZeroWidth docs pagesARead-onlyInspect
Enumerate all available docs pages, optionally filtered by product (e.g. 'compass', 'legal', 'overview'). Use this to discover what slugs exist before calling get_doc. No authentication required.
| Name | Required | Description | Default |
|---|---|---|---|
| product | No | Optional product slug filter (e.g. 'compass', 'legal', 'overview'). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, destructiveHint=false, and openWorldHint=false, so the safety profile is covered. The description adds context beyond them with 'No authentication required,' a genuinely useful operational fact for callers, though it says nothing about pagination or result size for a full enumeration.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight sentences with purpose, routing guidance, and the auth fact front-loaded in order of importance. No filler or repetition of the title.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only, single-optional-param listing tool whose annotations cover safety, the description is nearly sufficient. Without an output schema it could note the return shape (e.g. that results are slugs/pages), but the 'slugs' reference largely covers that, so only a small gap remains.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the schema already documents the single 'product' parameter. The description's example values ('compass', 'legal', 'overview') duplicate the schema description verbatim, adding no meaning beyond it. Baseline 3 applies when the schema does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Enumerate') and resource ('all available docs pages') with scope, plus the optional product filter. It names the sibling get_doc and frames itself as the discovery step before retrieval, letting an agent distinguish it from single-doc fetch without opening either schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly states when to use it: 'to discover what slugs exist before calling get_doc,' giving a clear dependency flow and naming the alternative. It does not address when NOT to use it or whether search_docs is a better discovery path, leaving a minor gap.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
search_docsSearch ZeroWidth docsARead-onlyInspect
Search ZeroWidth product documentation. Returns matching pages with title, slug, public URL, and a query-relevant snippet. Use this when the user asks about a ZeroWidth product (Compass, Workbench, Caliper, Prism, Ledger, Napkin, zv1), an API behavior, or a policy. No authentication required — the docs corpus is public.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | Max number of results. Defaults to 10. | |
| query | Yes | Search query — keywords or natural-language phrase. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, destructiveHint=false, and openWorldHint=false, so the safety profile is covered. The description adds useful non-annotation context: 'No authentication required — the docs corpus is public' and the shape of returned results (title, slug, public URL, snippet). It does not mention pagination or ordering, but it goes beyond the structured fields.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is four short sentences, front-loaded with the core action, followed by return information, usage trigger, and auth note. Every sentence earns its place, and there is no repetition or filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a two-parameter read-only search tool, the description covers purpose, return format, auth requirements, and usage triggers. It does not explain how this differs from search_workspace or list_docs/get_doc, and it lacks result-ordering or empty-result behavior, but it is otherwise sufficient for correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, and both parameters are documented in the schema itself. The description says the query returns a 'query-relevant snippet' but adds no syntax, format, or constraint details beyond what the schema already provides, so the baseline of 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: 'Search ZeroWidth product documentation.' It also names the return fields and the product scope with concrete examples (Compass, Workbench, Caliper, etc.), which lets an agent distinguish it from siblings like list_docs, get_doc, and search_workspace without opening schemas.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives clear usage context: 'Use this when the user asks about a ZeroWidth product ..., an API behavior, or a policy.' However, it does not name alternative tools (search_workspace, list_docs, get_doc) or state when not to use this tool, so it stops short of explicit routing guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
search_workspaceSearch the whole workspaceARead-onlyInspect
Finds entities across every tool by name in one call — Workbench flows, Compass pages, Caliper datasets, evals, rubrics, reviews, specs and sources (apps sending agent traces), Ledger entries, Napkin sketches and decks. Use it FIRST when the user names something without saying where it lives ('the onboarding flow', 'that invoice page'); reach for a tool's own list only when you already know the tool. Each hit carries its id, kind, and workspace-relative path, so the id feeds the matching *_get tool and the path makes a link. Results only include what the user can see, and only kinds this token may read.
| Name | Required | Description | Default |
|---|---|---|---|
| q | Yes | Case-insensitive substring matched against names/titles. | |
| kinds | No | Restrict to these kinds (flow, page, dataset, eval, entry, board). Omit to search everything. | |
| limit | No | ||
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and destructiveHint=false, so safety is covered. The description adds genuinely new behavioral context beyond that: results are filtered to what the user can see and to kinds the token may read, and it discloses the hit shape (id, kind, workspace-relative path).
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three dense, front-loaded sentences with no filler: purpose and coverage first, routing second, return/scope semantics last. Every clause carries information the agent can act on.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, so the description carries the return-value burden and does so by describing the id/kind/path tuple and how the id feeds *_get tools. Combined with the permission scoping note, an agent has everything needed to call and use this correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 75%, so most parameters are already documented structurally. The description reinforces name/title matching but adds no new syntax or format detail for kinds, limit, or workspace, so the baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Opens with a specific verb (Finds) and resource (entities across every tool by name), then enumerates the concrete kinds covered (flows, pages, datasets, evals, rubrics, etc.). An agent can immediately distinguish this cross-tool search from the many per-tool list siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly states when to use it ('Use it FIRST when the user names something without saying where it lives'), gives concrete examples ('the onboarding flow'), and names the alternative plus its selection condition ('reach for a tool's own list only when you already know the tool'). This is textbook when/when-not/alternative routing.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
44 tool updates
- First observed
comments_create - First observed
comments_list - First observed
comments_resolve - First observed
compass_changes_upsert - First observed
compass_gaps_create - First observed
compass_gaps_list - First observed
compass_gaps_resolve - First observed
compass_gaps_update - First observed
compass_inbox_list - First observed
compass_interview_invite - First observed
compass_interview_targets - First observed
compass_interviews_create - First observed
compass_interviews_get - First observed
compass_interviews_list - First observed
compass_interviews_revoke - First observed
compass_links_create - First observed
compass_links_delete - First observed
compass_links_update - First observed
compass_map_view - First observed
compass_opportunities_accept - First observed
compass_opportunities_create - First observed
compass_opportunities_delete - First observed
compass_opportunities_list - First observed
compass_opportunities_mark_implemented - First observed
compass_opportunities_propose - First observed
compass_opportunities_register_expectation - First observed
compass_opportunities_set_status - First observed
compass_opportunities_update - First observed
compass_page_links_list - First observed
compass_pages_create - First observed
compass_pages_delete - First observed
compass_pages_get - First observed
compass_pages_list - First observed
compass_pages_restore - First observed
compass_pages_search - First observed
compass_pages_update - First observed
compass_statuses_list - First observed
entity_tags_browse - First observed
entity_tags_get - First observed
entity_tags_set - First observed
get_doc - First observed
list_docs - First observed
search_docs - First observed
search_workspace
Related MCP Connectors
Read and manage the competitors, signals, and weekly briefs in your probr workspace.
- SotoGoalOAuthcom.sotogoal
Goal-centered work: read and write goals, tasks, notes and attachments, and decide what to do next.
- PanttsOAuthcom.pantts
Read and manage your Pantts boards — kanban, gantt, table and graph over one shared card model.
Read and update your standards.new workspace: records, documents, search and schema.
Related MCP Servers
- AlicenseNot gradedqualityBmaintenanceEnables managing work entries, statuses, cycles, and free-form notes between meetings, with tools for creating, updating, deleting, and presenting meeting markdown summaries.37 npmMIT
- AlicenseNot gradedqualityCmaintenanceEnables reading, writing, and searching a local markdown-based memory store using MCP tools, with safety checks and index consistency.MIT
- AlicenseNot gradedqualityAmaintenanceNotion-like workspace of pages and customizable databases with a remote MCP server (OAuth 2.1 + PKCE, scoped read/write tokens, full audit log). 14 tools to search, read, and write pages, database rows, and database schemas.74AGPL 3.0
- AlicenseAqualityAmaintenanceA local workspace whose accessible pages and MCP tools are generated from the same definitions, so AI agents can do everything a person can: create, update and find notes, tasks and projects, build and arrange pages, and read any page back as a screen reader gets it.36MIT
Glama MCP Gateway
Add one secure layer between your agents and this server.