Skip to main content
Glama

ZeroWidth Ledger

Server Details

Log decisions with expectations, add evidence, and track metrics in Ledger.

If you are the author of this connector, you can claim ownership by verifying the domain or GitHub account it belongs to. Claimed connector authors can inspect health checks, view analytics, and manage their listing.
Status
Healthy
Last Tested
Transport
Streamable HTTP · MCP 2025-11-25
URL

TDQS

A3.9/5.0

Scored across 40 tools

Disambiguation4/5

Most tools target a distinct resource+action, and descriptions explicitly steer between look-alikes (e.g. ledger_metric_starters_list tells you to reach for it before ledger_metric_presets_list, and feed vs metric vs reading is clearly delineated). The main soft spot is the presets/starters pair plus the dense cluster of ledger_entries_* lifecycle tools (create, draft, update, settle, retract), which an agent could still muddle without careful reading.

Naming Consistency3/5

The dominant pattern is namespace_noun_verb (comments_create, ledger_entries_list, ledger_metrics_record_reading), but the docs family breaks rank with verb_noun and no namespace (get_doc, list_docs, search_docs). There is also a singular/plural wobble (ledger_feeds_* vs ledger_feed_sources_list) that keeps it from being fully predictable.

Tool Count3/5

40 tools is heavy by the rubric's own bar (25+), though the surface spans roughly ten genuinely distinct subsystems (entries, feeds, metrics, presets, starters, plans, comments, tags, docs, search), so each tool has a plausible place. It reads as a large but not gratuitous surface, landing it on the borderline.

Completeness5/5

Coverage is unusually thorough: full lifecycle for entries (create/draft/update/settle/retract), metrics (create/update/archive + readings), feeds (create/update/delete/run/preview/backfill), plans (add/update/delete/list), comments (create/list/resolve) and tags (browse/get/set), plus cross-cutting search. No obvious dead ends for the stated memory-and-metrics domain — docs are intentionally read-only.

Available Tools

40 tools
comments_createComment on an entityAInspect

Posts a comment on a workspace entity — a new thread, or a reply when rootId is given. Use it to leave findings where the discussion already lives (an eval result on the flow being debated, a summary on a long thread). Mention people via mentionedUserIds (from workspace member ids) to ring their notification bell; never mention someone who didn't ask to be pulled in.

ParametersJSON Schema
NameRequiredDescriptionDefault
bodyYes
rootIdNoReply into this thread; omit to start a new one.
entityIdYes
workspaceNoWorkspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys.
approvalIdNo
entityKindYesWhat the thread hangs on.
mentionedUserIdsNo

TDQS

A3.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=false, destructiveHint=false and openWorldHint=false, so the write/safety profile is covered. The description adds value beyond that by disclosing a side effect annotations cannot express: mentioning users rings their notification bell, with an accompanying social caution. It omits permission/auth requirements for posting.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the core action and its two modes, then layers usage and mention etiquette compactly. Two sentences, minimal waste, though the trailing 'never mention someone who didn't ask' is advisory padding rather than invocation-critical information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 7-parameter mutation with 43% schema coverage and no output schema, the description covers the social/mention dimension well but leaves the entity-targeting parameters and the undocumented approvalId unexplained, which an agent would need to call this reliably in all cases.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is only 43%, so the description must compensate, and it partially does: rootId (reply target), mentionedUserIds (workspace member ids, notification behavior), and body are implied. However entityKind's 17-value enum, entityId, and especially approvalId are unexplained in both schema and description, leaving real gaps.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a precise verb and resource ('Posts a comment on a workspace entity') and distinguishes the two modes of operation: new thread vs. reply when rootId is given. It never names its closest siblings (comments_list, comments_resolve), so it stops just short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives concrete usage context ('leave findings where the discussion already lives') with two illustrative scenarios (eval result on a debated flow, summary on a long thread). It does not state when not to use it or point to a sibling alternative, so it lacks the explicit routing of a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

comments_listRead an entity's comment threadsA
Read-only
Inspect

Lists the comment threads on one workspace entity (open first, then resolved) with authors and timestamps. Read this before weighing in on contested work — the threads are where disagreement lives before it becomes a decision.

ParametersJSON Schema
NameRequiredDescriptionDefault
entityIdYes
workspaceNoWorkspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys.
entityKindYesWhat the thread hangs on.

TDQS

A3.9/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, destructiveHint=false, so the safety profile is covered. The description adds genuinely new behavior: threads are returned open-before-resolved and include authors and timestamps, which shapes how an agent interprets output. It stops short of pagination or volume limits.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences with the key facts front-loaded and no redundancy. The second sentence is motivational framing that carries mild value but is slightly softer than a hard routing rule.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a read-only list tool with no output schema and a mostly documented schema, the description covers ordering, content, and the entity scoping needed to call it. Pagination/result-size behavior is the only notable gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 67%, and the schema itself handles the workspace slug nuances and the enum list. The description only says 'one workspace entity', adding little beyond the structured fields, so the baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (lists) and resource (comment threads on one workspace entity), plus the ordering rule (open first, then resolved) and payload (authors, timestamps). This clearly separates it from comments_create and comments_resolve, which mutate rather than read.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The second sentence gives a soft situational cue ('read this before weighing in on contested work'), which implies when the tool is useful. However it names no alternatives and gives no explicit when-not or prerequisite guidance, so usage remains inferred rather than stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

comments_resolveResolve or reopen a threadA
Destructive
Inspect

Sets a comment thread's resolved state (rootId = the thread's root comment id). Resolve ONLY when the human asked or the thread's question is demonstrably settled — and say what settled it in a reply first. Reopening is for new evidence.

ParametersJSON Schema
NameRequiredDescriptionDefault
rootIdYes
resolvedYes
workspaceNoWorkspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys.
approvalIdNo

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true, readOnlyHint=false and openWorldHint=false, so the mutation/safety profile is covered. The description adds real behavioral context beyond that: the precondition (a reply explaining what settled the thread) and the reopen semantics. It doesn't clarify reversibility or how the state change affects existing replies, keeping it below 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tight sentences with no filler; the action and the primary precondition are front-loaded before the reopening clause. Every clause carries decision-relevant information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a small toggle tool with no output schema, the description covers the action, the target, and the conditions well. It leaves approvalId and workspace behavior unexplained, which matters for a destructive write, but overall an agent has enough to act correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is only 25%, and the description compensates for rootId only (root comment id). 'resolved' is implied by resolve/reopen framing, but 'workspace' is documented only in the schema and 'approvalId' is explained nowhere — a notable gap for a mutation tool with an approval parameter.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (sets resolved state) plus resource (comment thread), and covers both directions — resolve and reopen. The parenthetical 'rootId = the thread's root comment id' disambiguates the target, making it clearly distinct from comments_create/comments_list.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly gates the action: resolve ONLY when the human asked or the question is demonstrably settled, and reply first explaining what settled it; reopen is for new evidence. This is genuine when/when-not guidance with a required prior step.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

entity_tags_browseBrowse the workspace's tagsA
Read-only
Inspect

Without a tag: every tag in use across the workspace with how many entities carry it, most-used first — the vocabulary the team already organizes by. With a tag: everything filed under it across every tool, each with its kind, id, title, and path. Use it to reuse existing labels instead of inventing near-duplicates, and to answer 'show me everything about X' when X is a label.

ParametersJSON Schema
NameRequiredDescriptionDefault
tagNoA tag to expand into its items. Omit to list tags.
workspaceNoWorkspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys.

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnly/destructive/openWorld, so safety is covered, and the description adds real behavioral detail: result ordering (most-used first), entity counts per tag, and the per-item fields returned in tag mode (kind, id, title, path). No pagination or size-limit behavior is mentioned, keeping this from a 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded and cleanly parallel: 'Without a tag:' then 'With a tag:' then the usage clause. Every sentence carries distinct information with no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description correctly carries the return-value burden for both modes, and annotations cover the safety profile while the schema covers both parameters. Nothing an agent needs to select or call this tool is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3, but the description adds meaning beyond the schema by characterizing what each mode of the tag parameter actually returns, turning a bare 'omit to list tags' into the two distinct result shapes. The workspace parameter semantics remain entirely schema-borne.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific dual operation (list every tag in the workspace, or expand one tag into all entities filed under it) with the exact resource and scope. It is immediately distinguishable from the write-oriented siblings entity_tags_get and entity_tags_set because the browse mode and cross-tool aggregation are spelled out.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives concrete when-to-use guidance: reuse existing labels rather than inventing near-duplicates, and answer 'show me everything about X' when X is a label. It stops short of naming sibling alternatives or stating when-not-to-use (e.g. versus search_workspace), so it is clear context without explicit exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

entity_tags_getRead the tags on entitiesA
Read-only
Inspect

Returns the tags on a batch of entities of one kind — the labels galleries organize by. Ids come from the kind's list/get tool or from search_workspace. Use it before entity_tags_set so you replace the full set knowingly, and to answer 'what is this filed under'. Entities the user can't see are omitted.

ParametersJSON Schema
NameRequiredDescriptionDefault
idsYes
workspaceNoWorkspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys.
entityKindYesWhich kind the ids belong to.

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and destructiveHint=false, so safety is covered; the description adds the non-obvious behavioral fact that 'Entities the user can't see are omitted,' i.e. results are permission-filtered rather than erroring. It does not mention limits (e.g. the 100-id cap or ordering), but the value-add beyond annotations is genuine.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences, all front-loaded: what it returns first, then usage, then the permission caveat. No filler and each sentence carries distinct information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description carries the return-value burden and does state what comes back (tags) and the omission behavior. It leaves minor gaps — tag value format, ordering, and whether unknown ids are silently dropped or error — but nothing that blocks correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 67%, so the schema documents most params, but the description adds real provenance for the hardest parameter: ids 'come from the kind's list/get tool or from search_workspace.' That tells an agent where to obtain valid ids, which the bare array schema does not.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource with scope: 'Returns the tags on a batch of entities of one kind.' The parenthetical 'the labels galleries organize by' distinguishes tags from other entity metadata, and the named siblings (entity_tags_set, search_workspace) make the boundary clear.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives explicit usage contexts: 'Use it before entity_tags_set so you replace the full set knowingly, and to answer what is this filed under.' That is a real when-to-use plus a stated alternative. It does not mention entity_tags_browse, the other obvious read-side sibling, so it falls short of fully disambiguating the read alternatives.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

entity_tags_setSet an entity's tagsA
Destructive
Inspect

Replaces the FULL tag set on one entity (an empty list clears it). Read the current tags with entity_tags_get first and pass the merged list — this is not additive. Tags are lowercase letters, numbers, spaces, and hyphens; prefer labels already in use (entity_tags_browse) so the workspace's vocabulary stays small. The id comes from the kind's list/get tool or search_workspace.

ParametersJSON Schema
NameRequiredDescriptionDefault
tagsYes
entityIdYes
workspaceNoWorkspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys.
approvalIdNoApproval id from a prior needs_confirmation response. Omit on the first call.
entityKindYes

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true and readOnlyHint=false, and the description reinforces this with the replace-not-merge semantics and the empty-list-clears behavior. It does not mention the needs_confirmation/approvalId retry flow that the schema implies, which is the one notable behavioral gap.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three dense sentences, all front-loaded with the destructive replace semantics first, followed by workflow and vocabulary guidance. Every sentence earns its place; the parentheticals make it slightly busy but not wasteful.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema exists, and the description covers destructive semantics, prerequisites, id sourcing, and tag format well enough to invoke the tool correctly. The confirmation/approval retry path is left to the schema field description rather than the tool description.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is only 40%, and the description compensates by documenting the tag character set (lowercase letters, numbers, spaces, hyphens), the merge requirement for the tags array, and the origin of entityId. The workspace and approvalId parameters are only explained in the schema, so coverage is good but not complete.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Replaces the FULL tag set on one entity') and immediately clarifies the destructive scope with '(an empty list clears it)'. This cleanly separates it from the sibling readers entity_tags_get and entity_tags_browse.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives explicit workflow: read current tags with entity_tags_get first, pass the merged list, and warns 'this is not additive'. Also routes to entity_tags_browse for vocabulary reuse and names where the entityId comes from (kind's list/get tool or search_workspace), so the agent knows both prerequisites and alternatives.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_docFetch ZeroWidth doc by slugA
Read-only
Inspect

Fetch the full Markdown body of a specific docs page by its slug. Use this after search_docs when the user needs the complete content of a page. No authentication required.

ParametersJSON Schema
NameRequiredDescriptionDefault
slugYesPage slug. Accepts 'compass/api', '/compass/api', or 'docs/compass/api'.

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, openWorldHint=false, and destructiveHint=false, covering the safety profile. The description adds valuable context by stating 'No authentication required' and specifying the return body as full Markdown, which agents need to know beyond the structured hints.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short sentences, front-loaded with the core purpose ('Fetch the full Markdown body...'), followed by usage routing and a salient behavioral note. Every sentence earns its place with no redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple one-parameter read tool with full schema coverage and annotations covering safety, the description is complete: it states the return format, the required input type, usage context, and authentication status. Nothing an agent needs to invoke it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, and the single `slug` parameter is fully documented in the schema with accepted formats. The description adds no syntax or format details beyond what the schema provides, so it meets the baseline rather than exceeding it.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('Fetch'), resource ('full Markdown body of a specific docs page'), and retrieval key ('by its slug'). It distinguishes itself from the search-oriented sibling by positioning as the step after `search_docs` for complete page content.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly says to use this after `search_docs` when the user needs the complete content of a page, which names the alternative and the condition that selects it. It does not state when not to use it (e.g., for listing or metadata), but the context is clear enough for a simple read tool.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

ledger_entries_add_evidenceAttach evidence to a Ledger entryAInspect

Records what happened against an open decision — a manual observation the user reports ('the pilot team says triage feels faster'), an implementation note, or a reference to a Caliper run. Evidence is what settlement later reads, so attach it as it arrives and cite it in the lesson. Metric readings attach themselves via ledger_metrics_record_reading — don't duplicate them here. Fails with conflict on superseded entries. Ids come from ledger_entries_list. May return needs_confirmation.

ParametersJSON Schema
NameRequiredDescriptionDefault
dataNoStructured detail worth keeping with the summary.
kindNoUsually `manual`; `caliper_eval_run` with refId for a run; `implementation` when the change shipped.manual
refIdNoSource record id (a Caliper run id, …) when there is one.
entryIdYesThe entry the evidence is about.
summaryYesWhat happened, in one or two sentences.
workspaceNoWorkspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys.
approvalIdNoApproval id from a prior needs_confirmation envelope.

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare a non-destructive, non-open-world write, and the description adds real behavioral context beyond them: failure with conflict on superseded entries, a possible `needs_confirmation` envelope, and the overall importance of evidence for later settlement. It doesn't describe the success return shape, but with an envelope caveat and failure mode disclosed this is stronger than average.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core purpose and follows with routing, failure, and id-sourcing facts. Some narrative (the quoted user anecdote, 'cite it in the lesson') is looser than needed, but every sentence carries usable signal.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a mutation tool with no output schema, it covers routing, prerequisites, duplicate-avoidance, failure modes, and the confirmation envelope, which are the pieces an agent needs. It omits nothing critical, though it could state whether the appended entry becomes immutable.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents all seven parameters including the `kind` enum and `workspace` token rules. The description adds usage flavor (evidence kinds, refId for a Caliper run) but no syntax the schema lacks, so baseline 3 is correct.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource (records/attaches evidence against an open decision) and enumerates concrete kinds of evidence. It explicitly distinguishes itself from the sibling ledger_metrics_record_reading, so an agent can route without opening either schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives clear when-to-use ('attach it as it arrives and cite it in the lesson'), a when-not ('Metric readings attach themselves via ledger_metrics_record_reading — don't duplicate them here'), and prerequisite sourcing ('Ids come from ledger_entries_list'). Alternatives and conditions are named explicitly.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

ledger_entries_createRecord a decision, lesson, or observation in the LedgerAInspect

Records a memory entry. Every entry is a TITLE (summary: one short plain sentence) over a BODY (rationale: the detail, markdown welcome) — never put the detail in the title. Three kinds: decision — a change being made now; PRE-REGISTRATION IS THE POINT, so prediction (what we expect) must be written NOW, before any evidence exists, and never backfilled to match an outcome. lesson — a distilled belief the team already holds (lesson text required, no prediction). observation — a durable fact worth remembering (no prediction): something already true, never something planned. Ideas, pitches, backlog items, and upcoming work are NOT entries — a dated piece of work that carries out a decision is a plan item (ledger_plan_add on that decision), and a running list you keep across runs belongs in a Napkin doc or sheet. When a user states something durable about their business in conversation, offer to capture it as a lesson or observation. When a user shares MEETING NOTES, propose the decisions you find with ledger_entries_draft so they keep or drop each one in Ledger; use this tool for an entry they've asked you to record. Link the Compass pages the entry touches so future consult-before-acting finds it. May return needs_confirmation — summarize the entry and wait for approval.

ParametersJSON Schema
NameRequiredDescriptionDefault
kindNodecision = a change with a pre-registered expectation; lesson = a belief arriving already settled; observation = a fact that is already true (not a plan, idea, or pitch).decision
lessonNoWhat we learned — required for kind=lesson, optional for kind=observation, forbidden on decisions (their lesson is written at settlement).
summaryYesThe entry's TITLE: one short plain sentence naming what we did (decision) or the fact itself (lesson, observation). Plain text, no markdown, at most 280 characters — e.g. "Newsletter moves to a biweekly cadence". Everything longer goes in rationale.
rationaleNoThe entry's BODY: the detail under the title — why, context, lists, steps. Markdown is fine here and renders as formatted text.
workspaceNoWorkspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys.
approvalIdNoApproval id from a prior needs_confirmation envelope.
predictionNoDecisions ONLY. What we expect: a list of `claims` plus one shared `deadline`, or `freeform` for room decisions nothing can settle. A claim either REACHES a value (comparator + target, e.g. CTA clicks >= 400) or HOLDS one (`hold`, e.g. newsletter reads no more than 5% below the 28 days before this). Name every number the change is expected to move AND every number it shouldn't cost — a decision that claims only what it hopes will rise gets to pick its own evidence. Set `watch: true` on a number worth following that the decision isn't committing to. When the change aims at part of an event metric ("new users in Germany"), narrow the claim with `slice` ({ country: ["DE"] }) using label values from ledger_metrics_get — a claim on an event metric settles on the total counted since landing (reach) or events per day (hold).
compassPageIdsNoWhere — Compass page ids this decision touches.

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare readOnlyHint=false, destructiveHint=false, openWorldHint=false, so the write/safety profile is largely covered. The description adds genuinely non-annotation behavior: it may return `needs_confirmation`, in which case the agent should summarize and wait for approval, plus the pre-registration invariant (prediction written before evidence, never backfilled) that constrains correct use.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the title/body model and the three kinds, and each subsequent sentence carries a rule or a routing decision rather than filler. It is dense and uses heavy caps for emphasis, which is slightly noisy, but nothing is padding.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For an 8-parameter tool with nested prediction objects and no output schema, the description covers kinds, the pre-registration constraint, the approval flow, and the Compass linking purpose. Return-shape detail is the only meaningful omission, and the needs_confirmation envelope is at least flagged.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% so the baseline is 3, but the description adds framing the schema does not: the title/body split (summary as TITLE, rationale as BODY, never detail in the title) and the rule that a decision's lesson is written at settlement. The prediction block is well covered by its own schema description, so the added value is modest rather than transformative.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource (records a memory entry in the Ledger) and immediately breaks the resource into its three kinds (decision/lesson/observation) with a one-line definition for each. It further distinguishes itself from siblings by name (ledger_plan_add, ledger_entries_draft), so an agent can route without opening another schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives explicit when-to-use routing: use ledger_entries_draft for meeting notes so the user keeps/drops each proposal, use ledger_plan_add for dated work that carries out a decision, use this tool for an entry the user asked to record, and put running lists in a Napkin doc. It also names what is NOT an entry (ideas, pitches, backlog items). This is about as complete as usage guidance gets.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

ledger_entries_draftPropose Ledger entries for reviewAInspect

Writes up to 50 entries as DRAFTS: proposals that stay out of the record until a person confirms each one in Ledger's Drafts view. Use it for entries you found rather than were told — decisions in meeting notes, or a team's past changes read from its tracker, pull requests or launch posts (set fromHistory: true). For history: take each claim from what the source said AT THE TIME, set landedAt to when it shipped, attach the source URL, and leave rollout as full unless the source says otherwise. These read as low confidence because they were written down after the fact; say so plainly rather than overstating them. Tell the person how many drafts are waiting and that they review them under Decisions → Drafts. May return needs_confirmation.

ParametersJSON Schema
NameRequiredDescriptionDefault
draftsYes
workspaceNoWorkspace slug. Ignored for workspace API keys.
approvalIdNo
fromHistoryNoTrue when the drafts come from a team's past work, not the conversation.

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations state readOnlyHint=false, destructiveHint=false, openWorldHint=false, but the description adds the crucial behavioral fact that these are non-committing proposals reviewed in Decisions → Drafts, that confidence is low for historical claims, and that the call may return `needs_confirmation`. Those are traits the annotations cannot convey.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Purpose and safety are front-loaded in the first sentence, and the remaining sentences are operational instructions rather than filler. It is dense and fairly long, with some stylistic guidance ('say so plainly rather than overstating them') that is useful but could be tightened.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a write-of-drafts tool with no output schema, no nested-object flag, and 50% schema coverage, the description supplies the workflow, the history-specific field rules, the confidence caveat, the user-facing next step, and the possible `needs_confirmation` return. Nothing essential for correct invocation is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With 50% schema coverage, the description compensates well: it explains fromHistory semantics, prescribes landedAt ('when it shipped'), rollout defaults ('full unless the source says otherwise'), source URL attachment, and that prediction is decisions-only. It doesn't touch workspace or approvalId, but the high-value parameters are covered.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb, resource, and scope: 'Writes up to 50 entries as DRAFTS: proposals that stay out of the record until a person confirms each one in Ledger's Drafts view.' The 'draft' framing distinguishes it cleanly from ledger_entries_create, which commits records, so an agent can route without opening a schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It explicitly states when to use it ('entries you found rather than were told') and gives concrete trigger examples — decisions in meeting notes, a team's past changes from its tracker, PRs, or launch posts — plus the fromHistory condition. This is explicit when-to-use guidance with no inference required.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

ledger_entries_getRead one Ledger entry in fullA
Read-only
Inspect

One entry's full anatomy — the six fields, attached evidence (Caliper runs, measurements, observations), and the supersede chain (what replaced it, or what it replaced). Cite entry ids when telling the user about prior related decisions.

ParametersJSON Schema
NameRequiredDescriptionDefault
entryIdYesEntry id (from ledger_entries_list).
workspaceNoWorkspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys.

TDQS

A3.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, destructiveHint=false, and openWorldHint=false, so the safety profile is covered. The description adds real value beyond that by enumerating the returned content (six fields, Caliper runs, measurements, observations, supersede chain), which matters since there is no output schema.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two compact sentences with no filler. The most useful content (what the entry contains) is front-loaded, and the second sentence is a targeted usage tip rather than padding.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description responsibly summarizes the return payload, and annotations cover the safety profile while the schema fully covers parameters. Only the relationship to sibling tools (list, update, retract) is left implicit.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so both entryId and workspace are fully documented in the schema, including the personal-token workspace requirement. The description adds no parameter syntax or meaning beyond that, so the baseline of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

Names a specific verb and resource ('Read one Ledger entry in full') and then details exactly what the entry contains: the six fields, attached evidence, and the supersede chain. An agent can tell it apart from ledger_entries_list, though the description never explicitly contrasts with that sibling.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is only implied. The closing sentence ('Cite entry ids when telling the user about prior related decisions') advises how to use the output, not when to select this tool over ledger_entries_list or ledger_entries_update. No exclusions or prerequisites are stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

ledger_entries_listList Ledger decision entriesA
Read-only
Inspect

The workspace's memory: decisions (changes with pre-registered expectations), lessons (distilled beliefs), and observations (captured facts). CONSULT BEFORE ACTING — before proposing a flow change, prompt edit, or process decision, filter by the Compass page it touches (compassPageId) and check whether prior attempts exist and how they settled. Filter status=open for unsettled expectations awaiting evidence (each carries a derived lapsed flag — true when its deadline has passed; surface lapsed ones when the user asks what needs attention); kind=lesson for what the team already believes.

ParametersJSON Schema
NameRequiredDescriptionDefault
qNoFree-text search across an entry's summary, rationale, and lesson (case-insensitive). Use it to find the entry about a topic before acting.
kindNoFilter by species: decision (pre-registered expectations), lesson (distilled beliefs), observation (captured facts).
cursorNoPagination cursor from a prior page.
originNoFilter by who wrote it (manual, workbench, …).
statusNoFilter by lifecycle state (open = awaiting evidence).
workspaceNoWorkspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys.
compassPageIdNoCompass page id — every decision touching that workflow / system / person.

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare a safe read (readOnlyHint=true, destructiveHint=false, openWorldHint=false), so the bar is lower. The description adds genuine behavioral context beyond them: that status=open entries carry a derived `lapsed` flag and that lapsed entries should be surfaced when the user asks what needs attention — meaningful with no output schema to document it.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the ledger's purpose before the usage directives, and each clause maps to an actionable filter. The semicolon-heavy second half packs three filter recipes into one dense sentence, which slightly reduces scannability but doesn't waste words.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 7-param, zero-required list tool with annotations covering safety and no output schema, the description supplies the missing behavioral detail (the derived lapsed flag) and the consult-before-acting workflow. Return shape and pagination semantics remain undocumented, which keeps it short of a 5.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With schema coverage at 100%, the baseline is 3, but the description goes further by explaining how filters compose (use compassPageId to find prior attempts and how they settled; status=open means awaiting evidence) rather than restating field descriptions. It clarifies the intent behind combining parameters, which the schema alone does not.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description identifies the resource (the ledger: decisions, lessons, observations) and the operation by naming it a filtered list, but the listing verb is more implied than stated. It is distinguishable from siblings like ledger_entries_get or ledger_entries_settle through the filtering framing, though it never explicitly contrasts them.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives explicit when-to-use guidance ('CONSULT BEFORE ACTING — before proposing a flow change, prompt edit, or process decision') and concrete selection recipes for parameters (compassPageId to check prior attempts, status=open for unsettled expectations, kind=lesson for team beliefs). This is a strong routing instruction rather than a vague hint.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

ledger_entries_retractRetract a Ledger entryA
Destructive
Inspect

Takes an entry out of the curated ledger — THE remedy when you recorded something wrong (a decision that wasn't made, a duplicate, a fact the user corrects). Reversible from Ledger and fully audited, so it is safe to offer as soon as the user says 'that's not right'. Not for overturning a settled claim the team once believed — a human supersedes that. Fails with conflict if already retracted. Ids come from ledger_entries_list. May return needs_confirmation.

ParametersJSON Schema
NameRequiredDescriptionDefault
entryIdYesThe entry to retract.
workspaceNoWorkspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys.
approvalIdNoApproval id from a prior needs_confirmation envelope.

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructive=true/readOnly=false, but the description adds substantial context beyond them: reversibility from Ledger, full audit trail, the conflict failure mode when already retracted, and a possible `needs_confirmation` envelope. Reversibility does not contradict destructiveHint — it refines it.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded and largely earn-its-place, with the core remedy statement first and exclusions after. A couple of clauses (the quoted 'that's not right' reassurance, the 'THE remedy' emphasis) are slightly redundant with earlier sentences, adding density without new information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a mutation tool with no output schema, the description covers what an agent needs: the confirmation envelope it may receive, the conflict behavior, and the safety/reversibility profile. Nothing essential for correct invocation is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3, but the description adds real value by stating where entryId comes from ('Ids come from ledger_entries_list') and by tying `needs_confirmation` to the approvalId flow. It stops short of describing the workspace parameter's semantics, which the schema already covers.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Takes an entry out of the curated ledger') and immediately frames the purpose in operationally meaningful terms (correcting a mistaken record). It distinguishes itself from siblings like ledger_entries_settle and ledger_entries_update by scoping to retroactive removal of a wrongly recorded entry.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicit when-to-use ('THE remedy when you recorded something wrong — a decision that wasn't made, a duplicate, a fact the user corrects') with concrete triggers, plus an explicit when-NOT-to-use ('Not for overturning a settled claim the team once believed — a human supersedes that'). The alternative and the condition selecting it are named.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

ledger_entries_settleSettle an open Ledger decisionA
Destructive
Inspect

The ritual moment: evidence has landed, the expectation closes, the lesson is written. Only propose settlement when attached evidence actually answers the prediction — check ledger_entries_get first and cite the evidence in the lesson. lesson (what we now believe) is required and permanent; settlement happens exactly once. May return needs_confirmation.

ParametersJSON Schema
NameRequiredDescriptionDefault
lessonYesWhat we now believe — the distilled, citable lesson.
entryIdYesThe open entry to settle.
verdictNoDid the pre-registered expectation hold? `confirmed` every claim came out as hoped, `missed` none did, `mixed` some did and some didn't. Mixed is not a softer miss — "it worked and it cost us something" is usually the most informative result a change can produce, so use it rather than rounding to either side. Only when the entry carries an expectation and the evidence gives a clear answer.
workspaceNoWorkspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys.
approvalIdNo

TDQS

A3.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true and readOnlyHint=false; the description meaningfully extends this by stating the lesson is 'permanent' and that 'settlement happens exactly once', which makes the irreversibility concrete. It also discloses a `needs_confirmation` response path despite there being no output schema — useful granularity beyond the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

It is short and front-loaded, but the opening clause ('The ritual moment: evidence has landed, the expectation closes, the lesson is written') is atmospheric framing that consumes prime space without adding operational content. Trimming it would leave a purely actionable description.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 5-parameter mutation tool with no output schema, the description covers the important ground: prerequisites, irreversibility, one-shot settlement, and the confirmation path. It omits guidance on workspace/approvalId handling, which the schema partially covers, but nothing critical to calling it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 80%, so the schema already documents lesson, entryId, verdict and workspace. The description adds only that `lesson` is permanent and should cite evidence; verdict, workspace and approvalId semantics are left entirely to the schema. Baseline 3 is appropriate when the schema does most of the work.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description conveys the operation — closing an entry once evidence lands and writing a lesson — and distinguishes it from siblings like ledger_entries_add_evidence by requiring that evidence 'actually answers the prediction'. The opening metaphor ('the ritual moment') is evocative rather than literal, so the exact verb+resource must be inferred, but it is recoverable.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives an explicit precondition for use ('Only propose settlement when attached evidence actually answers the prediction') and routes the agent to `ledger_entries_get` to verify first. It stops short of naming the correct alternative when the precondition fails (e.g., add_evidence), so the exclusion is implied rather than fully specified.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

ledger_entries_updateEdit a Ledger entryA
Destructive
Inspect

Corrects an entry in place — its title (summary), body (rationale), the pre-registered prediction of an OPEN decision (null clears it), or its facet (filing category; editable even after settlement). Works on open decisions, and on lessons, observations, and beliefs (they carry no expectation, so fixing their wording is fine any time). Use to fix a typo, swapped fields, or a misrecorded detail. Never edits kind or lesson. Fails with conflict on a settled decision — its claim is corrected by a human superseding it in Ledger — and on superseded or retracted entries. Entry ids come from ledger_entries_list. May return needs_confirmation.

ParametersJSON Schema
NameRequiredDescriptionDefault
facetNoFiling facet; empty string = unassigned.
entryIdYesThe entry to edit: an open decision, or a lesson / observation / belief.
summaryNoNew title: one short plain sentence, no markdown, at most 280 characters.
rationaleNoNew body: the detail under the title; markdown is fine.
workspaceNoWorkspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys.
approvalIdNoApproval id from a prior needs_confirmation envelope.
predictionNoReplacement expectation (claims + deadline, or freeform), or null to clear it. Correcting a misrecorded expectation is fine; rewriting one to match what happened is not.

TDQS

A4.9/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only declare destructive/not-read-only; the description adds substantial context beyond them: the conflict failure mode on settled decisions, the null-clears-prediction semantics, that facet stays editable post-settlement, that kind/lesson are immutable, and that a needs_confirmation envelope may be returned. This is the behavioral detail an agent needs for a destructive in-place mutation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads what is corrected, then constraints, then failure modes, then the id source — a logical order of decreasing immediate relevance. Every clause carries a distinct rule (editable set, invariants, failure conditions, confirmation flow); no filler sentences.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 7-param destructive mutation with a nested prediction object and no output schema, the description covers the editable surface, immutability rules, failure/conflict behavior, id provenance, and the confirmation envelope. An agent has everything needed to invoke it correctly and to anticipate refusals.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3 and the schema already documents each field. The description still adds meaning beyond it by mapping summary→title, rationale→body, explaining null clears the prediction, and noting facet is editable even after settlement. Modest added value over an already-complete schema, hence 4 rather than 5.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('Corrects') plus resource ('an entry') and enumerates exactly which fields are editable (summary, rationale, prediction, facet). It also draws hard boundaries against siblings by declaring what it never edits (kind, lesson) and what it will not touch (settled, superseded, retracted entries). An agent can distinguish it from ledger_entries_settle/retract/create without opening a schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives explicit when-to-use ('fix a typo, swapped fields, or a misrecorded detail') and explicit when-not (settled decisions, superseded/retracted entries), with the reason for the refusal. It also routes the agent to ledger_entries_list for ids, removing the guesswork about where entryId comes from.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

ledger_feeds_backfillFill in a metric's history from its feedAInspect

Run the same frozen instruction over a past range, so a metric has history instead of starting the day it was set up. Offer this whenever you create a feed — a chart with a year behind it is worth far more than one that begins today, and an expectation can be judged against what normal looked like. Works when the source returns a series; a source that only ever reports 'right now' will write one point. Safe to repeat: overlapping ranges dedupe. Keep ranges to a few hundred points; a run that fills up says so. May return needs_confirmation.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYesFeed id, from ledger_feeds_list.
toNoEnd of the range. Defaults to now.
fromYesStart of the range, ISO date or instant.
workspaceNoWorkspace slug. Omit to use the pinned workspace.
approvalIdNo

TDQS

A3.9/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only declare mutation-safety (readOnly=false, destructive=false, openWorld=true). The description adds genuinely useful behavioral detail beyond them: idempotency via dedupe on overlapping ranges, size-limit feedback ('a run that fills up says so'), the single-point outcome for non-series sources, and a `needs_confirmation` return that ties to the approvalId param. This is well above what the annotations convey.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the purpose, then layers usage, constraints, and re-run safety. Mostly efficient, though the 'worth far more than one that begins today' pitch is mildly promotional filler that could be trimmed.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema exists, and the description compensates by disclosing the re-run/dedupe behavior, the range-size limitation, and the `needs_confirmation` outcome. Combined with the annotations, an agent has enough to call and interpret it, though nothing confirms what a successful backfill returns.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 80%, so the schema already documents id/from/to/workspace. The description implies the from-range concept and hints that approval/confirmation is involved (linked to approvalId), but adds no format or semantics beyond the schema. Baseline 3 is appropriate at this coverage level.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific action: run the frozen feed instruction over a past range to populate history. The verb+resource ('run' + 'frozen instruction'/'range') is concrete. It does not name the closely related sibling ledger_feeds_run, so it stops short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives a clear condition for invoking: 'Offer this whenever you create a feed,' with a rationale (a chart with history is more valuable, expectations can be judged). It does not explicitly distinguish when to use this versus ledger_feeds_run or ledger_feeds_preview, so no exclusions are covered.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

ledger_feeds_createSet up how a metric's numbers arriveAInspect

Freeze a tool call as a metric's standing source: it runs on the cadence you give and records what it finds, with no model involved. metric takes an id or a snake_case slug; an unknown slug starts tracking that metric. Confirm the mapping with ledger_feeds_preview first. Tell the user they can backfill history afterwards with ledger_feeds_backfill. May return needs_confirmation.

ParametersJSON Schema
NameRequiredDescriptionDefault
labelNoShort name for the feed, e.g. "Weekly signups".
metricYesMetric id, or a snake_case slug to start tracking.
mappingYes
scheduleYes
toolArgsNo
toolNameYesThe tool's own name on that server.
workspaceNoWorkspace slug. Omit to use the pinned workspace.
approvalIdNo
windowHoursNoHow far back each run looks. Sensible per-cadence default.
integrationIdYesFrom ledger_feed_sources_list.

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare this is a non-read-only, non-destructive, closed-world mutation. The description adds real behavioral context beyond that: the feed runs autonomously with 'no model involved', an unknown slug has the side effect of starting tracking for a new metric, and the call may return `needs_confirmation`. It does not, however, explain scheduling failure modes or what happens on a later conflicting feed.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the tool's purpose, then three compact functional sentences covering slug semantics, the preview prerequisite, and the backfill follow-up. Every sentence carries information, though the 'Tell the user they can backfill' clause is more agent-behavior guidance than tool description and slightly dilutes the focus.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 10-parameter mutation with nested mapping/schedule objects and no output schema, the description covers the important operational surface: cadence, the mapping-confirmation prerequisite, the slug side effect, and the `needs_confirmation` return signal. It leaves the nested mapping and schedule fields entirely to the schema, which is a modest gap given the nesting.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 60%, so the schema carries most of the burden. The description's only parameter statement ('metric takes an id or a snake_case slug; an unknown slug starts tracking that metric') largely restates the schema's own description for `metric`, adding the side-effect note but nothing about mapping, schedule, toolArgs, or windowHours. Baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Freeze a tool call as a metric's standing source') and immediately characterizes the mechanism (runs on cadence, records findings, no model involved). It is immediately distinguishable from siblings like ledger_feeds_preview and ledger_feeds_run, which are named in the same breath.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives explicit sequencing instructions: confirm the mapping with ledger_feeds_preview first, then tell the user about ledger_feeds_backfill afterwards. The prerequisite and follow-up alternatives are named with the condition that selects each, leaving nothing to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

ledger_feeds_deleteStop a metric feedA
Destructive
Inspect

Remove a feed. The readings it already wrote stay — the series is the record and outlives the instruction. Prefer pausing with ledger_feeds_update when the user might want it back. May return needs_confirmation.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYesFeed id, from ledger_feeds_list.
workspaceNoWorkspace slug. Omit to use the pinned workspace.
approvalIdNo

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true and readOnlyHint=false, so the deletion semantics are covered structurally. The description adds real context beyond them: readings already written are retained ('the series is the record and outlives the instruction') and that a needs_confirmation response is possible. It stops short of describing the confirmation/approval flow or permission requirements, which is the remaining gap.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three tight sentences, front-loaded with the action, then consequence, then the preferred alternative. Every sentence carries distinct information with no padding.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a destructive no-output-schema tool, it covers the critical agent-facing questions: irreversibility of the instruction vs. retention of data, the safer alternative, and a possible confirmation response. It could still say what confirmation requires or what happens to dependent feed sources, but nothing essential to correct invocation is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 67%; the schema already explains 'id' (from ledger_feeds_list) and 'workspace'. The description adds no parameter-level meaning and leaves 'approvalId' undocumented in both places, though the needs_confirmation mention hints at that flow without explaining it. Baseline 3 is appropriate given the schema does most of the work.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Remove a feed') and immediately distinguishes itself from the sibling update/pause path. An agent can separate it from ledger_feeds_update, ledger_metrics_archive, and ledger_plan_delete without opening any schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly names the alternative ('Prefer pausing with ledger_feeds_update') and the condition that selects it ('when the user might want it back'). This is a genuine when/when-not routing rule, not implied usage.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

ledger_feeds_listList metric feedsA
Read-only
Inspect

The standing instructions for how metrics' numbers arrive — which tool each one calls, how often, and how the last run went. Check before proposing a feed so you don't duplicate one, and consult when a user asks why a metric is stale: a feed with a failing last run is usually the answer.

ParametersJSON Schema
NameRequiredDescriptionDefault
metricIdNoNarrow to one metric (id or slug).
workspaceNoWorkspace slug. Omit to use the pinned workspace.

TDQS

A3.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, destructiveHint=false, openWorldHint=false, so the safety profile is fully covered. The description adds the useful fact that each feed carries a last-run status, which explains its diagnostic value, but says nothing about ordering, pagination, or result size. With annotations carrying the safety burden, this is adequate but not rich.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, no filler, with the resource definition front-loaded ahead of the usage guidance. The em-dash clause packing three attributes is dense but informative. Slightly indirect opening (a noun phrase rather than a verb phrase) costs it the top score.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema exists, so the description must convey what a feed record contains, and it does name the key fields (target tool, cadence, last-run outcome). With only two optional, fully documented parameters and complete safety annotations, an agent has what it needs to call this correctly; ordering and pagination behavior are the only real omissions.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so both parameters (metricId as id-or-slug narrowing, workspace with pinned-workspace default) are already fully documented in the schema. The description adds no parameter detail at all, which is the expected baseline when the schema does the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description defines the resource concretely and specifically: 'the standing instructions for how metrics' numbers arrive — which tool each one calls, how often, and how the last run went.' That is enough for an agent to distinguish metric feeds from sibling families like ledger_entries_*, ledger_feeds_create/run, and ledger_metrics_list. It never states the verb 'list' explicitly, relying on the title for that, which keeps it just short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives two concrete triggers: check before proposing a feed to avoid duplication, and consult when a user asks why a metric is stale because a failing last run is usually the cause. That is unusually actionable guidance for a list tool. It does not name competing tools (e.g. ledger_feed_sources_list or ledger_feeds_preview) as alternatives, so no explicit when-not routing.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

ledger_feed_sources_listList servers a metric feed could pull fromA
Read-only
Inspect

The connected servers this workspace exposes to you, with the id a feed needs. Their tools appear to you namespaced as ext__<name>__<tool> — call one directly to see what it returns before proposing a feed. A server the workspace has switched off for you is not listed and cannot be fed from here.

ParametersJSON Schema
NameRequiredDescriptionDefault
workspaceNoWorkspace slug. Omit to use the pinned workspace.

TDQS

A3.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With readOnlyHint/openWorldHint/destructiveHint already declaring a safe, closed-world read, the description still adds real behavioral context: the `ext__<name>__<tool>` namespacing, that server tools can be called directly to inspect returns, and that disabled servers are excluded and cannot be fed from. That goes beyond the annotations, stopping short of describing the return shape or ordering.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, front-loaded with what is returned and the key id, followed by practical usage detail. Dense but every clause carries information; minor polish could tighten the namespacing sentence.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

There is no output schema, so the description carries the return explanation and does so adequately — servers plus the feed-usable id. It is complete enough to invoke correctly, though it leaves return ordering/shape unspecified.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% and the single optional `workspace` parameter is fully documented in the schema (omit to use the pinned workspace). The description alludes to workspace scope but adds no syntax or format beyond the schema, so the baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource: it lists the connected servers exposed to the workspace, together with the id a feed needs. The feed linkage is clear from the content, though it does not explicitly name the sibling feed tools it precedes (ledger_feeds_create/propose), so sibling differentiation is implied rather than stated.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It offers contextual guidance ('call one directly to see what it returns before proposing a feed') and an exclusion (switched-off servers are not listed and cannot be fed from), which implies when this list is useful. However, it never explicitly says when to use this vs. other ledger_feeds_* tools or states prerequisites, leaving routing to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

ledger_feeds_previewTry a feed's mapping before creating itAInspect

Call a connected server's tool once and see what a mapping would pull out of the response — nothing is written and no feed is created. Send no mapping for a first look: you get a sample of the response plus the paths that hold numbers. Then send a mapping to confirm it finds the readings you expect. Always do this before ledger_feeds_create; proposing a feed whose mapping you haven't seen work is how a metric fills up with the wrong number. May return needs_confirmation.

ParametersJSON Schema
NameRequiredDescriptionDefault
mappingNoHow to read the number. Omit on the first look.
toolArgsNoArguments, exactly as the tool wants them. Use {{from}} / {{to}} (ISO instants) or {{date}} (YYYY-MM-DD) where a time range goes — those are substituted per run, and are what let one feed also backfill.
toolNameYesThe tool's own name on that server — NOT the ext__ namespaced form you call it by.
workspaceNoWorkspace slug. Omit to use the pinned workspace.
approvalIdNo
windowHoursNoHow far back the preview's window reaches. Default 24.
integrationIdYesFrom ledger_feed_sources_list.

TDQS

A4.6/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Adds real behavioral context beyond the annotations: it discloses the possible `needs_confirmation` return, the two-pass sampling behavior, and that the call is non-persistent ('nothing is written and no feed is created'). readOnlyHint=false sits in mild tension with that claim, since it invokes an external, open-world tool, but the description is explicit about what the ledger side will not do.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Four sentences, front-loaded with the operation and its non-persistence, then the workflow, then the failure mode. No sentence is filler; the warning about wrong metrics justifies its length.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 7-parameter, nested, mutating-adjacent call with no output schema, the description covers both the invocation workflow and the shape of what comes back (a response sample plus the paths holding numbers). The only thin spot is the confirmation flow's mechanics after `needs_confirmation`, which is named but not explained.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 86%, so the baseline is 3, but the description adds genuine meaning by explaining the role of `mapping` across the two passes ('Send no mapping for a first look... then send a mapping to confirm'). The toolArgs templating semantics are already in the schema, so this is reinforcement rather than new semantics.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a precise verb and resource: call a connected server's tool once to see what a mapping would extract, with the explicit scope 'nothing is written and no feed is created'. It is immediately distinguishable from ledger_feeds_create, which it names by contrast.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives an unambiguous when-to-use ('Always do this before ledger_feeds_create'), a two-phase protocol (send no mapping first, then send a mapping to confirm), and the rationale for the exclusion (an unseen mapping is how a metric fills with the wrong number).

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

ledger_feeds_runRun a metric feed nowAInspect

Run a feed immediately over its own window and record what it finds. Use it right after creating one to prove the mapping works — a run that comes back empty names the path that missed, which is what you fix with ledger_feeds_update. Does not move the feed's schedule. Re-running is safe: readings dedupe per time bucket. May return needs_confirmation.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYesFeed id, from ledger_feeds_list.
workspaceNoWorkspace slug. Omit to use the pinned workspace.
approvalIdNo

TDQS

A4.1/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With readOnlyHint=false and destructiveHint=false already declared, the description adds real behavioral context beyond annotations: it does not move the schedule, re-running is idempotent ('readings dedupe per time bucket'), and it may return `needs_confirmation`. That is precisely the mutation/idempotency/approval information an agent needs.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Four tight sentences, action front-loaded, and every clause adds distinct information (trigger, failure handling, schedule safety, idempotency, confirmation state). No filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema exists, so the description carries return-value burden; it names the two meaningful statuses (`empty`, `needs_confirmation`) but does not describe the normal success payload. Combined with the unexplained `approvalId`, one gap remains, but the core behavior is well covered.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 67%; `id` and `workspace` are documented in the schema, while `approvalId` is undocumented. The description's mention of a possible `needs_confirmation` return hints at the approval flow tied to `approvalId`, but never explains the parameter, so it only marginally compensates.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource with scope: 'Run a feed immediately over its own window and record what it finds.' The 'over its own window' framing implicitly separates it from ledger_feeds_backfill, though no sibling is named explicitly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives a clear trigger: 'Use it right after creating one to prove the mapping works.' It also routes the failure case to ledger_feeds_update. No explicit exclusions (e.g. when to prefer preview or backfill), but the context is concrete.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

ledger_feeds_updateFix or pause a metric feedA
Destructive
Inspect

Change a feed's arguments, mapping, cadence, or pause it. Reach for this when a feed's last run reports empty (the response shape moved, so the mapping needs a new path) or error. Changing the mapping changes what the number means, so say what you're changing and why. May return needs_confirmation.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYesFeed id, from ledger_feeds_list.
labelNo
enabledNofalse pauses it; the definition and history are kept.
mappingNo
scheduleNo
toolArgsNo
toolNameNo
workspaceNoWorkspace slug. Omit to use the pinned workspace.
approvalIdNo
windowHoursNo

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already flag readOnly=false, destructive=true, openWorld=false. The description adds value beyond them: it warns that changing the mapping alters what the recorded number means, and discloses a `needs_confirmation` return path. It does not detail reversibility, approval semantics, or permission requirements, so it is not complete.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Four tight sentences, front-loaded with the action and its scope, then trigger conditions and caveats. The 'say what you're changing and why' line is a soft instruction rather than voidless, but overall it earns its length.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a destructive, nested-schema mutation with 10 params and no output schema, the description covers the trigger and the confirmation return, which is genuinely useful. It nonetheless leaves most parameter meanings and the full behavioral contract (reversibility, approval flow) underspecified.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is only 30% across 10 params, so the description must compensate. It maps concepts to a few params (arguments→toolArgs, mapping→mapping, cadence→schedule, pause→enabled), but leaves label, workspace, approvalId, toolName, and windowHours unexplained. Partial compensation justifies the baseline 3, not more.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (change) and narrows the resource to a feed's arguments, mapping, cadence, or pause state. This cleanly separates it from siblings like ledger_feeds_create, ledger_feeds_run, ledger_feeds_delete, and ledger_feeds_list without opening any schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives an explicit trigger — reach for this when a feed's last run reports `empty` or `error` — and even diagnoses the `empty` cause (response shape moved). It stops short of naming alternative siblings (e.g. when to preview vs. run vs. recreate), so it is strong context without full when-not guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

ledger_metric_presets_adoptStart tracking metrics from the starting catalogAInspect

Creates the named presets as real metrics the workspace owns, tagged from the catalog. Safe to repeat: a preset the workspace already has comes back already_present rather than creating a second series, and a workspace's own renames and targets are never overwritten. Adopt only what the user agreed to — a metric nobody reads is noise on the Metrics tab, and eight thoughtful ones beat forty. Tell them the metrics start empty and the next step is a feed (ledger_feeds_create) or a first reading. May return needs_confirmation.

ParametersJSON Schema
NameRequiredDescriptionDefault
slugsYesPreset slugs from ledger_metric_presets_list.
workspaceNoWorkspace slug. Ignored for workspace API keys.
approvalIdNo

TDQS

A4.3/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Adds real behavior beyond the annotations: idempotent adoption returning `already_present`, protection of existing renames and targets from being overwritten, and a possible `needs_confirmation` result. These are exactly the side-effect and repeat-safety traits an agent needs and that readOnlyHint/destructiveHint alone do not convey.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Well front-loaded with the core action, then behavior, then guidance. Slightly chatty with the 'eight thoughtful ones beat forty' line, but it is short and each sentence carries information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description carries return-value explanation well (`already_present`, `needs_confirmation`), plus next-step and repeat-safety context. It omits the workspace parameter's meaning and the approval flow that `approvalId` implies, which is a modest gap for a 3-param write tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 67%, so the schema documents slugs (with maxItems 40) and workspace. The description clarifies that slugs come from the preset catalog and that metrics are tagged from it, but says nothing about the `workspace` or `approvalId` parameters, so it does not fully compensate for the gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource: 'Creates the named presets as real metrics the workspace owns, tagged from the catalog.' An agent can distinguish this from ledger_metric_presets_list (read) and ledger_metrics_create (raw creation) without opening any schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives clear adoption guidance ('Adopt only what the user agreed to') and names the next step (ledger_feeds_create) plus the source list implied by the slugs. It does not explicitly contrast with the sibling ledger_metric_starters_adopt, leaving the presets-vs-starters choice to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

ledger_metric_presets_listBrowse the starting catalog of metricsA
Read-only
Inspect

Ready-made metrics a business can start tracking, each with a unit, a cadence, which direction is good, a business surface (facet), and cross-cutting tags. Reach for this when a workspace has few or no metrics, when someone asks what they should be measuring, or when a decision needs a number to settle against and none exists. Filter by facet (where it lives in the business) or tag (what kind of number it is). Suggest a SMALL set — three to six that fit what you know about this business — and say in one sentence why each one, rather than listing the catalog. Adopt with ledger_metric_presets_adopt. These are starting points: a workspace renames and retargets them freely afterwards.

ParametersJSON Schema
NameRequiredDescriptionDefault
tagNoCross-cutting bucket, e.g. retention, cost, speed.
facetNoBusiness surface, matching the Ledger facet taxonomy.

TDQS

A4.6/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnly/openWorld=false/destructive=false, so safety is covered. The description adds real context beyond them: items are starting points that a workspace may rename and retarget, and adoption happens through a separate tool. It stops short of pagination or size limits, but for a small read-only catalog that is a minor gap.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with what the tool returns, then when to use it, then filtering, then delivery guidance. Every sentence carries information, though the block is dense enough that it spans several distinct concerns; still efficient for the amount of guidance an agent needs.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

There is no output schema, but the description compensates by enumerating the fields each preset contains and clarifying these are editable starting points. Combined with annotations covering the safety profile, an agent has everything needed to call and use this tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% with both enums documented, so baseline is 3. The description goes further by explaining the semantic distinction between the two filters — facet = where it lives in the business, tag = what kind of number it is — which helps an agent choose the right filter rather than just read the enum values.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb/resource ('ready-made metrics a business can start tracking') and enumerates what each item carries (unit, cadence, good direction, facet, tags). It also implicitly separates itself from ledger_metrics_list by scoping to workspaces with few or no metrics, so an agent can tell what it returns without opening a schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives three explicit triggering conditions ('few or no metrics', 'someone asks what they should be measuring', 'a decision needs a number to settle against and none exists') and names the downstream tool (ledger_metric_presets_adopt). It even steers the output behavior — suggest 3-6, not the whole catalog.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

ledger_metrics_archiveArchive or unarchive a metricA
Destructive
Inspect

Takes a metric out of the gallery (archived: true) or puts it back (false). Readings stay, and entries that settled against it still read correctly — this is the cleanup for a metric minted once and abandoned, or one the workspace stopped watching. metricId accepts the id or slug from ledger_metrics_list. May return needs_confirmation.

ParametersJSON Schema
NameRequiredDescriptionDefault
archivedYestrue to archive, false to restore.
metricIdYesMetric id or slug (from ledger_metrics_list).
workspaceNoWorkspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys.
approvalIdNoApproval id from a prior needs_confirmation envelope.

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations flag destructiveHint=true, and the description adds genuinely useful context beyond them: readings are preserved and settled entries still read correctly, plus the possibility of a `needs_confirmation` envelope. This clarifies the real blast radius of the 'destructive' flag, though it doesn't detail permission or rate-limit behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core action, then supporting consequences and closure conditions. Three sentences, essentially all earning their place, with only mild density in the middle sentence.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema, but the description covers the mutation semantics, data-retention behavior, and the confirmation flow. Combined with annotations covering the safety profile, an agent has enough to invoke it correctly; permissions and default workspace behavior are left to the schema.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so baseline is 3. The description repeats that `metricId` accepts an id or slug from ledger_metrics_list, which the schema already states, adding no new syntax or constraint detail.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Specific verb+resource: toggles a metric's gallery state via `archived` true/false. Clearly distinct from siblings like ledger_metrics_update or ledger_metrics_create, and the bidirectional semantics (archive vs. unarchive) are stated up front.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives concrete when-to-use guidance — 'cleanup for a metric minted once and abandoned, or one the workspace stopped watching' — which frames the intended scenario well. It stops short of naming an explicit alternative tool (e.g., why not ledger_metrics_update), so it doesn't reach the top bar.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

ledger_metrics_createStart tracking a metricAInspect

Create a metric — a number the workspace watches (triage time, weekly signups, cost per run). Check ledger_metrics_list first; names are unique per workspace. When a user says they want to track or measure something, offer this. May return needs_confirmation.

ParametersJSON Schema
NameRequiredDescriptionDefault
iconNoThe thing the number counts, e.g. phone for calls, landmark for profit.
nameYese.g. "Triage time".
unitNoDisplay unit — "min", "%", "$", "tickets/day".
levelNooutcome = what the business is judged on; driver = moves an outcome; activity = daily work.
drivesNoIds of existing metrics this one moves.
workspaceNoWorkspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys.
approvalIdNo
descriptionNoWhat the number means and where it comes from.

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=false and destructiveHint=false, so the mutation profile is known. The description adds two pieces of behavioral context the annotations cannot: the per-workspace uniqueness constraint on names and the fact that the call 'May return `needs_confirmation`', which prepares the agent for an unusual response shape.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Four short sentences, front-loaded with the verb and resource, then precondition, then trigger, then response caveat. Every sentence adds a distinct, actionable fact with no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a mutation tool with no output schema, the description covers purpose, prerequisite, trigger, and one return-value caveat, which is nearly everything needed. It omits permission requirements and what the created metric's identifier looks like, a minor gap given the annotations already cover the safety profile.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With 88% schema description coverage the schema already documents most parameters, so 3 is the baseline. The description earns above that by defining the `name` concept ('a number the workspace watches') and stating its uniqueness scope, information the schema itself does not convey.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource ('Create a metric') and grounds it with concrete examples (triage time, weekly signups, cost per run), so the agent knows exactly what is produced. It does not, however, differentiate from sibling creators like ledger_metric_presets_adopt or ledger_metric_starters_adopt, so a one-line distinction is missing.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

'Check ledger_metrics_list first; names are unique per workspace' gives a concrete precondition, and 'When a user says they want to track or measure something, offer this' supplies an explicit trigger condition. It stops short of naming when NOT to use it versus preset/starter adoption tools, so it is clear context without exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

ledger_metrics_getRead one Ledger metric with its readingsA
Read-only
Inspect

One metric's definition (unit, kind, direction, target, cadence) plus its readings newest first and the open decisions bound to it. Use before answering 'how is X trending?' or before recording a reading against it. metricId accepts the id or the snake_case slug from ledger_metrics_list. For an event metric whose readings carry labels, labels lists each label and its values, largest total first. Pass slice to get the series for part of it ("new users in DE"), or by to split the series by one label ("new users by country"); either returns series, summed per bucket.

ParametersJSON Schema
NameRequiredDescriptionDefault
byNoEvent metrics only. A label name to split the series by.
sliceNoEvent metrics only. Per label, the values to keep: { country: ["DE", "AT"] } keeps readings from DE or AT; several labels must all match. Names and values come from `labels`.
bucketNoSeries bucket when `slice` or `by` is set. Default week.
metricIdYesMetric id or slug (from ledger_metrics_list).
workspaceNoWorkspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys.

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover the safety profile (readOnlyHint, destructiveHint, openWorldHint), so the description is free to add behavior it uniquely knows: readings are ordered newest first, labels are ordered by largest total, and series results are summed per bucket with a week default. It does not mention result-size limits or pagination on the readings array.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The what-it-returns clause is front-loaded, followed by usage triggers, then parameter mechanics; every sentence carries distinct information with no filler. Density is high but appropriate for a tool with five parameters and no output schema.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description correctly describes the return payload (definition, readings, open decisions, labels, series) and its ordering. It omits how many readings are returned or whether they are capped, which is a minor remaining gap for a read tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3, but the description earns more by explaining that metricId accepts either an id or a snake_case slug, and by giving concrete slice/by examples ('new users in DE', 'new users by country') plus the interaction with bucket. That adds interpretive value beyond the schema's field-level text.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource (read one metric) and enumerates exactly what is returned: definition fields, readings newest-first, and open decisions bound to it. It is clearly distinguishable from ledger_metrics_list, since it also explains that metricId can come from that sibling.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives explicit trigger conditions: 'Use before answering how is X trending?' or before recording a reading against it, which routes the agent away from list/create siblings. It stops short of stating when-not-to-use it or naming the recording tool outright, so it falls just short of a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

ledger_metrics_listList the workspace's Ledger metricsA
Read-only
Inspect

The numbers the workspace watches — each with unit, latest reading, and how many open decisions are bound to it. Consult when a user mentions a number that sounds like a tracked metric, and before recording a reading.

ParametersJSON Schema
NameRequiredDescriptionDefault
workspaceNoWorkspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys.

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, destructiveHint=false, and openWorldHint=false, so the safety profile is covered. The description adds value beyond that by describing the shape of the returned data (unit, latest reading, bound-decision count). It omits pagination/result-size behavior, hence 4 rather than 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tight sentences: the first delivers the content/return summary, the second delivers the usage trigger. Zero filler and front-loaded with the most useful information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema exists, so the description appropriately compensates by naming the returned fields. With one optional, fully documented parameter and covered annotations, the definition is essentially complete; only pagination/ordering behavior is unstated.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and the single 'workspace' parameter is thoroughly documented in the schema (default workspace, personal tokens, API-key behavior). The description adds no parameter-level detail, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description frames the tool as listing 'the numbers the workspace watches,' which together with the title ('List the workspace's Ledger metrics') gives a clear verb+resource. It also enumerates the returned fields (unit, latest reading, open decision count). It does not explicitly distinguish itself from ledger_metrics_get, so it stops short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives concrete triggers – 'when a user mentions a number that sounds like a tracked metric' and 'before recording a reading' – which effectively routes the agent toward ledger_metrics_record_reading as the follow-up action. No explicit exclusions or named alternatives, so not a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

ledger_metrics_record_readingRecord a metric readingAInspect

Record one observation of a tracked metric. metricId accepts a metric id OR its snake_case slug from ledger_metrics_list; an unknown one is a not_found error — check ledger_metrics_list, create it with ledger_metrics_create, or pass createIfMissing: true to mint a "measure" metric at that slug in the same call (an "event" metric when the reading carries labels). The reading automatically lands as evidence on every open decision whose prediction is bound to this metric — so when a user reports a number ("triage is down to 12 minutes"), offer to record it. For event-kind metrics, omit value to count one occurrence. May return needs_confirmation.

ParametersJSON Schema
NameRequiredDescriptionDefault
atNoISO timestamp the reading is for. Defaults to now; set it when backfilling an earlier reading.
keyNoIdempotency key (e.g. "2026-w32") — a repeat write with the same key returns the original reading instead of doubling the series. Use for scheduled/recurring recordings. The key is per set of labels, so one key per day can cover every country.
noteNoWhere the number came from, if worth recording.
valueNoThe observed value, in the metric's unit. Omit for event-kind metrics to record one occurrence.
labelsNoEvent metrics only. What this count is broken down by, e.g. { country: "DE", plan: "pro" } — snake_case names, string values. Lets the metric be read and claimed by slice later. Refused on a measure.
metricIdYesMetric id or slug (from ledger_metrics_list).
workspaceNoWorkspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys.
approvalIdNo
createIfMissingNoMint the metric when the slug is unknown. Default false — an unknown slug is an error, so a typo can't quietly start a second series.

TDQS

A4.6/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Goes well beyond annotations by disclosing the not_found failure mode, the createIfMissing side effect (minting a measure vs event metric based on labels), and the cross-cutting effect that readings land as evidence on open decisions bound to the metric. Also flags a possible needs_confirmation return. The only untold part is idempotent behavior details, which the schema already covers.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the core action and then packs error handling, creation flow, and side effects into dense parentheticals. Every clause carries information, though the single paragraph is heavy and would benefit from light separation.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With 9 params, 89% schema coverage, nested labels, an annotation set, and no output schema, this description covers the required workflow (resolve-or-create), the event/measure distinction, and a non-obvious cross-tool side effect. Nothing an agent needs to invoke it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is high (89%), so baseline is 3, but the description adds real meaning: metricId accepts an id or slug and what happens on unknown values, and labels imply event-kind semantics vs a measure. This clarifies parameter interplay the schema states only in fragments.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Record one observation of a tracked metric'), which cleanly separates it from sibling read/create/list/archive metric tools. The scope is precise and an agent can identify its role without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly routes the agent: check ledger_metrics_list for the id/slug, use ledger_metrics_create to make one, or pass createIfMissing to mint inline. It also gives a triggering condition ('when a user reports a number... offer to record it') and the event-metric invocation pattern (omit value).

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

ledger_metric_starters_adoptStart from a starter metric treeAInspect

Creates a starter's metrics at their levels and links them. Pass slugs to keep only some of its metrics; links are made only where both ends exist. Safe to repeat and safe after a different starter: metrics the workspace already has are left as they are and get linked into the tree. Adopt only what the user agreed to. Tell them the metrics start empty and the next step is connecting a source (ledger_feeds_create) or a first reading. May return needs_confirmation.

ParametersJSON Schema
NameRequiredDescriptionDefault
slugsNoThe starter's metrics to keep. Omit for all of them.
starterYesStarter slug from ledger_metric_starters_list.
workspaceNoWorkspace slug. Ignored for workspace API keys.
approvalIdNo

TDQS

A4.3/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Adds substantial behavior beyond the annotations: idempotency ('safe to repeat and safe after a different starter'), non-destructive merge semantics ('metrics the workspace already has are left as they are'), and a possible `needs_confirmation` return. This is exactly the beyond-annotation disclosure the dimension rewards.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the core action, then scope, then idempotency, then next-step guidance. Dense but every clause carries information; only slightly crowded by the interleaved user-facing instruction.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema, so the mention of `needs_confirmation` is doing useful work, and the idempotency/empty-start caveats cover the main surprises. Missing coverage of `approvalId` and workspace scoping keeps it from a 5.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 75%, so the schema already documents `slugs` and `starter`. The description restates the slug-subset behavior but adds nothing on `approvalId` or `workspace`, both of which stay undocumented in the description. Baseline 3 fits when the schema carries most of the load.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('Creates a starter's metrics at their levels and links them') with the linking behavior included, which distinguishes it from a plain metric-create tool. An agent can tell what it does without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives clear operating context: 'Adopt only what the user agreed to', `slugs` to subset, and follow-up steps (connect a source or a first reading). It stops short of explicitly naming ledger_metric_presets_adopt as the alternative adopt path, so routing between the two adopt tools is left to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

ledger_metric_starters_listBrowse starter metric treesA
Read-only
Inspect

Starter trees by shape of business or team (subscription software, services firm, online store, sales team, service operation). Each lists its metrics with a level (outcome / driver / activity) and the links between them (which number moves which). Reach for this before ledger_metric_presets_list when a workspace has no metrics yet or asks how its numbers fit together: pick the starter that matches what you know about the business, describe its tree in a sentence or two, and offer to adopt it with ledger_metric_starters_adopt, dropping any metric that doesn't fit.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4.6/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, destructiveHint=false and openWorldHint=false, so the safety profile is covered. The description adds real value beyond this by disclosing the shape of the returned content (metrics tagged by level plus inter-metric links) and the intended follow-up adoption flow. It stops short of listing pagination or count limits, hence a 4 rather than 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the resource and its contents, then the routing condition, then the workflow. Every sentence earns its place, though the closing adoption advice ('describe its tree in a sentence or two... dropping any metric that doesn't fit') edges toward prescriptive verbosity for a read-only list tool.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema present, the description carries the burden of describing returns and does so (metrics, levels, links between them). Combined with the explicit when-to-use condition and the named follow-up tool, an agent has everything needed to select and invoke it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so there is nothing for the description to disambiguate; the baseline for a no-parameter tool is 4. The description's enumeration of business shapes usefully hints at the conceptual categories the caller will see, but these are not input arguments.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Names a specific resource (starter metric trees) and states exactly what they contain: business/team shapes, each metric's level (outcome/driver/activity), and the links between numbers. It distinguishes itself from the sibling ledger_metric_presets_list by naming it directly, so an agent can route correctly without opening either schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly states when to reach for this tool ('before ledger_metric_presets_list when a workspace has no metrics yet or asks how its numbers fit together') and names the alternative. It also spells out the downstream workflow — pick a matching starter, describe its tree, then offer ledger_metric_starters_adopt — leaving nothing to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

ledger_metrics_updateEdit a metric's definitionA
Destructive
Inspect

Corrects a metric's name, unit, description, kind (measure / event), direction (which way is good), target, cadence, icon, level, or the metrics it drives (its place in the metric tree). Readings are untouched. Only what you pass changes. The slug is not editable here — external writers address metrics by slug. Fails with conflict when a rename collides with a live metric. metricId is the id from ledger_metrics_list. May return needs_confirmation.

ParametersJSON Schema
NameRequiredDescriptionDefault
iconNoThe mark the metric wears in every tool: the thing it counts (phone for calls, landmark for profit). "" goes back to the catalog's icon.
kindNomeasure = a value each reading; event = a count of occurrences.
nameNo
unitNo"min", "%", "$", "tickets/day".
levelNooutcome = what the business is judged on; driver = a number that moves an outcome; activity = daily work. "" unplaces it.
drivesNoIds of the metrics this one moves. Replaces the current set; pass the full list. Loops are refused.
targetNoGoal value in the metric's unit; null clears.
cadenceNoHow often a reading is expected.
metricIdYesMetric id (from ledger_metrics_list).
directionNoup = higher is better, down = lower is better, none.
workspaceNoWorkspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys.
approvalIdNoApproval id from a prior needs_confirmation envelope.
descriptionNo

TDQS

A4.4/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Goes well beyond the destructiveHint=true annotation: it discloses partial-update behavior, that readings are unaffected, that the slug is immutable, the conflict-on-rename failure mode, and that a needs_confirmation envelope may be returned. These are exactly the behavioral traits an agent needs before invoking a mutating tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the verb and editable fields, then layers constraints (immutability, conflict, confirmation) in tight declarative sentences. Every clause carries information; nothing is padded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 13-parameter mutation tool with no output schema, it covers the essentials: scope of change, immutable fields, failure mode, and the confirmation flow. It could tie approvalId more explicitly to the needs_confirmation envelope, but the schema covers that parameter.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 85%, so the schema already documents most parameters (including kind, direction, drives, cadence). The description restates the field list and adds only the 'metric tree' framing for drives, so it adds marginal value beyond structured fields, which is the baseline-3 case.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (corrects/edits) and resource (a metric's definition) and enumerates the editable fields. It explicitly distinguishes itself from reading tools with 'Readings are untouched,' so an agent can separate it from ledger_metrics_record_reading without opening schemas.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides clear context: partial-update semantics ('Only what you pass changes'), the source for metricId ('from ledger_metrics_list'), and a hard exclusion (slug not editable). It does not explicitly say when to prefer this over ledger_metrics_create or ledger_metrics_archive, so it stops short of full when/when-not guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

ledger_plan_addAdd dated work to a Ledger decisionAInspect

Adds a plan item to an OPEN decision (the work that carries it out: a post, a launch step, an email). Omit entryId for a standalone date (a holiday, an event you're only watching). Items are all-day unless you pass startTime with timeZone (and optionally endTime), for a webinar, a scheduled post, or a launch at noon. If the date costs money or time and comes with an expectation, record it as a decision with ledger_entries_create instead. Use repeatWeeklyUntil for a weekly cadence: it writes one row per week (max 60), each movable on its own. Fails with conflict once the decision is settled. May return needs_confirmation.

ParametersJSON Schema
NameRequiredDescriptionDefault
urlNoWhere the work lives (draft, deck, published post).
noteNo
tagsNoWhat kind of work it is: blog, video, social, event… Reuse the workspace's existing tags (see ledger_plan_list).
dueOnYesA calendar day, YYYY-MM-DD.
titleYes
endTimeNoSame-day end, after startTime.
entryIdNoThe open decision this carries out.
timeZoneNoIANA zone the time is in (America/Chicago). Required with startTime; use the user's own zone.
startTimeNoOmit for an all-day item. 24-hour HH:MM in timeZone.
workspaceNoWorkspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys.
approvalIdNoApproval id from a prior needs_confirmation envelope.
ownerUserIdNoA workspace member's user id.
repeatWeeklyUntilNoAlso add a copy every 7 days through this day.

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already indicate it's not read-only and not destructive. The description adds crucial behavioral details: fails with conflict once decision is settled, may return needs_confirmation, and repeatWeeklyUntil writes one row per week (max 60), each movable. This goes beyond what annotations provide.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core action and immediately addresses entryId. The description is dense but well-structured, covering multiple scenarios efficiently. It could be slightly more concise, but every sentence adds necessary context for a tool with 13 parameters.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given 13 parameters, no output schema, and annotations that only cover safety hints, the description does a good job explaining key behaviors, failure modes, and parameter usage. It doesn't cover all parameters (e.g., workspace, tags), but those are documented in the schema. It's nearly complete for an agent to call correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 85%, so baseline is 3. The description adds value by explaining entryId's purpose ('the work that carries it out'), when to omit it ('standalone date'), the all-day vs timed distinction (startTime/timeZone/endTime), and repeatWeeklyUntil's weekly row behavior (max 60). Some parameters like workspace are not covered, but the description enhances understanding of key parameters.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource: 'Adds a plan item to an OPEN decision'. It also clarifies the conceptual model (plan item = work that carries out a decision) and distinguishes from the sibling ledger_entries_create by describing when to use which. An agent can immediately understand this adds a dated work item, not a decision or evidence.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly states when to use entryId vs omit it (standalone dates), when to use all-day vs timed, and when to redirect to ledger_entries_create instead. However, it doesn't mention ledger_plan_update or ledger_plan_delete as alternatives for editing/removing items, though those siblings exist.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

ledger_plan_deleteRemove a Ledger plan itemA
Destructive
Inspect

Removes one plan item while its decision is open, for an item added by mistake or work that's no longer planned. If the work was planned and then dropped, prefer ledger_plan_update with status 'skipped' so settlement can see it. Fails with conflict once the decision is settled. May return needs_confirmation.

ParametersJSON Schema
NameRequiredDescriptionDefault
itemIdYes
workspaceNoWorkspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys.
approvalIdNoApproval id from a prior needs_confirmation envelope.

TDQS

A4.6/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true and readOnlyHint=false, so the safety profile is known. The description adds real context beyond that: the operation is only valid while the decision is open, it fails with a conflict after settlement, and it may return a needs_confirmation envelope. It stops short of spelling out the confirmation/approval flow explicitly, keeping it at a 4.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three tightly written sentences: purpose, the preferred alternative, then the failure/confirmation behavior. No filler and the destructive scope is front-loaded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a destructive tool with no output schema, the description covers the important return states (conflict, needs_confirmation) and the mutation constraint. The only gap is the expected format/identity semantics of itemId, which is a minor omission.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 67%, with workspace and approvalId documented but itemId bare. The description compensates by linking the needs_confirmation return to the approvalId parameter, which is meaning the schema alone does not supply.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (removes) and resource (plan item) and adds the key scoping condition 'while its decision is open', which an agent can use to tell it apart from ledger_plan_update and ledger_plan_add without opening schemas.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly names the alternative and the condition that selects it: 'If the work was planned and then dropped, prefer ledger_plan_update with status \'skipped\''. It also states the failure condition (conflict once settled), so both when-to-use and when-not-to-use are covered.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

ledger_plan_listRead the Ledger calendar or one decision's planA
Read-only
Inspect

Pass entryId for one decision's plan items, or from + to (YYYY-MM-DD, at most about a year apart) for the calendar: decision spans (recorded day → expectation deadline) plus every dated item inside the window, including standalone dates with no decision. Filter the calendar by ownerUserId to answer 'what's mine this week' or to find tomorrow's items to draft. Never use it to compare or rank people's output.

ParametersJSON Schema
NameRequiredDescriptionDefault
toNoA calendar day, YYYY-MM-DD.
tagNoOnly items with this tag (blog, video, event…).
fromNoA calendar day, YYYY-MM-DD.
entryIdNoOne decision's plan.
workspaceNoWorkspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys.
ownerUserIdNoOnly items owned by this member.

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already establish the safe read-only, non-destructive, closed-world profile, and the description adds real behavioral detail beyond them: what the calendar actually returns (decision spans from recorded day to expectation deadline, plus standalone dated items) and the roughly one-year window limit. Pagination and result volume are not addressed, so not a 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the entryId-vs-from/to decision, then the return semantics, then the filtering use case and the exclusion. Dense but nearly every clause earns its place; the single long paragraph is slightly harder to scan than bulleted alternatives.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a read-only list tool with a fully documented schema and annotations covering safety, the description supplies the missing pieces: the two call modes, the window constraint, the returned item types, and the anti-use case. No output schema exists, so return-shape detail here is a bonus rather than a gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3, but the description adds meaning the schema lacks: entryId scopes to one decision's plan, from/to must be paired and kept to about a year, and ownerUserId is framed by its use case. Only `tag` goes unexplained beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific read operation on the ledger plan/calendar and defines its two modes (per-decision via entryId vs. calendar window via from/to) in the first sentence. It is clearly distinguishable from siblings ledger_plan_add/update/delete, which mutate the same resource.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly routes the caller: pass entryId for one decision, from+to for the calendar, and use ownerUserId for 'what's mine this week' or 'tomorrow's items to draft'. It also states a when-not — never use it to compare or rank people's output — so both selection and exclusion are covered.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

ledger_plan_updateMove, rename, re-own, or mark a Ledger plan itemA
Destructive
Inspect

Edits one plan item while its decision is open: move it (dueOn), rename it, change the owner (null clears), attach the link, or set status (planned | done | skipped). Mark done only when the user says it shipped; a link alone doesn't mean done. Fails with conflict once the decision is settled. May return needs_confirmation.

ParametersJSON Schema
NameRequiredDescriptionDefault
urlNo
noteNo
tagsNoReplaces the item's tags.
dueOnNoA calendar day, YYYY-MM-DD.
titleNo
itemIdYes
statusNo
endTimeNo
timeZoneNoRequired when setting startTime.
startTimeNonull makes it all-day again.
workspaceNoWorkspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys.
approvalIdNoApproval id from a prior needs_confirmation envelope.
ownerUserIdNo

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true, so the description's added value is the conflict behavior ('Fails with conflict once the decision is settled') and the possibility of a `needs_confirmation` return, which the agent must handle. It does not spell out that field overwrites are irreversible, so it stops short of a 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three dense sentences, each carrying distinct information (capabilities, guardrail, failure modes), with the core capability front-loaded. No filler or restatement of the tool name.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

There is no output schema, and the description covers the two return-path facts an agent needs (conflict on settled decisions, `needs_confirmation` envelope, which pairs with the approvalId parameter). It omits the workspace-token requirement, but that is already documented in the schema.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With 13 parameters and only 46% schema description coverage, the description partly compensates: it maps dueOn to move, title to rename, url to link, status to the three enum values, and adds the non-obvious 'null clears' rule for ownerUserId. Only `note` and `itemId` are left undocumented in both places, so the compensation is incomplete but substantial.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Opens with a specific verb+resource ('Edits one plan item') and enumerates the exact mutations supported (move via dueOn, rename, re-own, attach link, set status). It is trivially distinguishable from siblings ledger_plan_add, ledger_plan_delete, and ledger_plan_list.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives a real precondition ('while its decision is open') and a substantive when-to-use rule ('Mark done only when the user says it shipped; a link alone doesn't mean done'). It never names a sibling tool as the alternative, so it falls short of a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_docsList ZeroWidth docs pagesA
Read-only
Inspect

Enumerate all available docs pages, optionally filtered by product (e.g. 'compass', 'legal', 'overview'). Use this to discover what slugs exist before calling get_doc. No authentication required.

ParametersJSON Schema
NameRequiredDescriptionDefault
productNoOptional product slug filter (e.g. 'compass', 'legal', 'overview').

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, destructiveHint=false, and openWorldHint=false, so the safety profile is covered. The description adds context beyond them with 'No authentication required,' a genuinely useful operational fact for callers, though it says nothing about pagination or result size for a full enumeration.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three tight sentences with purpose, routing guidance, and the auth fact front-loaded in order of importance. No filler or repetition of the title.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a read-only, single-optional-param listing tool whose annotations cover safety, the description is nearly sufficient. Without an output schema it could note the return shape (e.g. that results are slugs/pages), but the 'slugs' reference largely covers that, so only a small gap remains.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the schema already documents the single 'product' parameter. The description's example values ('compass', 'legal', 'overview') duplicate the schema description verbatim, adding no meaning beyond it. Baseline 3 applies when the schema does the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('Enumerate') and resource ('all available docs pages') with scope, plus the optional product filter. It names the sibling get_doc and frames itself as the discovery step before retrieval, letting an agent distinguish it from single-doc fetch without opening either schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly states when to use it: 'to discover what slugs exist before calling get_doc,' giving a clear dependency flow and naming the alternative. It does not address when NOT to use it or whether search_docs is a better discovery path, leaving a minor gap.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

search_docsSearch ZeroWidth docsA
Read-only
Inspect

Search ZeroWidth product documentation. Returns matching pages with title, slug, public URL, and a query-relevant snippet. Use this when the user asks about a ZeroWidth product (Compass, Workbench, Caliper, Prism, Ledger, Napkin, zv1), an API behavior, or a policy. No authentication required — the docs corpus is public.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNoMax number of results. Defaults to 10.
queryYesSearch query — keywords or natural-language phrase.

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, destructiveHint=false, and openWorldHint=false, so the safety profile is covered. The description adds useful non-annotation context: 'No authentication required — the docs corpus is public' and the shape of returned results (title, slug, public URL, snippet). It does not mention pagination or ordering, but it goes beyond the structured fields.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is four short sentences, front-loaded with the core action, followed by return information, usage trigger, and auth note. Every sentence earns its place, and there is no repetition or filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a two-parameter read-only search tool, the description covers purpose, return format, auth requirements, and usage triggers. It does not explain how this differs from search_workspace or list_docs/get_doc, and it lacks result-ordering or empty-result behavior, but it is otherwise sufficient for correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, and both parameters are documented in the schema itself. The description says the query returns a 'query-relevant snippet' but adds no syntax, format, or constraint details beyond what the schema already provides, so the baseline of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: 'Search ZeroWidth product documentation.' It also names the return fields and the product scope with concrete examples (Compass, Workbench, Caliper, etc.), which lets an agent distinguish it from siblings like list_docs, get_doc, and search_workspace without opening schemas.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives clear usage context: 'Use this when the user asks about a ZeroWidth product ..., an API behavior, or a policy.' However, it does not name alternative tools (search_workspace, list_docs, get_doc) or state when not to use this tool, so it stops short of explicit routing guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

search_workspaceSearch the whole workspaceA
Read-only
Inspect

Finds entities across every tool by name in one call — Workbench flows, Compass pages, Caliper datasets, evals, rubrics, reviews, specs and sources (apps sending agent traces), Ledger entries, Napkin sketches and decks. Use it FIRST when the user names something without saying where it lives ('the onboarding flow', 'that invoice page'); reach for a tool's own list only when you already know the tool. Each hit carries its id, kind, and workspace-relative path, so the id feeds the matching *_get tool and the path makes a link. Results only include what the user can see, and only kinds this token may read.

ParametersJSON Schema
NameRequiredDescriptionDefault
qYesCase-insensitive substring matched against names/titles.
kindsNoRestrict to these kinds (flow, page, dataset, eval, entry, board). Omit to search everything.
limitNo
workspaceNoWorkspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys.

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and destructiveHint=false, so safety is covered. The description adds genuinely new behavioral context beyond that: results are filtered to what the user can see and to kinds the token may read, and it discloses the hit shape (id, kind, workspace-relative path).

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three dense, front-loaded sentences with no filler: purpose and coverage first, routing second, return/scope semantics last. Every clause carries information the agent can act on.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema exists, so the description carries the return-value burden and does so by describing the id/kind/path tuple and how the id feeds *_get tools. Combined with the permission scoping note, an agent has everything needed to call and use this correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 75%, so most parameters are already documented structurally. The description reinforces name/title matching but adds no new syntax or format detail for kinds, limit, or workspace, so the baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Opens with a specific verb (Finds) and resource (entities across every tool by name), then enumerates the concrete kinds covered (flows, pages, datasets, evals, rubrics, etc.). An agent can immediately distinguish this cross-tool search from the many per-tool list siblings.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly states when to use it ('Use it FIRST when the user names something without saying where it lives'), gives concrete examples ('the onboarding flow'), and names the alternative plus its selection condition ('reach for a tool's own list only when you already know the tool'). This is textbook when/when-not/alternative routing.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 40 tool updates
    • First observedcomments_create
    • First observedcomments_list
    • First observedcomments_resolve
    • First observedentity_tags_browse
    • First observedentity_tags_get
    • First observedentity_tags_set
    • First observedget_doc
    • First observedledger_entries_add_evidence
    • First observedledger_entries_create
    • First observedledger_entries_draft
    • First observedledger_entries_get
    • First observedledger_entries_list
    • First observedledger_entries_retract
    • First observedledger_entries_settle
    • First observedledger_entries_update
    • First observedledger_feed_sources_list
    • First observedledger_feeds_backfill
    • First observedledger_feeds_create
    • First observedledger_feeds_delete
    • First observedledger_feeds_list
    • First observedledger_feeds_preview
    • First observedledger_feeds_run
    • First observedledger_feeds_update
    • First observedledger_metric_presets_adopt
    • First observedledger_metric_presets_list
    • First observedledger_metric_starters_adopt
    • First observedledger_metric_starters_list
    • First observedledger_metrics_archive
    • First observedledger_metrics_create
    • First observedledger_metrics_get
    • First observedledger_metrics_list
    • First observedledger_metrics_record_reading
    • First observedledger_metrics_update
    • First observedledger_plan_add
    • First observedledger_plan_delete
    • First observedledger_plan_list
    • First observedledger_plan_update
    • First observedlist_docs
    • First observedsearch_docs
    • First observedsearch_workspace

Related MCP Connectors

Related MCP Servers

  • F
    license
    Not graded
    quality
    B
    maintenance
    Captures AI-assisted work into a searchable ledger and exposes it via MCP for querying past conversations and decisions.
    -
  • F
    license
    Not graded
    quality
    B
    maintenance
    Enables users to turn a decision into an append-only judgment record that keeps their own wording and premises separate from the AI's suggestions, then returns on a chosen event or fallback date so reality and later answers can be appended without rewriting the original. Runs over any MCP-capable assistant with records stored in a local project ledger, and never scores or profiles the person.
    672 npm
    1
    -
  • A
    license
    Not graded
    quality
    Not graded
    maintenance
    Enables users to record probabilistic judgments in an immutable ledger, automatically score them against outcomes, and analyze systematic biases across six classification layers. It supports natural-language interaction for logging predictions, checking randomness, and auditing data integrity while avoiding advice or replacing user judgment.
    Apache 2.0
Try in Browser

Glama MCP Gateway

Add one secure layer between your agents and this server.

Resources