Skip to main content
Glama

Server Details

Connect Claude, Cursor, and other MCP-compatible coding agents to DevTune AI search visibility data with the Streamable HTTP Model Context Protocol server.

DevTune exposes a Model Context Protocol (MCP) server so that AI agents can query your AI search visibility data directly. This lets agents like Codex, Claude, Cursor, and other MCP-compatible tools access your metrics and data directly.

Ownership verified
Status
Healthy
Uptime
99.6% over 22 days
Last Tested
Transport
Streamable HTTP · MCP 2025-11-25
URL

TDQS

B3.3/5.0

Scored across 66 tools

Disambiguation3/5

Many tools cover closely related metric retrieval (citation stats/analysis/top pages/domains, visibility summary/response/list, rollups vs page metrics vs audit summary), and devtune_get_citation_analysis is an explicit alias for top citation pages. Descriptions are detailed and often clarify scope, but the sheer number of get_* tools creates real misselection risk.

Naming Consistency5/5

Every tool uses the devtune_ prefix followed by a consistent snake_case verb_noun pattern (get_, list_, create_, update_, delete_, start_, send_, etc.). The convention is predictable throughout.

Tool Count1/5

66 tools is an extreme mismatch for a well-scoped MCP server; the rubric places 50+ tools at the lowest score. The surface is far larger than an agent can reasonably navigate, with many variants of metrics, actions, Dex, library, and audit operations.

Completeness3/5

Coverage is broad (metrics, audits, actions, Dex, Library, Knowledge Profile, prompt analysis, agents), but notable lifecycle gaps exist: no tool to accept or archive an action (suggested -> accepted transition is missing), and no delete operation for Library items. Agents may hit dead ends in core workflow transitions.

Available Tools

66 tools
devtune_acknowledge_retrieval_restrictionA
Idempotent
Inspect

Acknowledge an intentional retrieval restriction for only the exact URLs displayed in an audit finding. Requires explicit customer confirmation and actions.write. Preserves the finding, evidence and score. Never infer customer intent.

ParametersJSON Schema
NameRequiredDescriptionDefault
urlsYes
findingIdYes
auditRunIdYes
confirmedIntentYes

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already mark the tool as idempotent and non-destructive, and the description adds valuable beyond-annotation context: it requires actions.write, preserves the finding/evidence/score, and forbids inferring customer intent. This gives the agent a clear operational and safety profile.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three concise sentences front-load the purpose, then state prerequisites and behavior. There is no filler, redundancy, or unhelpful repetition of schema fields.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a mutation tool with no output schema, it covers confirmation, permissions, scope, and non-destructive behavior. It does not explicitly spell out the relationship between auditRunId and findingId or return behavior, but those gaps are minor given the self-explanatory names and idempotent annotation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With 0% schema description coverage, the description compensates for the most nuanced parameters: urls must be exactly the URLs shown in the audit finding, and confirmedIntent must reflect explicit customer confirmation rather than inference. auditRunId and findingId are left to their self-explanatory names.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('acknowledge'), a clear resource ('retrieval restriction'), and a precise scope ('only the exact URLs displayed in an audit finding'). This makes its purpose distinct from siblings like reconsider_retrieval_restriction.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It clearly states when to use the tool: with explicit customer confirmation, for exact displayed URLs, and with actions.write permission. It also gives a 'never infer customer intent' guardrail, though it does not explicitly name alternative tools or contrast with reconsider_retrieval_restriction.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

devtune_attach_outcome_urlAInspect

Attach a published URL to a completed action. Obtain human approval first. Measurement starts at the action completion date (the declared went-live date), independently of crawling or page-tracking allowance. publishedUrls replaces the published-URL list; retain existing entries. Existing measurement dates cannot be overwritten. Requires actions.write.

ParametersJSON Schema
NameRequiredDescriptionDefault
urlYes
actionIdYes
changeEventIdNo
publishedUrlsYes

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure. It discloses that measurement is tied to the action completion date, that publishedUrls replaces the existing list (with a warning to retain entries), and that existing measurement dates cannot be overwritten. It also notes the required permission (actions.write). These are significant behavioral details that an agent needs to know before calling this mutation tool. It does not mention reversibility or idempotency, but the coverage is solid.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is succinct and front-loaded with the core action. Each sentence adds essential information: the action itself, the human approval prerequisite, the measurement timing nuance, the replacement behavior for publishedUrls, the immutability of existing dates, and the permission requirement. There is no fluff or redundancy; every word earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a mutation tool with four parameters, no output schema, and no annotations, the description covers key contextual aspects: prerequisites (human approval), behavioral details (measurement timing, list replacement, immutability), and permissions. It does not describe return values or error conditions, but without an output schema these may not be critical. The description is adequate for an agent to invoke the tool correctly, though it could mention what happens if the action is not completed or if the URL already exists.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It specifically explains the semantics of publishedUrls (it replaces the list and existing entries must be retained), which is valuable. However, it does not explain the semantics of actionId, url, or changeEventId beyond the obvious from names and context. Since the schema has no descriptions, the description could have elaborated on these parameters, especially changeEventId, but it covers the most complex one. It adds partial value but not full compensation for the zero coverage.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource ('Attach a published URL to a completed action'), clearly stating what the tool does. It distinguishes itself from the many 'get_', 'list_', and 'update_' siblings by focusing on the unique action of attaching a URL to a completed action. No other sibling covers this operation, so purpose is unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides some usage context: it instructs to 'Obtain human approval first' and clarifies that measurement starts at the completion date, which are important prerequisites and behavior notes. However, it does not explicitly mention when to use this tool versus alternatives (though no direct alternatives exist), nor does it state when not to use it. The guidance is helpful but not exhaustive regarding selection criteria.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

devtune_claim_content_changeBInspect

Claim a detected change as a completed measured action. Obtain human approval first. Re-claim restores the original action and measurement without overwriting their content. Returns a conflict when the page has another measurement completed within 14 days of the change, or a concurrent measurement update requires a retry. Requires actions.write.

ParametersJSON Schema
NameRequiredDescriptionDefault
titleYes
descriptionYes
changeEventIdYes

TDQS

B3.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the burden of disclosing behavioral traits. It reveals the mutation nature ('claim', 'restores'), explains the re-claim side effect (preserves original content), lists conflict cases, and states the required permission (actions.write). This adds significant context beyond the schema, though it stops short of detailing the success response or other side effects.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is three tight sentences with no filler. The primary purpose is front-loaded, and the most consequential behavioral nuances (re-claim, conflicts, permission) are compactly conveyed. Minor density around the conflict sentence does not hurt clarity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a mutation tool with no annotations, no output schema, and three undocumented parameters, the description omits essential information: meaning of each parameter, what a successful response looks like, and how the 'claim' materially alters state beyond marking. It covers prerequisites and failure modes but is far from complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0% and the description provides no explanation of the three parameters (title, description, changeEventId). An agent cannot determine what each field represents or how they relate to the change being claimed, leaving a critical gap for correct invocation.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the specific action ('Claim a detected change') and defines it as marking a change as a completed measured action. It distinguishes the tool's purpose from that of siblings like devtune_unclaim_content_change by describing the claim and re-claim behavior explicitly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It provides a prerequisite ('Obtain human approval first') and describes conditions under which a conflict arises (14-day rule, concurrent update), implying when the tool is appropriate. However, it does not explicitly contrast with alternatives such as unclaim or other action tools, leaving some selection guidance to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

devtune_create_prompt_actionAInspect

Create or reuse an action from a tracked prompt. Only set generateBrief when the customer requests metered brief generation (reserves 2 credits). Saves the action even when brief generation fails. Does not publish content.

ParametersJSON Schema
NameRequiredDescriptionDefault
titleYes
promptIdYes
targetUrlNoExisting page to improve. Omit to create new content. Locked pages cannot be targeted.
descriptionYes
generateBriefNoDefaults to false. Also queue a brief using content credits.
acknowledgedHighStakesNoSet true only after the customer explicitly acknowledges changing a high-stakes target page.

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure. It discloses two key traits: saves the action even when brief generation fails, and does not publish content. It also mentions credit reservation for generateBrief, adding cost implications. This is substantial value beyond the schema, though it does not cover all potential edge cases (e.g., idempotency of reuse).

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is three sentences: the first states the core purpose, the second gives a critical parameter guideline, and the third discloses behavior. It is front-loaded with the primary action and contains zero fluff. Every sentence earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a create-action tool with 6 parameters and no output schema, the description covers the purpose, key parameter usage, and critical failure/success behavior. It lacks details on the response format or what 'reuse' entails, but given the simplicity and absence of an output schema, it is adequately complete. The description does not overpromise or leave the agent guessing on essential call decisions.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 50%, and the description adds meaningful context for generateBrief by clarifying the credit cost and the condition under which it should be set, complementing the schema's default and queue description. The other parameters (promptId, title, description) are self-explanatory strings, and targetUrl already has a schema description. The description effectively compensates for the partial coverage.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states a specific verb-resource pair: 'Create or reuse an action from a tracked prompt.' It identifies the input source (tracked prompt) and distinguishes from sibling tools like devtune_get_actions or devtune_get_prompt_actions by implying a write operation. The additional behavioral notes (saves even on failure, does not publish) further clarify its role.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides explicit conditional guidance: 'Only set generateBrief when the customer requests metered brief generation (reserves 2 credits).' It also clarifies that saving occurs even if brief generation fails and that content is not published, implying publishing is a separate step. However, it does not name alternative tools or state when not to use this tool, so it lacks explicit exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

devtune_decide_dex_handoffBInspect

Confirm or cancel one pending specialist-agent handoff in a Dex conversation.

ParametersJSON Schema
NameRequiredDescriptionDefault
decisionYes
threadIdYes
handoffIdYes

TDQS

B3.2/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden of behavioral disclosure. 'Confirm or cancel' implies a state-changing mutation, but the description does not disclose what confirming does (does it dispatch a specialist, consume the handoff?), what canceling reverts, whether the operation is idempotent, or what happens if the handoff is not actually pending. This is a consequential decision tool with no behavioral detail.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single efficient sentence with no filler or repetition. It is appropriately sized, though the brevity comes at the cost of missing parameter and behavioral detail. Concise but not comprehensive.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a mutation tool with three required parameters, no output schema, and zero annotations, the description is inadequate. It does not explain the parameters, the consequences of each decision value, or any prerequisites (e.g., the handoff must exist and be pending). An agent would not know the full effect of invoking it.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It implicitly maps to the decision enum ('confirm or cancel' corresponds to the decision property) and 'one pending' hints at handoffId, but threadId is entirely unexplained, and there is no clarity on how the three parameters relate. The description adds only minimal value over the raw schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('Confirm or cancel') and a specific resource ('one pending specialist-agent handoff in a Dex conversation'). It clearly distinguishes from siblings like devtune_list_dex_handoffs (which lists rather than decides) and devtune_start_dex_conversation. An agent can tell what this does without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The phrase 'in a Dex conversation' gives some contextual anchor, and 'pending' implies it is used after handoffs have been surfaced (e.g., via list_dex_handoffs). However, it provides no explicit guidance on when to use this tool versus alternatives, and no exclusions or when-not-to-use conditions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

devtune_decide_dex_memoryCInspect

Save or dismiss one durable memory offer from a Dex conversation. Saves verify the text against the persisted agent event.

ParametersJSON Schema
NameRequiredDescriptionDefault
taskIdYes
offerIdYes
decisionYes
threadIdYes
offeredTextYes

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden of behavioral disclosure. It mentions that saves verify the text against the persisted agent event, which is helpful, but it does not disclose side effects of dismissing, idempotency, reversibility, permission requirements, or what happens on repeated calls. This is a significant gap for a mutation tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two short sentences, front-loaded with the core purpose and a behavioral note. There is no fluff or repetition. It is appropriately sized for the information it conveys, though it omits important details that are penalized elsewhere.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool has 5 required parameters, no output schema, and no annotations, the description is far from complete. It does not explain parameter semantics, usage context, or expected behavior beyond the basic action. An agent would struggle to invoke it correctly without external knowledge, making the definition inadequate for a tool of this complexity.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must explain the parameters. It does not mention any of the five parameters (taskId, offerId, decision, threadId, offeredText) by name or explain their roles. The only hint is the verification note, which loosely relates to offeredText, but it is not explicit. The agent has no guidance on what values to provide or how they interact.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states a specific verb (save or dismiss) and a distinct resource (durable memory offer from a Dex conversation), which distinguishes it from sibling tools like devtune_decide_dex_handoff. It also adds a meaningful behavioral detail about verification on save, making the purpose unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives no explicit guidance on when to use this tool versus alternatives. It does not mention prerequisites, when to choose save vs dismiss, or direct the agent to other tools in edge cases. The usage context is only implied by the name and the action, which is insufficient given the large sibling set.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

devtune_delete_dex_threadAInspect

Delete your private Dex chat after stopping its live run. Billing and audit records are retained. Requires a user-attributed API key owned by the chat creator.

ParametersJSON Schema
NameRequiredDescriptionDefault
threadIdYes

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden of behavioral disclosure. It clearly indicates the destructive nature of the operation, notes that 'Billing and audit records are retained,' and specifies the ownership/auth prerequisite. This is strong transparency for a delete operation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is three short sentences with no filler. The primary action is front-loaded, followed by key retention and authorization details. Every sentence earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a single-parameter delete tool with no output schema and no annotations, the description covers the essential context: what is deleted, when it can be deleted, what is retained, and who is allowed to delete it. No critical information is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, but there is only one parameter, threadId, whose UUID format is documented in the schema. The description adds meaningful context by specifying 'your private Dex chat' and the ownership requirement, even though it does not explicitly name threadId as the target.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a precise action and resource: 'Delete your private Dex chat'. It clearly distinguishes this deletion tool from sibling tools like devtune_list_dex_threads, devtune_rename_dex_thread, and devtune_send_dex_message.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides clear usage context: it should be used 'after stopping its live run' and only with 'a user-attributed API key owned by the chat creator.' It does not explicitly name alternatives or when-not-to-use conditions, but no sibling deletion tool exists, so the guidance is sufficient.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

devtune_generate_action_briefAInspect

Trigger brief generation for a specific action. Returns whether the brief was already ready or has been queued or is already generating in the background.

ParametersJSON Schema
NameRequiredDescriptionDefault
actionIdYesUUID of the action to generate the brief for.
briefStyleNoOptional brief style override. Defaults to best_fit and must respect project content preferences.

TDQS

A3.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the disclosure burden. It clearly reveals that the tool may trigger background processing and reports whether the brief is ready, queued, or already generating. This gives meaningful behavioral context, though it does not mention failure/error behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise and front-loaded: it states the trigger action and then the value of the return across three states. No additional filler or redundant restatement appears.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple trigger tool with two documented parameters and no output schema, the description is sufficient for an agent to understand the invocation's effect and the possible return state. It covers the asynchronous background aspect that is not present in the structured fields.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is a 3. The description does not add further parameter-level semantics beyond the schema's UUID and style details, which is acceptable because the schema already documents both parameters completely.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description identifies the specific action (brief generation for an action) and clarifies the return semantics: it returns whether the brief is ready, queued, or already generating. This sufficiently distinguishes it from siblings like devtune_get_action_brief.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no explicit guidance on when to call this tool versus devtune_get_action_brief or other action-related siblings. The status-oriented return implies usage for requesting or checking generation, but no alternative or condition is named.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

devtune_get_404_demandAInspect

Get project-scoped 404 Demand from the current materializer window, separated into credible content-intent clusters and likely scanner or technical noise, with materialization state, machine-traffic sensor coverage, and inferred project relevance. Clusters prioritize project fit before request volume; unrelated demand remains visible but is excluded from automatic action suggestions.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4.5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden of behavioral disclosure. It discloses that demand is separated into content-intent and noise clusters, that project fit is prioritized over request volume, and that unrelated demand remains visible but is excluded from automatic action suggestions. It also mentions materialization state and sensor coverage, providing rich transparency about the tool's behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, dense paragraph that front-loads the core purpose ('Get project-scoped 404 Demand') before detailing the clustering and filtering logic. It is not overly verbose for the complexity of the tool, though it could be broken into sentences for readability. Every sentence adds value.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the absence of an output schema, the description adequately explains what the tool returns: clusters, materialization state, sensor coverage, relevance, and how unrelated demand is handled. This is sufficient for an agent to understand the tool's output. There are no parameters to document, so nothing is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

There are zero parameters, so the schema provides no parameter information. The baseline for zero parameters is 4, and the description correctly focuses on output and behavior rather than parameter details, as none exist. This is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states a specific verb ('Get') and resource ('project-scoped 404 Demand from the current materializer window'), and differentiates itself from sibling tools by focusing on 404 demand with clustering and noise filtering. It is unambiguous and not a tautology.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides clear context about when to use this tool: it is project-scoped and tied to the current materializer window. It does not explicitly mention alternatives or exclusions, but the specificity of the description makes usage conditions obvious, so no additional guidance is needed.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

devtune_get_action_briefAInspect

Get the stored brief for a specific action, including readiness status, whether research was cut short, settled credits used, generated timestamp, and the execution-ready markdown when available.

ParametersJSON Schema
NameRequiredDescriptionDefault
actionIdYesUUID of the action to load the brief for.

TDQS

A3.6/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description alone must disclose behavior. It mentions 'when available' for the markdown, a conditional, but does not describe read-only nature, error conditions (e.g., missing action), permissions, or any side effects. The description focuses on output contents rather than behavioral traits.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, well-structured sentence. It front-loads the core purpose ('Get the stored brief') then lists the returned fields concisely. No filler or repetition.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple single-parameter retriever with no output schema, the description provides sufficient context about what the response includes. It notes the markdown is present 'when available,' covering the main conditional. It does not mention error handling or preconditions, but these are minor for a 'get' tool with no explicit constraints.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and the single parameter actionId is fully described as 'UUID of the action to load the brief for.' The description adds no extra semantic value beyond 'specific action,' which is redundant. Baseline 3 is appropriate given the schema already covers the parameter.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states a specific verb and resource: 'Get the stored brief for a specific action.' It enumerates the key fields returned, distinguishing it from other get_* tools and from the sibling generate_action_brief, which creates rather than retrieves.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies retrieval via 'stored brief' and the name 'get' versus 'generate' in the sibling list, but it does not explicitly state when to use this tool over alternatives. There is no clear 'use this when' or 'do not use if' guidance beyond what is inferred from naming.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

devtune_get_actionsBInspect

Get the actions for a project. status is suggested (DevTune proposed it, nobody has decided), accepted (the team wants it done), live (the page is published and its measurement window is open) or archived (taken off the queue by the team); blocked is true on accepted work the team has marked blocked. Suggested actions expose the same nullable opportunity percentile, percentileStatus and percentileScope as REST (signal_family compares like signals; families are interleaved), Reach, Gap, Confidence, family (including citation defence), estimated state, measured facts, and locked-page steering source. A null percentile carries percentileStatus insufficient_reach when reach is null, sensor coverage is absent, or estimated reach is zero. It carries percentileStatus no_measured_gap when every product in the family is zero and reach is non-null and not estimated-zero. Numeric percentiles omit the status. percentilePoolSize counts rankable candidates in the full family before selection or pagination. A size of one means Only ranked opportunity of this type, never Top 1%. Older records omit the pool size; do not infer singleton status from percentile 100 or the returned page. Positive-product candidates are seated before zero-product candidates. Citation-defence normalization is matched_cohort_executions when the cohort rollup is published or window_executions when the matched cohort is not published yet. Both compare citations per successful prompt execution; matched-cohort loss reports one-window-only composition changes and states matched coverage of earlier citations. Suggested actions come first and follow persisted portfolioRank ascending, nulls last, and no action exposes a legacy scalar score. Context detail includes bounded audit assessment dates, factor checks and prompt-coverage limits, plus structured Google eligibility dimensions when entitled live evidence is available. Every action carries assignedTo (its owner's user id, null when unassigned) and ownerName; change either with devtune_update_action.

ParametersJSON Schema
NameRequiredDescriptionDefault
pageNoPage number. Defaults to 1.
statusNoFilter by action status. Omit it to get every status.
channelNoFilter by recommended channel.
pageSizeNoResults per page (1-100). Defaults to 25.
priorityNoFilter by priority level.
detailLevelNoResponse detail level. summary is default. context adds bounded why-now context, available opportunity evidence, metrics, top evidence with structured Google eligibility when available, the external pages winning its prompts (competitor and third-party, in the order briefs study them), brief readiness, and follow-up links.

TDQS

B3.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does so richly: it discloses sort order (suggested actions first, portfolioRank ascending, nulls last), null-percentile/status behavior, percentilePoolSize semantics, and citation-defence normalization rules. It stops short of permissions, rate limits, or failure modes, so it is not a 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The purpose is front-loaded in the first sentence, and the dense content is arguably justified given there is no output schema to explain return fields. However, it is a run-on wall of text with no structural separation, making it hard to scan for the actionable parts.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description correctly takes on explaining returned fields, ordering, and null-handling, so it is largely complete for a six-parameter filtered-read tool. The main gaps are the lack of differentiation from sibling list tools and any mention of auth or error behavior.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so every parameter is already documented, establishing a baseline of 3. The description adds real meaning for the status enum values (suggested/accepted/live/archived) and for context-level output, but it does not clarify channel, priority, or page semantics beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The opening sentence gives a clear verb+resource ('Get the actions for a project') and the status/field discussion confirms this is a read/list tool. However, it never distinguishes itself from close siblings like devtune_get_prompt_actions or devtune_create_prompt_action, so an agent must infer which action-type this returns.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no explicit when-to-use or when-not-to-use guidance, and no named alternative for retrieving other action variants. The only routing hint is a trailing 'change either with devtune_update_action,' which addresses mutation rather than when to call this tool.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

devtune_get_adoption_metricsAInspect

Get SDK/package adoption trends including npm weekly downloads and GitHub stars over time. Tracks adoption for your primary domain and competitors.

ParametersJSON Schema
NameRequiredDescriptionDefault
metricNoFilter by specific metric type.
windowDaysNoRolling window in days: 30 or 90. Defaults to 30.

TDQS

A4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden. It does disclose the data sources and competitive scope, but it does not mention default behavior (e.g., 30-day window), output format, or any constraints such as data availability. This is adequate but not rich.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tight sentences with the core action and metrics first, and the scope in the second. No filler or repetition of parameter names.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple tool with two optional parameters and no output schema, the description covers purpose, metrics, and scope. It could mention the output shape or default window, but nothing critical prevents a correct call.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3, but the description adds real meaning by mapping the metric enum to 'npm weekly downloads' and 'GitHub stars.' It does not add detail on the windowDays parameter, but the schema already explains the allowed values and default.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific resource ('SDK/package adoption trends') and concrete metrics ('npm weekly downloads and GitHub stars'), so an agent knows exactly what the tool returns. It also states the competitive scope, which distinguishes it from other read-only analytics tools among the many devtune_get_* siblings.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The phrase 'adoption trends' implies when the tool is appropriate, and the scope ('primary domain and competitors') is useful context. However, it gives no explicit when-not-to-use guidance or alternatives, even though sibling tools like devtune_get_competitive_position and devtune_get_visibility_summary could overlap in an analytics context.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

devtune_get_agent_runAInspect

Get one managed-agent run, including its draft output only after the run is complete.

ParametersJSON Schema
NameRequiredDescriptionDefault
runIdYesManaged-agent run UUID.

TDQS

A3.7/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the burden of behavioral disclosure. It notes that draft output is available only after the run is complete, which is a meaningful behavioral constraint. However, it does not mention whether the call is read-only, what happens for in-progress runs, error behavior, or any authorization needs. The single condition provided is helpful but limited.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, well-structured sentence that front-loads the primary purpose and appends a relevant condition. It contains no filler and every phrase adds value.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple get-operation with one parameter and no output schema, the description covers the main action and a key behavioral nuance. However, it lacks details about the response structure (what fields the run includes beyond the 'draft output'), error scenarios (e.g., run not found), or the state of the run if not yet complete. This leaves some gaps for an agent calling it without prior domain knowledge.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema fully describes the only parameter (runId as a UUID with a clear description), so schema coverage is 100%. The description adds no additional parameter-specific meaning beyond what the schema already provides, so the baseline of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the operation ('Get one managed-agent run') with a specific resource and includes a scoping condition ('including its draft output only after the run is complete'). This differentiates it from devtune_list_agent_runs, which lists runs, and from devtune_run_agent, which starts them. The purpose is unambiguous and specific.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies use when a single run is needed and draft output is required only post-completion, but it does not explicitly state when to use this tool instead of alternatives like devtune_list_agent_runs or how to obtain the runId. No explicit 'use when' or 'use instead' guidance is provided, leaving selection somewhat to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

devtune_get_ai_referralsAInspect

Get complete paginated AI-referral landing pages (including engagement and page value), overview, and the complete engines breakdown with selected-provider sessions, distinct selected conversions, provider provenance, coverage gaps and native metrics for the same fixed window. Sessions are not people; conversion rates are events per session. Missing provider measurements remain null.

ParametersJSON Schema
NameRequiredDescriptionDefault
cursorNoOpaque cursor from the previous response; cursors expire after ~48 hours.
pageSizeNoResults per page (1-100). Defaults to 25.
windowDaysNoFixed rolling window in days: 7, 30, or 90. Defaults to 30.

TDQS

A3.9/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden of behavioral disclosure. It explicitly mentions pagination, notes that missing provider measurements remain null, and clarifies that sessions are not people and conversion rates are per session. These are meaningful behavioral nuances beyond a simple 'get'. It does not state read-only status, but the verb 'get' implies it, and no side effects are suggested.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, dense sentence that front-loads the primary purpose and then lists the detailed components, followed by two clarifying caveats. It is efficient and avoids redundancy, though it packs many terms into one sentence that could be slightly restructured for easier scanning. It earns a 4 for being concise and well-ordered.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the absence of an output schema, the description enumerates the key returned data elements (landing pages, engagement, page value, overview, engines breakdown, etc.), which is critical for an agent to understand what it will receive. It also explains metric interpretation and null behavior. It does not cover error handling or rate limits, but for a read-only get tool with pagination, these are minor gaps. Overall, it provides sufficient context for correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so each parameter (cursor, pageSize, windowDays) already has a clear description in the schema. The tool description adds minimal parameter-related detail, only referencing 'fixed window' and 'paginated', which are already implied by the schema. Per the rubric, a baseline of 3 is appropriate when the schema fully documents parameters.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('Get') and resource ('AI-referral landing pages') and enumerates the full set of returned data components (engagement, page value, overview, engines breakdown, provider provenance, coverage gaps, native metrics). This clearly distinguishes it from sibling tools like devtune_get_adoption_metrics or devtune_get_traffic_summary, which focus on different data domains.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides clear context about what the tool returns but does not explicitly state when to use this tool over alternatives. There are no exclusion criteria or alternative tool names mentioned. However, the detailed scope ('AI-referral') implies its use case, and the caveat about sessions vs. people clarifies interpretation. It stops short of guiding tool selection.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

devtune_get_audit_extractionAInspect

Read DevTune's extracted content for one section finding on demand. Supply the domain, page path, contentHash, inputHash and locator from audit findings. Returns exact section source, nearby Markdown, source URL and separate extraction/assessment timestamps, or unavailable/stale. This is stored crawl output; it does not reproduce an AI platform renderer or verify unseen images.

ParametersJSON Schema
NameRequiredDescriptionDefault
domainYes
locatorYes
pagePathYes
inputHashYes
contentHashYes

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations exist, so the description must carry the behavioral burden. It does: clarifies stored crawl output, that it does not reproduce a renderer, does not verify unseen images, and can return unavailable/stale. It doesn't mention mutation/side effects, but 'Read' plus stored data implies a safe lookup.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Four tight sentences cover purpose, required inputs, return values, edge cases, and limitations with no redundant filler. Easy to parse and highly actionable.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a read-only retrieval tool with no output schema or annotations, the description covers inputs, outputs, staleness, and key limitations. It doesn't detail failure response formats or prerequisites for audit findings, but the essentials are all present.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema has 0% descriptions, so the description must compensate. It names all five required parameters and tells the caller where they come from ('from audit findings'). It doesn't define each hash or the locator pattern, but the names and schema patterns are identifiable and the provenance guidance is a meaningful addition.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a precise read operation on a single artifact ('extracted content for one section finding'), names the five identifying inputs, and enumerates the exact return payload. This clearly sets it apart from broader get/list/update siblings.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly tells the caller to supply the five fields from audit findings and sets expectations for unavailable/stale results. Could name sibling tools for routing, but the provenance instruction is strong enough.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

devtune_get_audit_findingsAInspect

Get paginated audit findings and independently paginated section assessments from the published Audit rollup. Section findings include exact excerpts, remedies, versions and coverage states; they do not change numerical scores.

ParametersJSON Schema
NameRequiredDescriptionDefault
offsetNo
pageSizeNoResults per page (1-100). Defaults to 25.
includeDismissedNo

TDQS

A3.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the transparency burden and does well: it discloses pagination (including independent pagination for section assessments), the source object, the categories of data included, and explicitly states that these findings do not alter numerical scores. It does not cover auth requirements or failure modes, but for a getter this is a strong disclosure.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, front-loaded with the core purpose, with no filler. The scope and the critical non-scoring behavior are stated compactly.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The high-level purpose is clear, but with no output schema and no annotations the description should also convey return shape, pagination parameter behavior, and what includeDismissed means. It partially describes returned contents, but an agent still lacks enough detail to invoke it confidently.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is only 33% because only pageSize has a description, and the description does not compensate by explaining offset or includeDismissed. It repeats the notion of pagination but adds no semantics for the undocumented parameters beyond what the schema already provides.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific verb and resource: 'Get paginated audit findings and independently paginated section assessments from the published Audit rollup.' The detail about exact excerpts, remedies, versions, and coverage states disambiguates it from the sibling devtune_get_audit_summary and other get_* tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives context (published Audit rollup, findings vs. numerical scores) that implies when to use it, but it never explicitly states when-not-to or names an alternative like devtune_get_audit_summary. The 'do not change numerical scores' note is an indirect exclusion rather than an explicit routing instruction.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

devtune_get_audit_summaryAInspect

Get Audit summaries and independently bounded trends. trend.date is UTC for citations and answer fetches; clicksDate and priorClicksDate are Pacific GSC dates ending at the source boundary. Read current/prior coverage from searchConsole. When isCurrent is false, report partial last-published evidence with asOfDate and its original dates; do not call it current. GSC values are ranking evidence only for measured or verified_zero periods; report partial covered ranges and never pad GSC through UTC today.

ParametersJSON Schema
NameRequiredDescriptionDefault
windowDaysNoFixed rolling window in days: 7, 30, or 90. Defaults to 30.

TDQS

A3.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description must carry behavioral disclosure. It clearly explains date boundaries, UTC vs. Pacific, partial evidence, and the rule not to call partial data current, nor pad GSC values. This is a rich behavioral contract for the tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is dense but not overly long; it packs essential caveats into a few sentences. It front-loads the core purpose and then adds important accuracy constraints. No fluff, but could be slightly more structured.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the single optional parameter and no output schema, the description explains key context: date formats, timezone boundaries, partial coverage, verification states, and GSC limits. An agent has enough information to call it correctly and avoid misrepresenting data.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% for the only parameter, windowDays, which is well documented in the schema with enum values. The description does not add new meaning about the parameter beyond the schema, so baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states that the tool returns audit summaries and trends, with specific field semantics. It distinguishes it from related tools like get_audit_findings by focusing on summary-level trends. However, it does not explicitly name a sibling tool as an alternative.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage context through coverage and date handling rules, but does not explicitly state when to use this tool versus alternatives. It gives guidance on interpretation (e.g., when isCurrent false), but the when-to-use vs. when-not-to-use is absent.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

devtune_get_citation_analysisAInspect

Backward-compatible alias for top citation pages in AI search results. Prefer devtune_get_top_citation_pages for new workflows. Returns the same coverage state: an unavailable window has not been rolled up yet, so an empty list is not a measured zero.

ParametersJSON Schema
NameRequiredDescriptionDefault
cursorNoOpaque cursor from the previous response.
searchNoOptional prefix search within citation URLs and domains.
pageSizeNoResults per page (1-100). Defaults to 25.
windowDaysNoRolling window in days: 30 or 90. Defaults to 30.

TDQS

A3.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the burden of behavioral disclosure. It does add meaningful context: the tool is a backward-compatible alias, and it explains the coverage-state semantics (empty list is not a measured zero). However, it doesn't disclose pagination behavior, rate limits, or what 'citation analysis' specifically returns beyond top pages. The alias relationship is useful but the behavioral detail is thin.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is compact at three sentences and front-loads the most important fact (backward-compatible alias) before the usage directive. The coverage-state sentence is slightly dense but earns its place by explaining a non-obvious semantic. No wasted words.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple list tool with 100% schema coverage and no output schema, the description is mostly adequate. It explains the alias relationship and the coverage-state caveat, which are the non-obvious parts. However, it doesn't describe the return shape or how the alias differs behaviorally from the preferred tool, and with no annotations, an agent might not know whether this is a read-only operation (though the name suggests it is).

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents all four parameters (cursor, search, pageSize, windowDays). The description adds no parameter-specific meaning beyond the schema, but the baseline of 3 is appropriate because the schema does the heavy lifting. The description's mention of 'coverage state' relates to windowDays but doesn't add syntax or format details.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly identifies this as a backward-compatible alias for top citation pages in AI search results, with a specific verb ('get') and resource ('citation analysis' / top citation pages). It distinguishes itself from the preferred sibling devtune_get_top_citation_pages, though it doesn't fully explain what 'citation analysis' means beyond that alias relationship.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly says to prefer devtune_get_top_citation_pages for new workflows, which is a clear usage directive. It also explains the coverage-state semantics (unavailable window not rolled up yet), which helps an agent interpret results. However, it doesn't state when to use this alias over the preferred tool (e.g., legacy workflows only) or mention alternatives like devtune_get_top_citation_domains.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

devtune_get_citation_statsAInspect

Get aggregate citation statistics for a project, including total citations, unique pages, unique domains, average position, top-three count, and citation class totals. Returns a coverage state: an unavailable window has not been rolled up yet and carries no stats, which is not the same as a measured zero.

ParametersJSON Schema
NameRequiredDescriptionDefault
windowDaysNoRolling window in days: 30 or 90. Defaults to 30.

TDQS

A3.6/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure. It adds a valuable nuance about the coverage state, clarifying that an unavailable window has no stats rather than a measured zero. However, it does not mention side effects, permissions, or other behavioral traits, though for a get operation these may be assumed.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two sentences with no unnecessary words. The primary purpose is front-loaded in the first sentence, and the second sentence adds a meaningful behavioral nuance. Every sentence earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple get tool with one parameter and no output schema, the description is fairly complete. It covers the return content and a key behavioral detail. It does not explicitly mention the default window or that it operates on the current project, but these are either in the schema or implicit. Overall, adequate for the tool's complexity.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema fully describes the single parameter 'windowDays' with its enum values and default. The description adds no additional meaning beyond the schema, so the baseline of 3 is appropriate given 100% schema coverage.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the verb 'Get' and the resource 'aggregate citation statistics' with a specific list of metrics (total citations, unique pages, unique domains, average position, top-three count, citation class totals). It is specific and informative, but it does not explicitly differentiate from sibling tools like devtune_get_citation_analysis, which may overlap in purpose.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is provided on when to use this tool versus alternatives such as devtune_get_citation_analysis or devtune_get_top_citation_domains. The description only explains what the tool does, not when to choose it over siblings, leaving the agent to infer usage context.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

devtune_get_competitive_positionAInspect

Get competitive positioning data showing primary, competitor, and other citations. The default aggregates only currently active project platforms. shareOfVoice compares primary with competitor citations; primaryShareOfAllCitations also includes other citations.

ParametersJSON Schema
NameRequiredDescriptionDefault
topicIdNoFilter by topic UUID.
groupingNoTime grouping. Defaults to daily.
platformNoFilter by AI platform.
windowDaysNoRolling window in days: 30 or 90. Defaults to 30.
includeHistoricalPlatformsNoInclude a requested inactive platform’s historical data. Defaults to false.

TDQS

A3.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full disclosure burden. It goes beyond the schema by explaining the default scope (only currently active platforms) and by defining the semantics of shareOfVoice vs primaryShareOfAllCitations. This is valuable behavioral context, though it stops short of describing return format or pagination.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three concise sentences, with the core purpose front-loaded and no filler. The metric definitions are compact and directly useful.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description explains the key domain concepts well and works well with the fully documented schema. However, there is no output schema and the description does not explain the shape of the response or provide explicit differentiation from similar tools, which leaves some gaps for an agent.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the parameter descriptions already carry most of the meaning. The tool description reinforces the default behavior around active platforms, which relates to includeHistoricalPlatforms, but it does not add substantial information about the parameters beyond what the schema provides.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('competitive positioning data') and describes what is shown: primary, competitor, and other citations. It also mentions the metric names, which helps separate this from generic citation tools, though it does not explicitly distinguish it from similar siblings like devtune_get_citation_analysis.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Does not provide guidance on when to use this tool versus alternatives, nor does it mention exclusions or preconditions. The only usage-related statement is about default aggregation over active platforms, which is more of a behavioral default than guidance for tool selection.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

devtune_get_content_changesAInspect

List detected content changes, their manual claim tokens and measured-action IDs. existingMeasurementId links a page measurement completed within 14 days of the change without implying the change is claimed. Fixed reporting windows; 14-day comparisons. Path filters apply to list totals; summary covers the full window.

ParametersJSON Schema
NameRequiredDescriptionDefault
pathNo
cursorNoOpaque cursor from the previous response; cursors expire after ~48 hours.
pageSizeNoResults per page (1-100). Defaults to 25.
windowDaysNoFixed rolling window in days: 7, 30, 90, or 365. Defaults to 30.

TDQS

A3.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries full responsibility for behavioral disclosure. It reveals important semantics: existingMeasurementId does not imply a claim, and path filters affect list totals while summary covers the full window. These are non-obvious behaviors that an agent needs to interpret results correctly. However, it does not mention error handling, rate limits, or whether the operation is read-only (though 'List' implies read-only). The provided details are valuable and go beyond typical descriptions.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is three sentences with no filler. It front-loads the core purpose, then adds essential behavioral details. The structure is logical: what the tool returns, the meaning of existingMeasurementId, and scoping behavior. It is concise yet information-dense, though it could be slightly more structured with bullets, but overall it is efficient.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given there is no output schema and no annotations, the description should ideally describe the response format, error conditions, or authentication requirements. It does explain the key output elements (claim tokens, IDs) and window behavior, but it does not mention pagination behavior beyond cursor (which is in schema), error handling, or any side effects. For a low-complexity list tool, this is moderately complete but leaves some gaps an agent might need.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema descriptions cover 75% of parameters (cursor, pageSize, windowDays), leaving 'path' undocumented. The description compensates by explaining path filter behavior ('Path filters apply to list totals; summary covers the full window') and window semantics ('Fixed reporting windows; 14-day comparisons'). This adds meaning beyond the schema's enum and default descriptions, particularly for path and the interpretation of windowDays.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the verb 'List' and the resource 'detected content changes', and specifies the returned items (manual claim tokens and measured-action IDs). This makes the purpose unmistakable. It does not explicitly name a sibling alternative, but the content type (content changes with claim tokens) is distinct from other get_* tools like get_content_gaps or get_actions, so an agent can infer differentiation without explicit naming.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives operational constraints (fixed reporting windows, 14-day comparisons, path filter behavior) but does not explicitly state when to use this tool over alternatives or when not to use it. There is no mention of prerequisites or context like 'use when you need to audit content changes for a page'. The guidance is implied from the purpose but not explicit, so it earns a mid-range score.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

devtune_get_content_gapsAInspect

List measured topic, source, and brand content gaps from fixed 30/90-day rollups. Sources and brands the primary brand owns are excluded, as are topics the primary brand holds at least half the citations on, because those are presence rather than a gap. Every count is a citation split into primary and competitor; the rollups hold no mention facts. Coverage reports the covered-day count and the asOfDate the window ends on, so read the counts against availableDays rather than the requested window. Pipeline scores, generated summaries, prompt counts, and evidence are not part of this measurement contract.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNoResults per page (1-100). Defaults to 25.
offsetNoZero-based result offset. Defaults to 0.
searchNoCase-insensitive search on the display label.
gapTypeNoFilter by rollup key type.
windowDaysNoRolling window in days: 30 or 90. Defaults to 30.

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Given no annotations, the description carries full burden. It discloses the fixed window rollups, the exclusion logic, that counts are citations not mentions, and that coverage is measured against availableDays not the requested window. This prevents incorrect interpretation of results and is highly transparent about the data's limitations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is dense but organized: it starts with the main purpose, then exclusions, then data semantics, then contract scope. Every sentence adds critical information for correct usage. It is somewhat long but every part is necessary given the tool's complexity, so it earns a 4.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool is inherently complex with subtle data semantics (rollups, exclusions, citation vs mention). The description covers all the essential details an agent needs to correctly interpret results. No output schema exists, but the description explains what the data represents (citations, coverage days, asOfDate), making it complete for correct usage.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so schema already documents all parameters. The description adds context about gapType meaning but doesn't go into detail about each parameter beyond what schema provides. It does clarify that windowDays only accepts 30/90 and defaults, which is in schema. The description's contribution is minimal beyond the schema, so a 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool measures specific content gap types (topic, source, brand) from fixed 30/90-day rollups. It distinguishes itself by specifying what is excluded and that it reports citations split by primary/competitor. This clearly differentiates from siblings like devtune_get_citation_analysis and devtune_get_citation_stats.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly states what is excluded (owned sources/brands, topics with majority citations) and what is not part of the contract (pipeline scores, summaries). This guides an agent to know when to use this tool vs others in the same domain, though it doesn't name alternatives directly.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

devtune_get_dex_eventsCInspect

Read an ordered page of persisted Dex reply and lifecycle events for one project-scoped conversation thread.

ParametersJSON Schema
NameRequiredDescriptionDefault
afterNo
pageSizeNo
threadIdYes

TDQS

C2.9/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It states the operation is a read ('Read'), that events are 'persisted', and that it returns an 'ordered page'. This discloses it is a non-mutating, paginated read. However, it does not mention error behaviors, ordering specifics, or what happens for invalid thread IDs. Basic transparency is present but incomplete.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, tightly written sentence that front-loads the action and resource. There is no fluff or redundant wording. It earns its place but could have been expanded with more behavioral detail without losing conciseness.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with three parameters and no annotations or output schema, the description is too thin. It does not specify what constitutes a 'Dex reply' or 'lifecycle event', how pagination works (cursor vs offset), or what the response shape is. An agent would need to infer too much about ordering, filtering, and usage. The description is insufficient for reliable invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It mentions 'ordered page' and 'conversation thread', which hints at pageSize and threadId, but it does not explain the 'after' parameter (presumably a cursor) or explicitly map parameters to their roles. The description adds minimal value over the schema's type constraints, leaving the agent to guess parameter semantics.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action (read), the resource (persisted Dex reply and lifecycle events), and the scope (one project-scoped conversation thread). It is specific enough to distinguish from the many sibling get_* tools, though it doesn't explicitly name an alternative. The verb+resource combination is unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool versus other tools like list_dex_threads or get_agent_run. The description only explains what it does, not when it should be chosen. It implies usage for reading events for a thread but lacks any explicit exclusions or alternative routing.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

devtune_get_direct_liftAInspect

Get complete-population GA4 Direct traffic and resolver-owned AI-correlated lift classifications from fixed-window rollups, with bounded page details. Direct lift is correlated evidence, not attributed traffic, and is never added to measured referral totals.

ParametersJSON Schema
NameRequiredDescriptionDefault
pageLimitNoMaximum pages returned with their classification windows (1-25). Defaults to 10; aggregate totals are not capped.
windowDaysNoFixed rolling window in days: 7, 30, or 90. Defaults to 30.

TDQS

A4.4/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure. It explains that the data comes from 'fixed-window rollups', is 'complete-population', includes 'bounded page details', and emphasizes the non-additive nature of lift. This provides meaningful context about how the tool behaves without relying on structured metadata.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two sentences with no redundancy. The main purpose is front-loaded, and the second sentence adds an essential nuance about the nature of the lift metric. Every phrase earns its place without any filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description explains the domain-specific concepts of 'Direct lift' and clarifies its relationship to measured referrals, which is crucial for correct usage. However, since there is no output schema, it does not describe the return structure beyond 'bounded page details', which is a minor gap given the overall clarity.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% and both parameters (pageLimit, windowDays) already have descriptive text in the schema. The description adds no extra meaning to the parameters beyond what the schema provides, only indirectly referencing 'bounded page details' and 'fixed-window rollups'. This matches the baseline for high schema coverage.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb 'get' and names the exact resource: 'complete-population GA4 Direct traffic and resolver-owned AI-correlated lift classifications'. It clearly distinguishes this from siblings by explaining it is 'correlated evidence, not attributed traffic' and 'never added to measured referral totals', which sets it apart from tools like get_ai_referrals or get_traffic_summary.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides a clear when-not by stating the output is 'never added to measured referral totals', implying it should not be used for referral attribution. It also clarifies the semantic of 'lift' as correlated evidence, which guides selection. However, it does not name any specific sibling tools or give explicit conditions for when to use this over alternatives.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

devtune_get_google_eligibilityCInspect

Get Google eligibility dimensions, confirmed transitions, and sitemap submission evidence from fixed-window facts.

ParametersJSON Schema
NameRequiredDescriptionDefault
windowDaysNoFixed rolling window in days: 7, 30, or 90. Defaults to 30.
evidenceLimitNoMaximum page, transition, and sitemap evidence rows returned per list (1-25). Defaults to 10; aggregate totals are not capped.

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are supplied, so the description bears full responsibility for disclosing behavior. It only states the source and categories of retrieved data; there is no information about side effects, read-only expectation, computation involved, rate limits, or output fidelity. For a read-style tool this is minimal and leaves behavioral assumptions to the agent.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, efficient sentence that front-loads the verb and clarified resource. It is not redundant and earns its words. It loses a point because the three data categories are packed together without any formatting that could make them scan more clearly.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema and no annotations, the description must carry return-format expectations. It lists the three data categories being returned, which provides a basic notion of output content. However, it leaves terms like 'confirmed transitions' and 'sitemap submission evidence' undefined, and does not specify the structure or aggregation, so completeness is just adequate.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema already provides full coverage of the two parameters with descriptions, defaults, constraints, and enums. The description's phrase 'fixed-window facts' minimally reinforces the windowDays parameter but adds no meaningful insight beyond the schema. Baseline 3 is appropriate when schema covers 100% of parameters.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb and identifies three concrete data resources: Google eligibility dimensions, confirmed transitions, and sitemap submission evidence. It also names a source scope, fixed-window facts, which distinguishes it from sibling analytics tools. However, it does not explicitly contrast it with a sibling or scope out what it does not return.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to choose this tool over the many other devtune_get_* siblings. The description gives no selection criteria, prerequisites, or examples of contextual use. An agent is left to infer its purpose from the name alone.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

devtune_get_intelligence_forecastBInspect

Forecast short-term primary citations and presence-rate direction using recent project visibility trends. Presence-rate values are ratios from 0 to 1 and are null when no response has usable citation data.

ParametersJSON Schema
NameRequiredDescriptionDefault
windowDaysNoRolling window in days: 30 or 90. Defaults to 30.
horizonDaysNoForecast horizon in days: 7, 14, 30, 60, or 90.

TDQS

B3.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden. It discloses that presence-rate values are ratios from 0 to 1 and can be null when data is unusable, which is useful output context. However, it does not state whether this is a read-only operation, whether it requires specific permissions, or any rate limits or side effects.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two sentences, front-loaded with the core purpose and followed by a crucial data caveat. It is efficient and wastes little space.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with no annotations, no output schema, and two optional parameters, the description is adequate but not complete. It explains the forecast output format (ratios, nulls) but omits usage context, read/write nature, and how the window and horizon parameters affect results. An agent can call it, but with some uncertainty about the semantics.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3. The description adds no additional parameter meaning beyond what the schema provides – it doesn't explain how windowDays and horizonDays interact or the implications of different values.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states a specific verb (forecast) and resources (short-term primary citations and presence-rate direction) using recent project visibility trends. It distinguishes itself from siblings like devtune_get_citation_analysis or devtune_get_visibility_summary by being predictive rather than descriptive. However, it doesn't explicitly contrast itself against those siblings.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no explicit guidance on when to use this tool versus alternatives. The phrase 'short-term' is the only temporal hint, but no conditions or exclusions are provided. An agent must infer that this is for forecasting, but when to choose it over other intelligence or citation tools is unclear.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

devtune_get_knowledge_profileBInspect

Read the project Knowledge Profile and its pending, approved, and dismissed proposals.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNo
cursorNoOpaque cursor from the previous response.

TDQS

B3.4/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

There are no annotations provided, so the description must disclose behavior. It says 'Read' which implies read-only, but does not mention side effects (e.g., whether it marks proposals as viewed), nor any privacy or scoping details (e.g., whether it requires a project ID or works on the current project). No output format is described.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single sentence, concise and front-loaded with the main action and resource. It is not overly verbose, though it could list the categories more clearly.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given that the tool has no output schema and annotations are absent, the description leaves gaps: pagination (limit/cursor) is not fully explained, and the response structure is unknown. However, the tool is fairly simple, so a 3 is appropriate.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 50% (cursor is described, limit is not). The description does not explain 'limit' or 'cursor' further, but the schema already covers cursor. For 'limit', the agent has to infer it is pagination size. The description adds little beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the verb 'Read' and the resource 'Knowledge Profile', and explicitly mentions the three categories of proposals (pending, approved, dismissed). This is specific and distinguishes it from sibling tools like devtune_review_knowledge_proposal and devtune_update_knowledge_profile.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives no explicit guidance on when to use this tool versus its siblings. However, the context of 'read' and the mention of proposals implies it is for viewing state before making decisions, but this is not stated. No alternatives or exclusions are mentioned.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

devtune_get_outcomes_ledgerBInspect

Get measured actions with authoritative improved, mixed, no_clear_change, declined, still_measuring and unmeasured results, the evidence behind each result (counts, active days, exposure and the chance of a shift that large), and a conclusive-measurement improvement rate. Reporting windows filter actions; comparisons use fixed 14-day before/after windows. Also returns analytics-attributed conversions for fixed 30 and 90 day windows.

ParametersJSON Schema
NameRequiredDescriptionDefault
cursorNoOpaque cursor from the previous response; cursors expire after ~48 hours.
statusNo
pageSizeNoResults per page (1-100). Defaults to 25.
windowDaysNoFixed rolling window in days: 7, 30, 90, or 365. Defaults to 30.
interventionIdNoInclude 7, 14, 30 and 90 day impact windows (citations, presence and answer fetches) for this completion-based measurement.

TDQS

B3.2/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description must disclose behavioral traits. It describes the output content but does not state whether the operation is read-only, whether pagination is involved (though the schema mentions cursor), or any side effects. It also doesn't mention data freshness, rate limits, or whether results are computed on-demand. The description focuses on 'what' not 'how' or 'side effects'.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two sentences, front-loaded with the core purpose. It is information-dense without being verbose. The first sentence lists outcome categories and evidence, the second explains windows and conversions. It is well-structured for the amount of detail, though slightly complex due to the long list.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity and the absence of an output schema, the description provides a good overview of what is returned (result categories, evidence, improvement rate, conversions). However, it does not explain the structure of the response, pagination details (though cursor is in schema), or how the fixed windows interact with the windowDays parameter. It is adequate but leaves some gaps for an agent to infer.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 80% (4 of 5 parameters have descriptions). The description adds context about reporting windows and fixed comparison windows, which clarifies windowDays. It also mentions analytics-attributed conversions for 30/90 days, but does not map that to a specific parameter. It does not elaborate on status values or interventionId beyond the schema. Overall it adds a little but does not fully compensate for the missing 20%.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool retrieves measured actions with specific outcome categories and evidence details. It names the resource (outcomes ledger) and the exact data returned, which distinguishes it from other get_* siblings that focus on traffic, citations, or visibility. The mention of fixed comparison windows and conversions adds specificity.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explains what the tool does but provides no explicit guidance on when to use it versus alternative tools. It does not mention any prerequisites, conditions, or exclusions. An agent would not know if this is the right tool for a given query without prior knowledge of the other 50+ tools.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

devtune_get_page_guard_railsCInspect

Get paginated page guard-rail classifications from the same bounded rollup used by DevTune.

ParametersJSON Schema
NameRequiredDescriptionDefault
offsetNo
pageSizeNoResults per page (1-100). Defaults to 25.

TDQS

C2.8/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full disclosure burden. It does reveal the pagination trait and a 'bounded' data source, but omits return format, what 'bounded' concretely means, rate limits, and any consistency guarantees. For an unannotated read tool this is thin coverage.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single efficient sentence with no filler, and the key trait 'paginated' is front-loaded. It is concise rather than under-specified, though it could have used its brevity to carry more useful context.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple two-parameter read-only getter this is minimally adequate, but the description never clarifies what 'guard-rail classifications' are or what the 'bounded rollup' refers to. With no output schema and no annotations, an agent gets little beyond the fact that results are paginated.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 50% (offset is entirely undocumented, while pageSize has defaults). The description does not compensate: 'paginated' implies the offset/pageSize pair but adds no meaning beyond the schema, leaving the offset parameter's semantics unexplained.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('Get') and resource ('page guard-rail classifications') and correctly discloses pagination. However, among roughly fifty devtune_get_* siblings, it offers no differentiation — 'from the same bounded rollup used by DevTune' is vague since DevTune is the product itself, making the qualifier circular rather than distinguishing.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to select this tool versus the many sibling getters (get_page_metrics, get_content_changes, etc.). The phrase 'same bounded rollup used by DevTune' hints at a shared data source but names no alternative and gives no exclusion conditions, leaving the choice entirely to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

devtune_get_page_metricsAInspect

Get unified page-level search, citation, referral, citability, and guard-rail metrics from Audit rollups. searchConsole carries current and prior project totals, availability, covered ranges, and Pacific source dates ending at D-3. Page coverage belongs to the published Audit snapshot and may lag the project snapshot. Report isCurrent false as partial last-published evidence with asOfDate and its own source dates. Page GSC columns are a content-page subset and are ranking evidence only when availability is measured or verified_zero. Report partial coverage with its covered ranges; compare periods only when both are verified.

ParametersJSON Schema
NameRequiredDescriptionDefault
cursorNoOpaque cursor from the previous response.
sortKeyNo
categoryNo
pageSizeNoResults per page (1-100). Defaults to 25.
windowDaysNoFixed rolling window in days: 7, 30, or 90. Defaults to 30.
sortDirectionNo

TDQS

A3.6/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden. It discloses that data may be partial, that isCurrent may be false, and that coverage ranges should be reported. It explains the semantic meaning of the returned data, such as asOfDate and source dates. This is transparent about the data's nature, though it doesn't cover pagination or error behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single dense paragraph, not front-loaded with the most critical info (though it starts with the core purpose). It packs many details about data semantics and conditions, which is valuable but makes the text long and somewhat difficult to parse quickly. It could benefit from structured bullet points, but every sentence contributes meaning.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity (6 params, no output schema, no annotations), the description thoroughly explains data freshness, coverage, and validity conditions, giving the agent a clear picture of what the response represents. It hints at response fields (isCurrent, asOfDate, GSC columns). However, it lacks any guidance on pagination or sorting via cursor and sortKey, and it doesn't describe the full response structure.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema has descriptions only for pageSize and windowDays (50% coverage). The tool description adds no information about any parameter, including cursor, sortKey, category, and sortDirection, which are undocumented. With 50% coverage, the description should compensate but does not, leaving parameter semantics incomplete.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool retrieves unified page-level metrics (search, citation, referral, citability, guard-rail) from Audit rollups. This is specific and distinguishes it from other metric tools in the sibling list by focusing on page-level and Audit rollup origin, though it does not explicitly name a contrasting sibling.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides detailed guidance on data freshness (D-3 Pacific source dates), coverage caveats (published Audit snapshot may lag), and when page GSC columns are valid (only when availability is measured or verified_zero). It instructs to report partial coverage and compare periods only when verified. This gives clear usage context, though it doesn't explicitly point to alternative tools.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

devtune_get_page_trackingAInspect

Read saved published URLs, page-tracking capacity and crawl availability for a completed action. Tracking capacity does not block outcome measurement. plan_limit means upgrade or replace a tracked page; The subscription plan determines the crawl allowance. Never ask the customer to resubmit a saved URL. choices includes paused pages which can be restored. reservedPages counts selected URLs awaiting a crawl. Supply replacementUrl to read replacementImpact and disclose its affectedActions before requesting approval. Omit actionId for project-wide tracking management. allowanceStatus gives the daily status: at_limit or beyond_plan, and the AI-cited pages the allowance leaves untracked; its counts are null when not measured. Requires outcomes.read.

ParametersJSON Schema
NameRequiredDescriptionDefault
afterNo
offsetNo
searchNo
actionIdNo
replacementUrlNo

TDQS

A3.6/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does substantial work: it discloses the auth requirement ('Requires outcomes.read'), a non-obvious invariant ('Tracking capacity does not block outcome measurement'), null behavior ('its counts are null when not measured'), and an operational warning ('Never ask the customer to resubmit a saved URL'). It stops short of describing pagination or the full return shape.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The purpose is front-loaded, but the body is a stream of terse, loosely related clauses that jump from capacity semantics to field definitions to a disclosure rule without grouping. Every sentence carries information, yet the structure makes it hard to scan and the fragmentary style hurts readability.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a no-annotation, no-output-schema, 5-parameter read tool, it supplies a surprisingly complete picture: return-field meanings (plan_limit, choices, reservedPages, allowanceStatus), the permission needed, and edge-case behavior. The main gap is the undocumented pagination parameters and no explanation of paging behavior.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% across 5 parameters, so the description must compensate. It explains actionId (omit for project-wide) and replacementUrl (unlocks replacementImpact), but leaves the three pagination/search parameters (after, offset, search) entirely undocumented anywhere. It covers part of the gap, not all of it.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource: 'Read saved published URLs, page-tracking capacity and crawl availability for a completed action.' It is clearly a read of tracking state and distinct from the mutation sibling devtune_replace_tracking_page, though it bundles several resources into one grab-bag sentence rather than naming a single crisp object.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides situational cues such as 'Omit actionId for project-wide tracking management' and 'Supply replacementUrl to read replacementImpact', which imply usage contexts. However, it never states when NOT to use this tool or names alternative siblings (e.g. replace_tracking_page), so routing guidance is only implied.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

devtune_get_project_usageAInspect

Read whether the project is paused and its usage: active prompts, credits spent this credit period, AI Traffic events this UTC month, and tracked pages against the page limit. Each figure carries its own period. A paused project starts no probes, agents, crawls, audits or data syncs, drops new AI Traffic, and gives credentials read access only, so writes are refused until it is resumed in the app. Its prompts and project slot do not count against the account while paused. Credits include credits held by running work until it settles. trackedPages is null until the first page-tracking status is computed. Available to every credential.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4.4/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does so richly: a paused project starts no probes/agents/crawls/audits/syncs, drops new AI Traffic, gives read-only credentials so writes are refused, and its prompts/slot stop counting. It also explains that credits include credits held by unsettled running work and that trackedPages is null until first computed.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the read target, then the paused-project consequences and credit/null caveats. Dense but every sentence is load-bearing; slightly long, though nothing is wasted.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema exists, so the description must explain return values, and it does: it enumerates the figures, their distinct periods, the null convention, and the paused-state effects. An agent has everything needed to call and interpret it.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Zero parameters, so the baseline is 4. The description additionally clarifies field-level semantics of the returned figures (per-figure periods, null trackedPages), which adds value beyond the empty schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource: reading project pause state plus usage figures (active prompts, credits, AI Traffic events, tracked pages). No sibling tool covers project-level pause/usage, so an agent can distinguish it at a glance.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description makes clear what the tool returns but offers no explicit when-to-use condition or named alternative; usage is implied by the resource. 'Available to every credential' gives a permission note but no routing guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

devtune_get_prompt_actionsAInspect

List existing and suggested actions whose stored evidence includes a prompt. Returns action status and brief readiness.

ParametersJSON Schema
NameRequiredDescriptionDefault
promptIdYes

TDQS

A3.7/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the burden of disclosing behavior. It does say this is a listing operation and mentions the returned information, which suggests a read-only action. It does not address pagination, ordering, empty results, authorization, or whether 'existing and suggested' are returned as separate categories.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences with no filler. The main operation is front-loaded, and the return-value sentence adds useful information without redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a single-parameter read-only tool, the description covers the core purpose and a brief return summary. However, there is no output schema and no annotation context, so the vague 'brief readiness' and the lack of any mention of alternatives or parameter details leave minor but real gaps for an agent deciding whether and how to invoke it.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate for the undocumented promptId parameter. The tool name and the phrase 'includes a prompt' make the parameter's role reasonably clear, but the description never explicitly states that promptId selects the prompt whose actions will be listed, leaving some inference to the agent.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('List'), a precise resource ('existing and suggested actions whose stored evidence includes a prompt'), and the return content ('action status and brief readiness'). This filter by prompt evidence distinguishes it from siblings like devtune_get_actions, so an agent can tell what this tool is for.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The intended use is implied: call this when you have a promptId and need actions whose evidence involves that prompt. However, there is no explicit guidance about when NOT to use it or how it compares with alternatives such as devtune_get_actions or devtune_get_action_brief.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

devtune_get_prompt_set_analysisA
Read-onlyIdempotent
Inspect

Read project-scoped analysis state, assessments, proposals and credit estimate/balance. Free read; never starts analysis, settles credits or applies proposals. Supply the returned run.id on subsequent pages. Requires visibility.read. Interactive review is available in-app; Slack execution is unsupported.

ParametersJSON Schema
NameRequiredDescriptionDefault
runIdNo
offsetNo
pageSizeNo

TDQS

A4.4/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint and idempotentHint, but the description adds significant context: it never starts analysis, settles credits, or applies proposals (reinforcing side-effect-free), requires visibility.read (auth), details pagination via supplying the returned run.id, and notes Slack execution is unsupported. This goes well beyond the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is three concise sentences with no repetition. It front-loads the core purpose, then adds behavioral constraints and prerequisites. Every sentence contributes distinct information, making it efficient and well-structured.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a read-only tool with 3 optional parameters and no output schema, the description covers purpose, side effects, auth requirements, pagination, and environment limitations. An agent has enough information to call it correctly, including how to navigate subsequent pages via run.id.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate for all parameters. It only mentions run.id for pagination and does not explain offset or pageSize at all. Since these are optional and self-explanatory as pagination controls, the lack of explicit guidance is a notable gap given the low coverage.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool reads project-scoped analysis state, assessments, proposals, and credit estimate/balance. It specifies the verb 'Read' and the resource, and explicitly distinguishes it from actions like starting analysis or applying proposals, separating it from sibling tools such as devtune_start_prompt_set_analysis.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives clear context: it is a free read, never starts analysis, settles credits, or applies proposals, implying when to use it for read-only needs. It also mentions requiring visibility.read and that Slack execution is unsupported. However, it does not name specific alternative tools, only implies them, so it lacks explicit when-not alternatives.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

devtune_get_recent_answer_fetchesBInspect

Get up to 20 latest answer-fetch hourly records in the selected UTC window. Each item is one rollup record, not an individual event or a whole-window total. hour_bucket is the start of an hour. verified_event_count confirms the peer address; failed_event_count contradicts the claimed operator; unknown_no_address_count means the sensor supplied no address; unknown_no_contract_count means the operator publishes no supported identity-checking contract; unknown_check_unavailable_count means a published check could not complete. The before-verification remainder is event_count minus those five counts.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNoMaximum hourly records, 1-20. Defaults to 20.
windowDaysNoFixed rolling window in days: 7, 30, or 90. Defaults to 30.

TDQS

B3.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It extensively defines the meaning of output fields (e.g., verified_event_count, unknown_no_address_count), which adds real semantic transparency. However, it omits operational behavior such as read-only nature, required permissions, rate limits, or side effects.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is front-loaded with the core purpose and then defines each count. While dense and lengthy, every sentence defines a field that is otherwise undocumented (no output schema), so the length is justified. Structure is clear despite being a single paragraph.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given no output schema and no annotations, the description does the heavy lifting by explaining the return record structure and each count's meaning. It is largely complete for interpretation, though it could clarify the overall response format or empty-state behavior.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already fully documents both parameters (limit and windowDays). The description adds no parameter-level meaning beyond what the schema provides. Baseline 3 is appropriate when the schema does the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('Get'), resource ('answer-fetch hourly records'), and scope ('up to 20 latest ... in the selected UTC window'). It also clarifies the unit of return ('one rollup record, not an individual event or a whole-window total'), but does not explicitly differentiate itself from any sibling tool.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no explicit guidance on when to use this tool versus alternatives. The description explains what the records contain, but never states conditions, exclusions, or sibling tools to prefer. Usage is only implied by the tool name and scope.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

devtune_get_search_rollupsAInspect

Get materialized fixed-window Search Console, legacy GA4, or selected-provider analytics aggregate totals. Missing or stale freshness pointers serve the last published rollup with its own source dates. Retained evidence is partial with isCurrent false and status unmeasured; report asOfDate and say the numbers are not current. Never-published projects return unavailable coverage with null totals. GSC source dates end at Pacific today minus three days. Report that returned finality boundary. Read with per-contributor coverage metadata. Current totals require every eligible contributor to prove full-window coverage. Stale published totals remain visible without a verified current comparison. Use measurement.availability and explanation to distinguish partial evidence, unavailable verification, verified zero, and complete totals. GSC storedEvidence reports row-presence dates, including legacy and disconnected history; it does not verify continuous coverage. Always report the returned source window. The analytics source also returns paginated daily session aggregates; limit and offset apply only to those rows. Return pagination.generation on every continuation; data changes or UTC midnight require restarting at offset 0.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNo
offsetNo
sourceYesAggregate source: Search Console, legacy GA4, or the selected analytics provider.
generationNo
windowDaysNoRolling window in days: 30 or 90. Defaults to 30.

TDQS

A3.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure. It thoroughly explains fallback to last published rollup, stale data semantics, never-published null totals, GSC date boundary, coverage metadata, and pagination rules. It even details output nuances like measurement.availability and pagination.generation. This is far beyond average, though some phrasing (e.g., 'Retained evidence is partial with isCurrent false and status unmeasured') is cryptic and could be clearer, preventing a perfect score.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single dense paragraph with no section breaks, bullets, or front-loading of the most critical information. It repeats concepts (e.g., 'Always report the returned source window' and 'Report that returned finality boundary' are nearly redundant). While every sentence carries information, the lack of structure makes it harder for an agent to parse quickly. It would benefit from segmentation by topic.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given there is no output schema and no annotations, the description does an impressive job covering edge cases (stale, never-published, partial coverage), pagination continuation, date boundaries, and the meaning of availability indicators. It even warns about GSC storedEvidence not verifying continuous coverage. However, it leaves some concepts vague (e.g., 'per-contributor coverage metadata' and the generation parameter's purpose) and does not define what 'unavailable coverage' precisely implies beyond null totals. Still, for a complex tool it is remarkably complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is only 40%, so the description must compensate. It clarifies that limit and offset apply only to analytics paginated rows, which is a key semantic the schema does not convey. It also maps the source enum to human-readable names. However, it does not explain the generation parameter at all, and windowDays is left entirely to the schema (which already has a clear description). The description adds some value but leaves notable gaps, justifying a 3.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb ('Get') and a precise resource ('materialized fixed-window Search Console, legacy GA4, or selected-provider analytics aggregate totals'). It clearly distinguishes this tool from the many sibling get_* tools by naming the exact data domain and its scoped window. Even without reading the schema, an agent knows what data this returns and for which sources.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives substantial context about when the tool is applicable: it handles stale or missing freshness pointers, never-published projects, and covers all three sources. However, it never explicitly states 'use this when you need aggregate totals' versus alternatives, and it doesn't name any sibling to rule out. The guidance is implicit in the resource description and edge-case behavior, not an explicit when-to-use directive.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

devtune_get_timeline_eventsAInspect

Get project timeline events as context for interpreting Visibility trends. Sources are external, team, deterministic, or detected. Detected facts use citation_share_change and return stable aggregate evidence { platform, domain, direction, currentShare, baselineShare, absoluteChange, relativeChange, observationStart, observationEnd, projectCount, projectFraction } inside detectedEvidence; relativeChange is null when the domain has no baseline citations. Project detected evidence uses the same platform, domain, direction, shares and observation dates, plus { scope: "project", domainRole, brandName, incidentId, baselineStart, baselineEnd, cohortPromptCount, activePromptCount, status, resolvedAt }, without projectCount or projectFraction. domainRole is own, competitor or other; brandName is null for other. Project incidents must meet their detector configuration absolute-share floor, currently 0.01 (one percentage point). Shares are fractions. Status is active or resolved; resolvedAt is null while active. High-confidence active incidents and movement-ended history surface automatically, unless an overlapping pending or active citation exclusion hides them. External facts use platform_model_change, platform_outage, data_collection_gap, search_algorithm_update, or devtune_methodology_change with evidence { title, description, sourceUrl }. Prompt additions, removals, and restores return summary counts and an optional shared topic, never prompt names. Tracked-URL changes use kind tracked_urls_changed, one per day, with evidence { changeCount, brands, brandCount, requestedBy }: the number of tracked URL, brand or classification-rule changes that day and up to five affected brand names. Tracking changes apply to citations collected from that day on; earlier citations keep the labels they were recorded with. Measured-action facts use source measured_action and kinds change_detected, change_claimed, measurement_opened and outcome_measured. Their measuredAction evidence includes actionId and actionTitle when attributed, nullable interventionId and url, changeEventId for changes and results, measurementOpenedAt for openings and results, and windowClosedAt and result for results only. Results are improved, mixed, no_clear_change, declined, still_measuring or unmeasured; an outcome_measured event requires the canonical measured_at timestamp from Results. These facts ignore topic filters and have no platform keys. No causal link to metric movements is implied.

ParametersJSON Schema
NameRequiredDescriptionDefault
windowDaysNoRolling window in days: 30 or 90. Defaults to 30.

TDQS

A3.9/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden and does so richly: it discloses that high-confidence active incidents surface automatically unless hidden by an overlapping citation exclusion, that measured-action facts ignore topic filters and have no platform keys, that tracking changes apply only to citations collected from that day on, and that no causal link to metric movements is implied. These are non-obvious traits the agent could not derive from the schema.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The purpose is correctly front-loaded in sentence one, and little of the text is filler given there is no output schema. But the remainder is a single dense paragraph of backticked field lists and kind enumerations with no segmentation, making it hard to scan and locate a specific fact source.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description effectively serves as the return-value contract: it enumerates all four fact sources, the discriminated kinds, the exact evidence object shapes, status/resolution semantics, the 0.01 absolute-share floor, and edge cases such as null brandName and null relativeChange. An agent has enough to interpret the payload without further documentation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and the single windowDays parameter (30/90, default 30) is fully documented in the schema. The description adds nothing about the window's effect on results, so the baseline of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The opening sentence gives a specific verb and resource ("Get project timeline events") plus the reason it exists ("as context for interpreting Visibility trends"), so the agent knows what it returns and why. It does not, however, name a sibling it should be preferred over (e.g. get_visibility_summary, get_content_changes), so differentiation is left implicit.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is only implied through the stated purpose (interpreting Visibility trends); there is no explicit "use this when / not when" and no alternative tool is named despite ~60 siblings covering overlapping visibility, content-change and measured-action data. The agent must infer the call condition.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

devtune_get_top_citation_domainsAInspect

Get top citation domains in AI search results. Returns cursor-paginated domain rows ordered by citation count. Pass sourceVertical to scope the list to one vertical, for example the community platforms AI cites. Returns a coverage state: an unavailable window has not been rolled up yet, so an empty list is not a measured zero.

ParametersJSON Schema
NameRequiredDescriptionDefault
cursorNoOpaque cursor from the previous response.
searchNoOptional prefix search within citation URLs and domains.
pageSizeNoResults per page (1-100). Defaults to 25.
windowDaysNoRolling window in days: 30 or 90. Defaults to 30.
sourceVerticalNoOptional source vertical scope. "community" covers Reddit, Hacker News, GitHub, Stack Overflow, Stack Exchange and Dev.to, including their subdomains.

TDQS

A3.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden of behavioral disclosure. It discloses cursor-based pagination, ordering by citation count, and crucially explains the coverage state (unavailable window not rolled up, empty list not a measured zero). This is valuable interpretive context beyond the basic read operation. It does not detail the exact response fields, but the provided semantics are sufficient for safe use.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences with no filler. The purpose is front-loaded, and the coverage-state caveat is placed at the end as an important interpretation note. Every sentence earns its place and the structure is clear and efficient.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description should fully explain the response shape. It mentions cursor-paginated domain rows and a coverage state, but does not specify the fields within each row or the exact structure of the coverage state. This leaves some ambiguity about how to consume the output. Given the tool's moderate complexity (5 optional params), a bit more detail would improve completeness.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3. The description adds minor value by giving an example of sourceVertical ('community' covers specific platforms), but this is already implied in the schema's enum description. No new meaning is added for cursor, pageSize, or windowDays beyond what the schema already states.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource: 'Get top citation domains in AI search results.' It clarifies that it returns domain rows, not pages, distinguishing it from sibling devtune_get_top_citation_pages. The ordering by citation count is also explicit, making the purpose unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives context on how to use parameters (e.g., pass sourceVertical to scope) but does not explicitly state when to use this tool versus alternatives like devtune_get_top_citation_pages or devtune_get_citation_analysis. No exclusions or alternative routing are mentioned, so the agent must infer usage from the name and description.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

devtune_get_top_citation_pagesAInspect

Get top citation pages in AI search results. Returns cursor-paginated page rows ordered by citation count. Pass sourceVertical to scope the list to one vertical, for example the community platforms AI cites. Returns a coverage state: an unavailable window has not been rolled up yet, so an empty list is not a measured zero.

ParametersJSON Schema
NameRequiredDescriptionDefault
cursorNoOpaque cursor from the previous response.
searchNoOptional prefix search within citation URLs and domains.
pageSizeNoResults per page (1-100). Defaults to 25.
windowDaysNoRolling window in days: 30 or 90. Defaults to 30.
sourceVerticalNoOptional source vertical scope. "community" covers Reddit, Hacker News, GitHub, Stack Overflow, Stack Exchange and Dev.to, including their subdomains.

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries full responsibility. It discloses cursor-paginated rows, ordering by citation count, and the important nuance that an empty list may reflect an unavailable window rather than a measured zero. This goes beyond a basic description and prevents a common misunderstanding.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences, each earning its place: purpose, pagination/ordering, sourceVertical usage, and coverage state. No fluff, front-loaded with the core purpose.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a read-only list tool with 5 optional parameters and no output schema, it explains pagination, filtering, window selection, and a key semantic nuance (coverage state). It doesn't describe the exact structure of returned page rows, but for a cursor-paginated list that is often sufficient; the main purpose and behavior are well covered.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema already describes all parameters (100% coverage), so the baseline is 3. The description adds meaningful semantics for sourceVertical (explains it scopes to a vertical with an example) and clarifies the cursor-pagination model. It doesn't add much for other parameters, but the added context justifies a 4.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: 'Get top citation pages in AI search results.' It clearly distinguishes this from sibling tools like devtune_get_top_citation_domains by specifying 'pages', and further describes pagination and ordering criteria.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It provides specific usage guidance for sourceVertical with an example ('for example the community platforms AI cites') and explains the coverage state to prevent misinterpretation of empty results. However, it does not explicitly mention alternative tools or conditions for when not to use it, which is a minor gap.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

devtune_get_traffic_platformsAInspect

Get LLM traffic grouped by AI platform, traffic type and bot class over a fixed rolling window, so a platform's training crawls and answer fetches are separate rows. verifiedEvents confirms the peer address; failedEvents contradicts the claimed operator; unknownNoAddressEvents means the sensor supplied no address; unknownNoContractEvents means the operator publishes no supported identity-checking contract; unknownCheckUnavailableEvents means a published check could not complete. The before-verification remainder is totalEvents minus those five counts.

ParametersJSON Schema
NameRequiredDescriptionDefault
domainNoFilter by specific domain.
windowDaysNoRolling window in days: 30 or 90. Defaults to 30.

TDQS

A3.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden, and it discharges most of it: it defines the row semantics and precisely unpacks five event counters, including the arithmetic relationship (totalEvents minus the five counts = before-verification remainder). It does not cover permissions, rate limits, or ordering, but for a read-only analytics tool the metric semantics are the behavior that matters and they are disclosed well.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the purpose, followed by dense but necessary counter definitions. Every sentence adds distinct meaning and there is no filler. The counter glossary is long but justified given there is no output schema to carry those semantics.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description is the only source of return semantics, and it does explain the row grouping and every counter plus the derived remainder. It still omits the possible values of platform/traffic type/bot class and any sorting or pagination behavior, which a fully complete definition would include.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so both parameters (domain, windowDays) are already documented with their enum and default in the schema. The description's phrase 'over a fixed rolling window' restates windowDays without adding enumeration values, format, or edge cases. Baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (Get) and resource (LLM traffic) plus the grouping dimensions (AI platform, traffic type, bot class) and the fixed rolling window. An agent can tell it produces grouped bot-classified rows. However it never distinguishes itself from closely-named siblings like devtune_get_traffic_summary or devtune_get_ai_referrals, so routing is left to inference.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is an implicit hint about why grouping matters (training crawls and answer fetches as separate rows), but no explicit when-to-use, when-not-to-use, or which sibling to pick instead. With ~40 sibling analytics tools, the absence of any routing guidance is a real gap.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

devtune_get_traffic_summaryAInspect

Get current website traffic totals, LLM bot and AI-referral subsets, and configuration-backed accepted-measurement Active state for snippet, edge middleware, Cloudflare, Search Console, GA4, and PostHog sensors. Sensor state is project-wide even when domain filters the traffic totals.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A3.9/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are present, so the description carries the behavioral burden, and it does disclose one non-obvious behavior: sensor Active state is project-wide even when domain filtering affects traffic totals. It stops short of explaining 'accepted-measurement Active' semantics or any freshness guarantees, but for a zero-argument get operation the core behavioral profile is reasonably clear.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two focused sentences with the main resource front-loaded in the first sentence and an important scoping caveat in the second. The phrase 'configuration-backed accepted-measurement Active state' is dense, but no sentence is wasted.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description usefully enumerates the main return components: totals, bot/AI-referral subsets, and sensor Active state. It also clarifies the project-wide scope of sensor state. It omits output format and time-window details, but for a parameterless read-only tool it is largely sufficient.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema has zero parameters, so there is no parameter-documentation burden and the baseline is 4. The description alludes to domain filtering, but that filtering is not exposed as a parameter and therefore does not create a schema/description mismatch.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific verb ('Get with a current resource') and spells out exactly what is returned: website traffic totals, LLM bot and AI-referral subsets, and per-sensor Active state for six named measurement sources. This level of detail distinguishes it from sibling tools like devtune_get_traffic_platforms or devtune_get_ai_referrals despite not explicitly naming those alternatives.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance, no stated alternatives, and no exclusions comparing this tool to its many siblings. The only contextual note is about scope ('Sensor state is project-wide even when domain filters the traffic totals'), but it does not help an agent decide which traffic-related tool to select.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

devtune_get_visibility_responseAInspect

Read one retained AI response using an executionId from devtune_list_visibility_responses, with structured sponsored ads labelled separately from the organic answer. Returns null if the execution is outside this project. retained=false means the raw response has expired, not that no ads were displayed. Use visibility summary for rollup-backed advertising rates.

ParametersJSON Schema
NameRequiredDescriptionDefault
executionIdYes

TDQS

A5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden and handles it well: it explains the null return for out-of-project executions, clarifies that retained=false means expiration rather than absence of ads, and reveals that sponsored ads are labelled separately from the organic answer.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences, no filler. The core action is front-loaded, followed immediately by the two most decision-relevant caveats and the routing hint. Every sentence earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a single-parameter read tool with no output schema, the description is complete: it states what is returned, how results are structured, what null means, what retained=false means, and which sibling to use for aggregate metrics. An agent has enough context to call it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It does by telling the agent exactly where to obtain executionId ('from devtune_list_visibility_responses'), which adds provenance beyond the bare UUID schema field.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource: 'Read one retained AI response using an executionId'. It clearly distinguishes the tool from devtune_list_visibility_responses (bulk listing) and devtune_get_visibility_summary (rollup rates), so an agent can identify the correct tool without inspecting schemas.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It explicitly states the executionId comes from devtune_list_visibility_responses, tells the agent when to use this tool versus the summary alternative ('Use visibility summary for rollup-backed advertising rates'), and includes a null-return caveat for executions outside the project.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

devtune_get_visibility_summaryAInspect

Get AI search visibility summary metrics for a project. The default aggregates only the project’s currently active platforms and returns measurement metadata; use only measurement.reportablePlatforms for current per-platform reporting. This excludes configured onboarding platforms that have no measurements in the window. measurement.dataStatus=no_data means the platform was not measured in the window, not that visibility was zero. Inactive historical data is excluded unless explicitly requested. Organic percentage fields use 0 to 100; citationPresence and shareOfVoice are null when no response has usable citation data. advertising reports separate forward-only ads detected, measuredResponses coverage, and primaryRate/competitorRate fractions from 0 to 1, null when unmeasured. Never combine advertising with organic Presence.

ParametersJSON Schema
NameRequiredDescriptionDefault
topicIdNoFilter by topic UUID.
platformNoFilter by AI platform (e.g. chatgpt, perplexity, gemini).
windowDaysNoRolling window in days: 30 or 90. Defaults to 30.
includeHistoricalPlatformsNoInclude data from a requested inactive platform. Defaults to false and should only be enabled when the user explicitly asks for historical platform data.

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description fully discloses behavior: dataStatus=no_data means unmeasured not zero, null semantics for citationPresence/shareOfVoice, exclusion of inactive historical data, and the advertising vs organic distinction. This is comprehensive behavioral transparency.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is dense but every sentence conveys essential information for correct usage and interpretation. It is front-loaded with the main purpose and then details key constraints, though it could be broken into clearer sections for readability.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the complexity of the tool and no output schema, the description covers critical edge cases: data status, null values, percentage ranges, historical data handling, and the advertising/organic separation. It lacks explicit mention of the return object structure, but provides sufficient interpretive guidance for an agent.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so parameters are already documented. The description adds value by clarifying windowDays default, includeHistoricalPlatforms default and when to enable it, and the meaning of platform filter values. This goes beyond the schema's basic descriptions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool retrieves AI search visibility summary metrics for a project, and adds specifics like active platforms and measurement metadata. It differentiates itself from siblings like get_visibility_response by focusing on summary-level metrics and field semantics.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It provides explicit guidance on when to use measurement.reportablePlatforms, warns against combining advertising with organic Presence, and explains the includeHistoricalPlatforms parameter usage. This gives clear direction on correct invocation and interpretation.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

devtune_get_what_worksAInspect

Get project channel evidence from canonical closed 14-day improvements. Samples count improvements, not all attempts; before/after evidence is not causal attribution.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNoMaximum recommendations to return. Defaults to 5.
windowDaysNoRolling window in days: 30 or 90. Defaults to 30.

TDQS

A3.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure. It explicitly discloses that results are a sample, not all attempts, and that the before/after evidence is not causal attribution. This is valuable transparency about the nature of the data and its limitations, going beyond what a schema could convey.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two sentences with zero redundancy. The main purpose is front-loaded, and the caveat follows succinctly. Every word earns its place, making it highly efficient for an agent to parse quickly.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description is adequate for a simple two-parameter query tool. It explains the purpose and key limitations but does not describe the return format or how the results are structured. Since there is no output schema, a brief mention of the output would improve completeness, but for a tool with this simplicity, it is minimally sufficient.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema covers both parameters (limit and windowDays) with clear descriptions and defaults. The description adds no additional parameter semantics, but since schema coverage is 100%, it does not need to. The baseline of 3 is appropriate because the schema already provides the necessary meaning.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the verb (get) and the resource (project channel evidence from canonical closed 14-day improvements), which is specific and unambiguous. It does not explicitly differentiate from sibling tools, but the focus on 'improvements' and 'what works' distinguishes it enough. The caveat about sampling adds clarity but the purpose is clear without it.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool versus alternatives. It does not mention any exclusions, prerequisites, or conditions that would help an agent decide between this and the many other get_* tools in the sibling list. The description simply states what it does without any 'use this when...' context.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

devtune_list_agent_runsBInspect

List project managed-agent runs with status, timing, settled credit consumption, token usage, and agent identity.

ParametersJSON Schema
NameRequiredDescriptionDefault
cursorNoOpaque cursor from the previous response.
agentKeyNoFilter to one roster or custom agent key.
pageSizeNoResults per page (1-100). Defaults to 25.

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the burden. It discloses the returned fields and implies pagination via cursor/pageSize, but it does not state whether results are ordered, whether agentKey is an exact match or partial, or whether this is a read-only operation. The description is accurate but not deeply transparent about behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single sentence that front-loads the main action and resource, then lists the key returned fields. It is concise and structured well, though it could be slightly more explicit about pagination behavior.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a list tool with no output schema and no annotations, the description covers the main purpose and returned fields, but it lacks details on ordering, default behavior, and how to use the cursor. It is adequate for basic invocation but leaves some gaps for an agent needing to paginate or filter effectively.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents all three parameters. The description adds context about what fields are returned but does not add meaning beyond the schema for the parameters themselves. Baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('List') and resource ('project-managed agent runs') and enumerates the fields returned (status, timing, credit consumption, token usage, agent identity). It is clear what the tool does, though it does not explicitly distinguish it from the sibling devtune_get_agent_run, which likely retrieves a single run.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies a listing use case and the parameters (cursor, agentKey, pageSize) suggest pagination and filtering, but it does not explicitly state when to use this tool versus devtune_get_agent_run or other list tools. No exclusions or alternative routing are provided.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

devtune_list_agentsAInspect

List the code-owned and custom managed agents available to a project using display-only definition fields. Includes the one-time account trial credit allowance and scheduling restriction when trialing. Each agent carries an evidence verdict naming any source it cannot read yet, so a caller can see what a run would be missing before paying for one.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden, and it does so well: it states the operation is display-only, surfaces trial-related constraints, and reveals the returned evidence-verdict field. It does not explicitly say there are no side effects or cover auth/rate limits, but the display-only framing is a strong implicit safety signal.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences, front-loaded with the core action, then adding the two distinct behavioral details (trial constraints and evidence verdicts). No filler or repetition of the tool name.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-parameter, no-output-schema listing tool, the description covers the action, the scope, the information returned per agent, and the caller's motivation (see what a run would miss before paying). Nothing material is missing for a correct call.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

There are zero parameters, so there is no schema to supplement. The description instead clarifies what the parameterless call returns (agent list plus trial credit, scheduling restriction, and evidence verdicts), which is the only semantic support an agent needs.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb ('List') and a concrete resource ('code-owned and custom managed agents available to a project'), and clarifies it uses display-only definition fields. This makes it easy to distinguish from siblings like devtune_list_agent_runs (runs, not agents) and devtune_run_agent (execution, not listing).

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides clear context for when this is useful: before paying for a run, to inspect trial credit, scheduling restrictions, and evidence of inaccessible sources. It doesn't explicitly name alternatives or state when not to use it, but the intended pre-run inspection use case is evident.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

devtune_list_dex_handoffsAInspect

List pending specialist-agent handoffs for one project-scoped Dex conversation thread.

ParametersJSON Schema
NameRequiredDescriptionDefault
threadIdYes

TDQS

A3.8/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. 'List' suggests a read-only operation, but this is not explicitly stated, and there is no disclosure of side effects, return format, or any constraints beyond 'project-scoped.' The term 'pending' is undefined, leaving ambiguity about what constitutes pending.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single sentence that is concise and front-loaded with the key action and scope. No filler or redundant wording; every part earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple tool with one parameter and no output schema, the description is largely sufficient. It states the purpose and scope. However, it could be improved by clarifying what 'pending' means and what the response contains (e.g., a list of handoff objects). Given the tool's low complexity, this is a minor gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It mentions the thread scope ('project-scoped Dex conversation thread'), which gives context for threadId, but it does not explain the parameter's format, constraints, or how to obtain it. The description adds minimal meaning beyond the parameter name itself.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb 'List' with a clear resource 'pending specialist-agent handoffs' and scopes it to 'one project-scoped Dex conversation thread.' This distinguishes it from sibling list tools like devtune_list_agent_runs or devtune_list_library, which target different resources.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies the tool is used when you need to see pending handoffs for a specific thread, but it does not explicitly contrast with alternatives like devtune_decide_dex_handoff (which likely creates handoffs) or devtune_get_dex_events (which might include handoff events). No when-not-to-use guidance is provided.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

devtune_list_dex_threadsAInspect

List the bounded recent Dex conversation history for the project, including current task, settled credits, and truncation state.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden. It adds useful behavioral context by noting the history is 'bounded' and includes 'truncation state,' which informs the agent about potential limits and pagination. However, it does not explicitly state that the operation is read-only or mention any side effects, authentication, or rate limits. It is not contradictory but leaves some behavioral ambiguity.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, dense sentence that gets straight to the point. Every phrase adds value: 'bounded recent' sets the scope, 'for the project' identifies the context, and the listed fields tell the agent what to expect. There is no filler or redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a read-only list tool with no parameters and no output schema, the description covers the essential information: what it returns and its scope. It could be slightly richer by explicitly stating it is a read-only operation or what 'bounded' means in practice (e.g., number of threads), but it is otherwise sufficient for an agent to invoke it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

There are zero parameters and the schema reflects that, so no parameter documentation is needed. The baseline for 0 params is 4, and the description adds no parameter-specific meaning (there are none to add). It correctly implies no inputs are required for the call.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('list') and a clear resource ('bounded recent Dex conversation history'), and it names the included fields (current task, settled credits, truncation state). This clearly distinguishes it from sibling tools like devtune_get_dex_events, which focus on events rather than conversation history, though it does not explicitly name a sibling to avoid.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is provided on when to use this tool versus alternatives. It simply describes what it does without indicating scenarios (e.g., 'use this to inspect the current task or check truncation') or excluding cases where another tool would be more appropriate. The agent is left to infer usage from the name and description alone.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

devtune_list_libraryCInspect

List project Library items. Current versions are returned by default.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNo
cursorNoOpaque cursor from the previous response.
includeSupersededNo

TDQS

C2.8/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden of behavioral disclosure. It does add one useful behavior—'Current versions are returned by default'—but it omits other relevant behavior such as pagination behavior, ordering, whether superseded items are included by default, and any permissions or side effects.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two short sentences with no filler. It front-loads the core purpose and adds a meaningful default-behavior note, making every word count.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with three optional parameters and no output schema, the description is thin. It does not explain pagination flow, what fields are returned, how to handle the cursor, or how listing differs from searching the library. An agent would need to infer too much from parameter names and sibling names.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is only 33%, with only 'cursor' described in the schema. The description partially compensates by implying that includeSuperseded defaults to false ('Current versions are returned by default'), but it does not explain the 'limit' parameter or its maximum, nor does it clarify pagination semantics beyond the cursor's schema note.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb and resource: 'List project Library items.' It clearly identifies the object and action, and the note about current versions adds useful scope. However, it does not explicitly distinguish itself from sibling tools like devtune_search_library or devtune_read_library_item.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is provided about when to use this tool versus alternatives such as devtune_search_library. The description only states what the tool does, not the conditions under which it should be preferred over other library-related tools.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

devtune_list_visibility_responsesAInspect

List the latest 20 successful AI executions in this project, optionally filtered by prompt or platform. Pass an executionId to devtune_get_visibility_response to inspect retained answer and ad evidence. This is an operational listing, not a measurement sample; use visibility summary for rates.

ParametersJSON Schema
NameRequiredDescriptionDefault
platformNo
promptIdNo

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are present, so the description carries the behavioral burden. It discloses the 20-item limit, the 'successful' filter, project scoping, and the operational vs. measurement distinction. It does not explicitly state whether the call is read-only or describe pagination, but for a listing tool the key behavioral traits are covered.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences, each earning its place: the core operation, the follow-up path, and the crucial distinction from measurement tools. The most important information is front-loaded with no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given only two optional simple parameters and no output schema, the description covers the essential context: what is listed, filter options, the 20-item limit, and how to proceed with a returned executionId. It does not describe the exact output fields, but the pointer to executionId gives enough operational grounding for correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It adds that platform and promptId are optional filters and explains their role ('optionally filtered by prompt or platform'). However, it doesn't precisely map 'prompt' to the promptId parameter or clarify whether filters combine, leaving some semantic ambiguity.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific action: 'List the latest 20 successful AI executions in this project', naming the resource, scope, count, and success filter. It also distinguishes itself from related tools by explicitly contrasting with devtune_get_visibility_response and visibility summary, so an agent can tell it apart from siblings without opening schemas.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives explicit routing guidance: pass an executionId to devtune_get_visibility_response for deeper inspection, and use visibility summary for rates. The line 'This is an operational listing, not a measurement sample' clearly states when this tool is appropriate and when it is not.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

devtune_read_library_itemAInspect

Read one Library item, including its content, provenance, version state, and grounding availability.

ParametersJSON Schema
NameRequiredDescriptionDefault
itemIdYes

TDQS

A3.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full disclosure burden. It states it 'reads' and lists what is included, implying a safe read operation, but it does not mention any prerequisites, error behavior, or side effects. For a read tool, the description is adequate but not rich in behavioral context.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, front-loaded sentence with zero filler. It efficiently conveys the action and the key aspects of the returned data, making it easy to scan.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple read tool with one parameter and no output schema or annotations, the description is largely complete. It explains what the tool returns and implies the required input. The only minor gap is not stating behavior for missing items, but this is not critical for a read operation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema has one required parameter, itemId, which is a UUID. The description does not mention this parameter, and schema description coverage is 0%, yet the parameter's purpose is self-evident as an identifier. The description adds no extra meaning, but the low complexity and intuitive nature of 'itemId' make this acceptable.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the verb 'Read' and the resource 'Library item', and enumerates the content returned (content, provenance, version state, grounding availability). This distinguishes it from list/search tools implicitly through the singular 'one' and 'read' action, but it does not name a specific sibling alternative.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage when you have an itemId and need full details on a single library item, but it does not explicitly state when to use this instead of list/search or mention any exclusions. The context is inferable from 'read one' but lacks explicit guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

devtune_reconsider_retrieval_restrictionA
Idempotent
Inspect

Reconsider an intentional retrieval restriction for only the exact URLs displayed in an audit finding. Requires explicit customer confirmation and actions.write. Preserves the finding, evidence and score. Never infer customer intent.

ParametersJSON Schema
NameRequiredDescriptionDefault
urlsYes
findingIdYes
auditRunIdYes
confirmedIntentYes

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond the annotations (readOnlyHint=false, idempotentHint=true, destructiveHint=false), the description adds important behavioral context: it requires explicit confirmation and the actions.write permission, preserves the finding/evidence/score, and prohibits inferring customer intent. This meaningfully helps the agent understand safety and side effects without contradicting the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is three compact sentences with zero filler. The core action is front-loaded, followed by permission/confirmation requirements, then a preservation guarantee and a policy constraint. Every sentence earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the four required parameters, the annotations, and the absence of an output schema, the description covers the essential prerequisites: exact URL scope, customer confirmation, required permission, and preservation of audit artifacts. It does not describe the exact outcome or return value of 'reconsider', but the action and constraints are sufficiently specified for correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema provides 0% description coverage, so the description must carry parameter meaning. It clarifies that 'urls' must be the exact URLs from the audit finding and that 'confirmedIntent' corresponds to explicit customer confirmation. The remaining parameters (auditRunId, findingId) are reasonably self-explanatory from their names and the phrase 'audit finding', though they are not explicitly mapped.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('Reconsider') and a specific resource ('an intentional retrieval restriction') with a clear scope ('only the exact URLs displayed in an audit finding'). This distinguishes it from the sibling 'acknowledge_retrieval_restriction' tool and makes the tool's purpose immediately actionable.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives clear context for when to use the tool: only for exact URLs in an audit finding, only with explicit customer confirmation, and never by inferring intent. It does not explicitly name sibling alternatives or state when-not-to-use cases, but the guidance is strong enough to route an agent correctly.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

devtune_rename_dex_threadCInspect

Rename one project-scoped Dex conversation thread.

ParametersJSON Schema
NameRequiredDescriptionDefault
titleYes
threadIdYes

TDQS

C2.7/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description must disclose behavioral traits. It identifies 'rename' as a mutation but does not mention side effects, permission requirements, reversibility, or what the tool returns. An agent has no idea what happens after the rename or whether it is destructive.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single efficient sentence with no wasted words. It is front-loaded with the action and resource. However, the brevity sacrifices essential detail, so while structure is clean, it is not fully adequate.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a mutation tool with two required parameters, no annotations, and no output schema, the description is severely incomplete. It does not explain parameter semantics, expected behavior, error conditions, or return value, leaving an agent to rely entirely on the schema and guess at side effects.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, and the description does not mention either parameter (threadId, title). The agent must inspect the schema to understand that threadId identifies the thread and title is the new name, and the maxLength constraint is not surfaced anywhere.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb 'rename' and a clear resource 'project-scoped Dex conversation thread', which distinguishes it from sibling operations like delete or list. It is concise and unambiguous about what the tool does.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool versus alternatives such as devtune_delete_dex_thread or devtune_list_dex_threads. No context is provided about prerequisites, typical scenarios, or conditions that would make rename the right choice.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

devtune_replace_tracking_pageAInspect

Pause replaceUrl to make room for url, which must be a saved published URL or a previously paused page. Obtain explicit human approval for the selected page first. Preserves history and the completed action. First call devtune_get_page_tracking with replacementUrl and disclose replacementImpact. Pass its token as impactToken after human approval. Omit replaceUrl to restore a paused page into an available slot; otherwise reverse by swapping the URLs. This controls audits and content-change detection; outcome measurement uses the action completion date. Use operation=pause to stop tracking url and release its slot without a replacement; obtain the disclosure for that URL first. Omit actionId to manage project tracking; restoring without an action requires a previously paused URL. Requires actions.write.

ParametersJSON Schema
NameRequiredDescriptionDefault
urlYes
actionIdNo
operationNo
replaceUrlNo
impactTokenNo

TDQS

A4.4/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full behavioral burden, and it discloses important side effects: it requires explicit human approval, preserves history and the completed action, releases slots, and requires actions.write. It also explains downstream effects on audits and content-change detection, which goes well beyond a simple mutation statement.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The text is long and dense, but virtually every sentence adds a distinct rule, prerequisite, or mode. It would be more scannable with bullets or clearer separation of the replace, restore, and pause flows, so it is not maximally structured, though it avoids filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a five-parameter mutation tool with no annotations and no output schema, the description covers prerequisites, permissions, parameter roles, and workflow sequencing well. It does not state expected return values or error behavior, which an agent may need when interpreting the result of the call.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It meaningfully explains `url`, `replaceUrl`, `impactToken`, and `operation=pause`, but `actionId` is only obliquely addressed via 'Omit actionId to manage project tracking'. The default or `track` operation mode is not explicitly described, leaving some ambiguity.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description conveys a specific mutation operation: pausing or replacing the currently tracked URL (`replaceUrl`) with `url`, and also supports restore and pause modes. It is distinct from the many read-only `devtune_get_*` siblings, though the opening phrase 'Pause replaceUrl to make room for url' is awkward and lacks a plain one-line statement of the tool's core purpose.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives explicit sequencing and conditions: 'First call devtune_get_page_tracking with replacementUrl', 'Pass its token as impactToken after human approval', 'Omit replaceUrl to restore a paused page', and 'Use operation=pause to stop tracking url'. It also states the required permission and clarifies when actionId should be omitted, making when-to-use guidance strong.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

devtune_review_knowledge_proposalBInspect

Approve, dismiss, or restore a Knowledge Profile proposal. Approval is the only transition that changes the profile.

ParametersJSON Schema
NameRequiredDescriptionDefault
decisionYes
proposalIdYes

TDQS

B3.3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the transparency burden almost entirely. It does disclose a meaningful side effect: 'Approval is the only transition that changes the profile.' However, it omits auth requirements, reversibility, state constraints, or consequences of dismiss/restore beyond not changing the profile.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two concise sentences, front-loaded with the action and resource. The second sentence adds critical behavioral differentiation without bloat. No redundant or filler content.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool is simple with two params and no output schemalexplicit statement about what happens on dismiss/restore or what the response contains, but the core invocation is understandable.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema gives enum values for decision, and the description adds the important semantic that only approve changes the profile. However, proposalId receives no additional explanation, though its meaning is inferable from the tool purpose.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description has a specific verb phrase, 'Approve, dismiss, or restore a Knowledge Profile proposal,' which clearly identifies the action and target. It does not explicitly distinguish itself from sibling tools like devtune_update_knowledge_profile, but the resource (proposal vs profile) makes the intent fairly clear.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is provided about when to use this tool versus approving through another workflow or updating the profile directly. The usage context is only implied by the action verbs; there are no explicit when-to-use or when-not-to-use conditions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

devtune_review_prompt_set_proposalA
Destructive
Inspect

Apply, undo, dismiss or restore one analyzer proposal after explicit human approval of that operation. Read the proposal first and show the human the proposed change. Never infer approval from a read or analysis request. Preserves quotas and stale/conflict checks. Requires actions.write and a user-attributed credential with devtune.manage. Interactive review remains in-app; Slack execution is unsupported.

ParametersJSON Schema
NameRequiredDescriptionDefault
operationYes
proposalIdYes
humanApprovedYes

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond the destructiveHint annotation, the description discloses permission requirements (actions.write, devtune.manage credential), quota preservation, stale/conflict checks, and a platform limitation. It also explains the non-obvious rule that a read/analysis request must never be taken as approval. No contradictions with annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The most important information (operations + human approval gate) is front-loaded, and every subsequent sentence adds distinct value: permissions, stale/conflict behavior, platform restriction. Nothing is redundant or filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a destructive, non-idempotent mutation with no output schema, the description covers the approval workflow, required credentials, side-effect protections, and unsupported invocation context. An agent has enough to decide when and how to call it safely.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description carries the burden. It maps the operation enum by listing the four values, identifies proposalId as 'one analyzer proposal', and conveys that humanApproved must be an explicit true approval. It does not elaborate the precise effect of each operation, but the names plus the approval context make the semantics clear.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with four specific operations ('Apply, undo, dismiss or restore') on a concrete resource ('one analyzer proposal'), immediately distinguishing this from the sibling devtune_review_knowledge_proposal by resource type and from read-only analysis tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It clearly states the required precondition ('after explicit human approval'), instructs the agent to read the proposal first and show the change, and warns against inferring approval from reads/analysis. It also names an environment exclusion (Slack unsupported). It does not explicitly name sibling alternatives, but the usage context is unambiguous.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

devtune_run_agentAInspect

Trigger a metered manual managed-agent run through DevTune’s evidence- and credit-checked dispatch path. Refuses with agent_evidence_unavailable, before any credits are spent, when the project has nothing this agent can read.

ParametersJSON Schema
NameRequiredDescriptionDefault
modelIdNoOptional available model-catalog identifier.
agentKeyYesRoster or custom agent key to run.
instructionsNoOptional focus for this run.
requestIdempotencyKeyYesRequired caller-generated UUID for idempotent retries.

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the behavioral burden, and it does disclose concrete behavior: the run is metered, goes through evidence/credit checks, and refuses with agent_evidence_unavailable before spending credits when no readable evidence exists. It stops short of describing success-side effects such as persistence, async completion, or what a returned run object contains.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, no filler: the first states the action and constraints, and the second states the key failure mode. The important refusal/cost behavior is front-loaded after the core trigger.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple 4-parameter, no-output-schema tool, the description supplies the essential guardrail and cost behavior an agent needs before invoking. It does not mention how to obtain/verify the resulting run, but sibling tools like get_agent_run and list_agent_runs fill that gap implicitly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the schema already documents agentKey, requestIdempotencyKey, modelId, and instructions. The description adds context about evidence availability but does not deepen the meaning of any individual parameter, fitting the baseline of 3.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a precise verb-resource pair—'Trigger a metered manual managed-agent run'—and frames it as the dispatch action, which separates it from the many get/list siblings. The refusal condition adds a specific behavioral signature rather than a generic 'run an agent' phrase.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It is clearly about initiating a run, so an agent can infer when to call it, but it never explicitly says when to prefer it over related tools like get_agent_run or list_agent_runs, nor does it state when not to use it. Usage is implied rather than stated as a rule.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

devtune_search_libraryCInspect

Hybrid-search the project Library. Current versions are returned by default.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNo
queryYes
cursorNoOpaque cursor from the previous response.
includeSupersededNo

TDQS

C2.8/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden of behavioral disclosure. It does mention that 'current versions are returned by default,' which hints at the includeSuperseded parameter behavior. However, it does not explicitly state whether the operation is read-only, nor does it describe side effects, pagination behavior, or sorting. The single behavioral note adds some value but leaves significant gaps.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is extremely concise and front-loaded, which is positive. However, it is under-specified: two short sentences omit essential parameter details and usage context. While each sentence earns its place, the brevity compromises completeness.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a search tool with 4 parameters, no output schema, and no annotations, this description is incomplete. It does not explain how to construct a query, how pagination works, what the response looks like, or the effect of limit/cursor. An agent would need to infer too much.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 25% (only cursor has a description). The description barely compensates: it alludes to includeSuperseded via 'current versions are returned by default' but does not explain query (search term) or limit (result cap). With three of four parameters undocumented in both schema and description, the tool is under-specified for effective invocation.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('hybrid-search') and resource ('project Library'), which clearly indicates the tool's purpose. It distinguishes from siblings like list_library and read_library_item by implying a search functionality, though it does not explicitly name alternatives. Overall, it is unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no guidance on when to use this tool versus alternatives. It does not mention when to use list_library instead, nor does it specify any exclusions (e.g., 'to browse all items, use list_library'). The context is implied but the user is left to infer.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

devtune_send_dex_messageBInspect

Send an idempotent message to a Dex thread, steering its active task or resuming a metered task after expiry.

ParametersJSON Schema
NameRequiredDescriptionDefault
messageYes
threadIdYes
messageIdYes

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the behavioral disclosure burden. It does disclose the idempotency trait and the effect on the thread's task state, which is valuable beyond the schema. However, it omits other behaviors such as error handling, side effects on non-metered threads, or authentication requirements, leaving gaps for a mutation tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single efficiently packed sentence that front-loads the core action and purpose. It avoids redundancy, though it uses jargon ('Dex', 'metered task') that might require domain knowledge. Overall, it earns its place without unnecessary verbosity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a mutation tool with no annotations and no output schema, the description is too brief. It does not mention prerequisites (e.g., thread existence), what happens on failure, or the expected response. Given 0% schema coverage and the need to guide an agent, this is a significant gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It only implicitly references 'message' and idempotency, but it does not explain the semantics of 'threadId' or 'messageId' (e.g., that messageId is the idempotency key). An agent has insufficient guidance to construct correct inputs from the description alone.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action ('send an idempotent message') and the resource ('Dex thread'), with a specific purpose ('steering its active task or resuming a metered task after expiry'). This is distinct from the sibling 'devtune_start_dex_conversation' which initiates a new conversation, and the verb+resource construction is unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage for existing threads by describing 'steering' and 'resuming' actions, but it does not explicitly name alternatives or state when not to use it. Compared to siblings like 'start_dex_conversation', the usage context is implied rather than explicit, so it falls short of a 4.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

devtune_start_dex_conversationCInspect

Start an idempotent, metered Dex conversation for the project using the same web-channel contract as the in-app dock.

ParametersJSON Schema
NameRequiredDescriptionDefault
messageYes
messageIdYes

TDQS

C2.7/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are present, so the description carries the full burden of behavioral disclosure. It does mention 'idempotent' (safe to retry) and 'metered' (likely cost/rate implications), which are valuable traits. However, it does not disclose side effects such as whether a new thread is created, what the response contains, or any prerequisites like authentication or existing thread context.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, focused sentence that front-loads the key action and qualities. It is concise and avoids redundancy. However, it could be slightly restructured to include usage guidance without becoming verbose, so it loses a point for not maximizing the space.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool that starts a conversation, the description is incomplete: it does not explain what a 'Dex conversation' is, what the 'web-channel contract' entails, what the expected output is (e.g., a thread ID), or how to handle the required messageId (e.g., for idempotency). With no output schema and no annotations, the agent is left with insufficient context to call it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% – the schema provides no descriptions for 'message' or 'messageId'. The description does not mention either parameter or explain their purpose or format. Since the description must compensate for the schema's lack of detail and it fails to do so, this is a significant gap for an agent trying to construct a valid call.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('Start') and resource ('Dex conversation'), and adds useful qualifiers ('idempotent, metered'). However, it doesn't explicitly differentiate this from the sibling devtune_send_dex_message or other conversation-related tools, so while the purpose is clear, the boundary to alternatives is implied rather than stated.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides minimal context about when to use this tool ('same web-channel contract as the in-app dock') but offers no explicit guidance on when to choose it over devtune_send_dex_message, devtune_list_dex_threads, or other siblings. No exclusions or conditions are mentioned, leaving the agent to infer appropriate usage.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

devtune_start_prompt_set_analysisA
Destructive
Inspect

Obtain explicit human approval of the credit estimate before starting or retrying analysis and pass that exact value as approvedEstimatedCredits. A changed estimate returns a conflict requiring renewed approval. Reserves estimated credits and settles actual usage; an active run is reused. Produces proposals without applying them. Requires actions.write and a user-attributed credential with devtune.manage. humanApproved attests to approval obtained from the human, not a decision the agent can make.

ParametersJSON Schema
NameRequiredDescriptionDefault
humanApprovedYes
approvedEstimatedCreditsYes

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond the annotations (readOnly=false, destructive=true, idempotent=false), the description discloses meaningful behavior: it reserves estimated credits and settles actual usage, reuses active runs, produces proposals without applying them, and requires specific permissions. It also clarifies that humanApproved must attest to genuine human approval, which is critical semantic context.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is dense but efficiently structured, leading with the primary mission and then layering constraints, side effects, and permission requirements. Every sentence adds unique value, with no filler or redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Despite lacking an output schema, the description provides enough context for correct invocation: prerequisites, parameter semantics, side effects, error condition, and post-condition. The agent knows what triggers a conflict, what happens to credits, and what the tool does not do (apply proposals).

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With 0% schema description coverage, the description fully carries parameter meaning. It explains that approvedEstimatedCredits must be the exact pre-approved estimate and that a changed estimate causes a conflict requiring renewed approval. It also defines humanApproved as evidence of human approval, not an agent decision.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the verb and resource: start or retry prompt set analysis after obtaining human approval of the credit estimate. It also distinguishes this from the read-only sibling get_prompt_set_analysis by emphasizing the action of starting/retrying an analysis run.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description clearly indicates when to use it: before starting or retrying analysis, and it notes that an active run is reused. It does not explicitly name alternatives or when not to use it, but the context is clear and actionable.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

devtune_unclaim_content_changeAInspect

Unclaim an unchanged manual claim using its current claimToken from devtune_get_content_changes. Obtain human approval first. Archives the action and removes its measurement from totals; preserves history. Rejects edited, linked or legacy claims. Requires actions.write.

ParametersJSON Schema
NameRequiredDescriptionDefault
claimTokenYes
changeEventIdYes

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations supplied, the description carries the full burden of behavioral disclosure. It covers key side effects: archives the action, removes its measurement from totals, preserves history, rejects edited/linked/legacy claims, and mandates the 'actions.write' permission. It does not clarify what happens on failure or whether the unclaim is idempotent, but the provided details are substantial and non-contradictory.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is four compact sentences: the core operation, the prerequisite for approval, the state-change consequences, rejection criteria, and the required permission. There is no filler or repetition of schema/annotation fields; every sentence earns its place and the purpose is front-loaded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool is a mutation with no annotations and no output schema, so the description should provide enough operating context. It covers approval, side effects, rejection criteria, and permissions, but it leaves two notable gaps: the meaning of changeEventId (also visible in parameter_semantics) and what the tool returns or how success/failure is signaled. Given the overall simplicity of a two-parameter action, this is a reasonable but incomplete description.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate for both parameters. It explains claimToken at a high level (must be the current token derived from devtune_get_content_changes), but changeEventId is not described at all. With only two parameters, leaving half of them semantically undocumented is a significant gap for an agent trying to invoke the tool correctly.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific verb ('Unclaim') and resource ('an unchanged manual claim'), and clarifies the provenance of the token with 'using its current claimToken from devtune_get_content_changes.' It is clearly distinct from sibling tools like claim_content_change, and the scope ('unchanged manual claim') prevents confusion with list or retrieval tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives strong usage context: requires a human approval first, obtains the claim token from get_content_changes, and rejects edited, linked, or legacy claims. It effectively tells an agent when to use the tool and when not to, though it does not explicitly name the sibling alternative for re-claiming content (e.g., devtune_claim_content_change), which would make the guidance fully explicit.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

devtune_update_actionA
Idempotent
Inspect

Rename an action, set its owner, or both. Works on accepted and live actions; restore an archived action first. DevTune's suggestions, including dismissed and snoozed ones, keep their title and have no owner, so accept one first. assignedTo is a member's user id or sign-in email; pass null to unassign. Owner ids and names are on devtune_get_actions. Returns the action with assignedTo and ownerName. Does not notify the owner or mark the brief out of date.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With readOnlyHint=false, idempotentHint=true, and destructiveHint=false already declared, the description adds meaningful context the annotations don't: 'Does not notify the owner or mark the brief out of date,' plus the return shape. This is real side-effect disclosure rather than a restatement of the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The purpose is front-loaded and every sentence carries information (preconditions, parameter format, side effects, return value). It is dense but skimmable, with no filler sentences, though it is longer than strictly necessary.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a mutation tool with annotations covering the safety profile and no output schema, the description is nearly complete: it covers preconditions, parameter format, side effects, and the return value ('the action with assignedTo and ownerName'). Little an agent needs to invoke it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema is an empty object (0 params), which sets a baseline of 4, and the description adds genuine semantics for the implied fields: 'assignedTo is a member's user id or sign-in email; pass null to unassign' and a pointer to devtune_get_actions for valid owner ids. This compensates for the undocumented schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb+resource pair: 'Rename an action, set its owner, or both,' which clearly states the mutation scope. It is distinct from sibling reads like devtune_get_actions and creates like devtune_create_prompt_action, though it never names a sibling to reinforce the distinction.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives explicit preconditions and negative cases: 'Works on accepted and live actions; restore an archived action first' and 'DevTune's suggestions... accept one first.' This tells the agent when the tool will and won't succeed, but it does not name an alternative tool for those preconditions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

devtune_update_knowledge_profileAInspect

Apply an immediate human-authored edit to one Knowledge Profile field. System and agent changes must still use proposals.

ParametersJSON Schema
NameRequiredDescriptionDefault
valueYes
targetFieldYes

TDQS

A3.9/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and it does reveal key behavior: the edit is immediate, human-authored, limited to one field, and system or agent changes are not permitted. It stops short of disclosing whether the existing value is overwritten, whether the change is reversible, or what the response looks like, but the most critical governance behavior is clearly stated.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two short sentences with no filler. The main action and immediate/human constraint are front-loaded, and the second sentence earns its place by adding an essential exclusion. Every word contributes to correct tool selection.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The schema is simple, but there are no annotations and no output schema, so the description needed to explain parameter meaning and side effects. It explains the high-level routing decision well but leaves targetField semantics, value expectations, and post-edit behavior unspecified, making the overall picture incomplete for an agent that must actually invoke it.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description needed to compensate for the schema's silence, but it does not. It never mentions targetField or value, does not explain the three allowed field names, and gives no guidance on how value should be formatted. The phrase 'one field' is the only implicit hint and is insufficient.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('Apply') with a clear resource ('Knowledge Profile field') and a precise scope ('one field'). It also distinguishes itself from proposal-based updates by saying the edit is immediate and human-authored, which separates it from siblings like review_knowledge_proposal.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly states when to use the tool ('immediate human-authored edit') and when not to ('System and agent changes must still use proposals'). This gives the agent a clear exclusion rule and points to the proposal workflow as the alternative, even without naming the sibling tool directly.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

devtune_update_library_itemCInspect

Change a Library item’s agent availability or manually detach/relink a Superseded version.

ParametersJSON Schema
NameRequiredDescriptionDefault
itemIdYes
successorIdNo
groundingEnabledNo

TDQS

C2.7/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden of behavioral disclosure. It discloses that the tool mutates a Library item and can manually detach/relink a Superseded version, but it does not state whether changes are reversible, what side effects occur on linked versions, whether permissions are required, or what the response contains. For a mutation tool with zero annotation coverage, this is a significant gap.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single sentence with no filler and front-loads the primary purpose. It is concise, though it could be slightly more structured by separating the two operation types.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a mutation tool with no annotations, no output schema, and 0% schema description coverage, the description is incomplete. It does not explain the effect of successorId null, whether groundingEnabled is required for the availability change, or what happens to a Superseded item when relinked. An agent would need to open the schema and guess at semantics.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate, but it only loosely maps to the parameters: 'agent availability' likely maps to groundingEnabled, and 'detach/relink a Superseded version' maps to successorId. itemId is implied as the target. The description does not explain the semantics of successorId null vs a UUID, nor the meaning of groundingEnabled, leaving the agent to infer parameter meaning from names alone.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('Change') and resource ('Library item'), and names two distinct operations: changing agent availability and detaching/relinking a Superseded version. It is clear enough to distinguish from siblings like devtune_read_library_item and devtune_upload_library_item, though it does not explicitly name a sibling alternative.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives no guidance on when to use this tool versus alternatives such as devtune_update_knowledge_profile or devtune_upload_library_item. It implies usage through the operation description but does not state prerequisites, when not to use it, or which sibling covers related update scenarios.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

devtune_upload_library_itemBInspect

Upload a text document into the project Library, optionally replacing an existing upload as its next version.

ParametersJSON Schema
NameRequiredDescriptionDefault
titleYes
contentYes
replaceItemIdNo

TDQS

B3/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It only mentions 'upload' and optional replacement, but does not disclose permissions, whether old versions are preserved, idempotency, or what happens on failure. Critical behavioral context for a write operation is missing.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

One sentence with no waste. The purpose is front-loaded, and the optional replacement behavior is mentioned early. Excellent conciseness.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given there are no annotations and no output schema, the description is too sparse. It omits error handling, size limits (though schema has maxLength), auth requirements, and the exact semantics of replaceItemId. An agent would lack sufficient context to use this tool reliably, especially when deciding between upload and update.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It mentions 'text document' and 'optionally replacing' but does not explain the individual parameters (title, content, replaceItemId) or their constraints. While parameter names are self-explanatory, the description adds minimal value beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a clear action (upload) and resource (text document into Library), and mentions optional replacement as next version. However, it does not differentiate from the sibling devtune_update_library_item, leaving ambiguity about when to upload vs update.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Implies usage: create a new library item or replace an existing one as a new version. But it does not explicitly state when to use this over alternatives like update_library_item, nor any exclusions or prerequisites.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 1 tool update
    • Addeddevtune_update_action
  2. 1 tool update
    • Addeddevtune_get_project_usage
  3. 1 tool update
    • Changeddevtune_get_actions1 field changed
      • changedInput schema / properties / detailLevel / description
        Previous value: -"Response detail level. summary is default. context adds bounded why-now context, available opportunity evidence, metrics, top evidence with structured Google eligibility when available, brief readiness, and follow-up links."New value: +"Response detail level. summary is default. context adds bounded why-now context, available opportunity evidence, metrics, top evidence with structured Google eligibility when available, the external pages winning its prompts (competitor and third-party, in the order briefs study them), brief readiness, and follow-up links."
  4. 1 tool update
    • Changeddevtune_get_outcomes_ledger1 field changed
      • changedInput schema / properties / interventionId / description
        Previous value: -"Include 7, 14, 30 and 90 day impact windows for this completion-based measurement."New value: +"Include 7, 14, 30 and 90 day impact windows (citations, presence and answer fetches) for this completion-based measurement."
  5. 1 tool update
    • Changeddevtune_get_actions3 fields changed
      • changedInput schema / properties / status / description
        Previous value: -"Filter by public action status. Use active for recommendations and backlog/in_progress/blocked/done/canceled for adopted work. The aliases open, completed, and dismissed map to backlog, done, and canceled."New value: +"Filter by action status. Omit it to get every status."
      • changedInput schema / properties / status / enum
        Previous value: -[
        -  "active",
        -  "backlog",
        -  "in_progress",
        -  "blocked",
        -  "done",
        -  "canceled",
        -  "open",
        -  "completed",
        -  "dismissed"
        -]New value: +[
        +  "suggested",
        +  "accepted",
        +  "live",
        +  "archived"
        +]
      • removedInput schema / properties / surface
        Removed value: -{
        -  "description": "Filter to the recommendation feed or the adopted backlog surface.",
        -  "enum": [
        -    "recommendation",
        -    "backlog"
        -  ],
        -  "type": "string"
        -}
  6. 1 tool update
    • Changeddevtune_get_outcomes_ledger1 field changed
      • changedInput schema / properties / status / enum
        Previous value: -[
        -  "measured",
        -  "in_flight",
        -  "blind",
        -  "unmatched_url",
        -  "sensor_gap",
        -  "win",
        -  "loss",
        -  "still_measuring",
        -  "unmeasured",
        -  "conclusive"
        -]New value: +[
        +  "measured",
        +  "in_flight",
        +  "blind",
        +  "unmatched_url",
        +  "sensor_gap",
        +  "improved",
        +  "mixed",
        +  "no_clear_change",
        +  "declined",
        +  "still_measuring",
        +  "unmeasured",
        +  "conclusive"
        +]
  7. 1 tool update
    • Changeddevtune_get_outcomes_ledger1 field changed
      • addedInput schema / properties / interventionId
        Added value: +{
        +  "description": "Include 7, 14, 30 and 90 day impact windows for this completion-based measurement.",
        +  "format": "uuid",
        +  "type": "string"
        +}
  8. 1 tool update
    • Changeddevtune_replace_tracking_page3 fields changed
      • removedInput schema / properties / configuredPageLimit
        Removed value: -{
        -  "maximum": 10000,
        -  "minimum": 50,
        -  "type": "integer"
        -}
      • changedInput schema / properties / operation / enum
        Previous value: -[
        -  "track",
        -  "pause",
        -  "set_limit"
        -]New value: +[
        +  "track",
        +  "pause"
        +]
      • addedInput schema / required
        Added value: +[
        +  "url"
        +]
  9. 2 tool updates
    • Addeddevtune_get_page_tracking
    • Addeddevtune_replace_tracking_page
  10. 2 tool updates
    • Addeddevtune_get_visibility_response
    • Addeddevtune_list_visibility_responses
  11. 1 tool update
    • Changeddevtune_update_knowledge_profile1 field changed
      • changedInput schema / properties / targetField / enum
        Previous value: -[
        -  "positioning",
        -  "products_services",
        -  "tone",
        -  "saved_instructions"
        -]New value: +[
        +  "company_context",
        +  "tone",
        +  "saved_instructions"
        +]
  12. 2 tool updates
    • Addeddevtune_acknowledge_retrieval_restriction
    • Addeddevtune_reconsider_retrieval_restriction

Related MCP Servers

  • A
    license
    A
    quality
    C
    maintenance
    Enables brand visibility monitoring across major AI platforms like ChatGPT, Claude, Gemini, and Perplexity. It allows users to track visibility scores, analyze competitor data, and receive actionable insights to improve AI-generated brand recommendations.
    16
    22 npm
    1
    MIT
  • A
    license
    Not graded
    quality
    B
    maintenance
    Enables tracking competitor websites, changelogs, blog feeds, and pricing pages with meaningful diffs, classification, and Markdown digests via MCP tools for listing, adding, removing competitors, running checks, and retrieving digests or changes.
    MIT
  • A
    license
    Not graded
    quality
    B
    maintenance
    Browse IndustryLens's published competitive-intelligence reports and head-to-head competitor comparisons from any AI agent — real, source-backed data.
    MIT
Try in Browser

Glama MCP Gateway

Add one secure layer between your agents and this server.

Resources