iGods GEO Visibility Tool (GVT)
Server Details
GEO visibility tests, score trends, sitemap discovery, and domain monitoring for AI agents.
- Status
- Healthy
- Uptime
- 99.9% over 22 days
- Last Tested
- Transport
- Streamable HTTP · MCP 2025-11-25
- URL
TDQS
Scored across 19 tools
Each of the 19 tools targets a distinct resource and action: running tests vs. batches, schedule management, baseline/latest/trend retrieval, batch status/summary, knowledge/fix/prompt reference, test visibility, and deletion. Descriptions explicitly cross-reference neighboring tools, making boundaries clear and preventing misselection.
Every tool follows a consistent gvt_verb_noun pattern (gvt_run_visibility_test, gvt_list_baselines, gvt_get_batch_summary, gvt_schedule_test), with only justified merged operations like gvt_schedule_test using modes described in its purpose.
19 tools are reasonable for a feature-rich visibility analytics service covering tests, batches, schedules, prompts, knowledge, and fixes, though it sits near the upper end of the typical 3-15 range.
The surface covers the full lifecycle: discovery (prompts, sitemap urls, tests, schedules), execution (single/batch tests), monitoring (status, summary, latest, baseline, trend), maintenance (delete, visibility toggle), and reference (knowledge, fixes). No obvious gaps within the stated domain.
Available Tools
19 toolsgvt_delete_testDelete a testADestructiveIdempotentInspect
Permanently delete a test session and all its associated results. This action cannot be undone. Use gvt_set_test_visibility instead if you only want to hide a test from public view. Deleting a test does not delete its schedule (if one exists): recurring runs continue until removed via gvt_schedule_test delete mode.
| Name | Required | Description | Default |
|---|---|---|---|
| tid | Yes | The 8-character test ID of the test to delete. |
Output Schema
| Name | Required | Description |
|---|---|---|
| message | No | Human-readable confirmation |
| success | No | True when the test was deleted |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already indicate destructive and non-read-only behavior, but the description adds significant context: the action cannot be undone, associated results are deleted, and existing schedules are preserved. These are non-obvious behaviors that the agent cannot infer from the schema or annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences with no wasted words. The first sentence states the core action and scope, the second flags permanence, and the third covers both the alternative and the schedule caveat. Every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description is complete for a destructive single-parameter tool. It covers the action, side effects, reversibility, and alternative routing. An output schema exists, so explaining return values is unnecessary, and the annotation profile already covers safety semantics.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The single parameter 'tid' is already fully documented in the schema as 'The 8-character test ID of the test to delete.' The description does not add further meaning about the parameter beyond what the schema provides, so the baseline score applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states a specific action, resource, and scope: 'Permanently delete a test session and all its associated results.' It explicitly distinguishes itself from the sibling gvt_set_test_visibility, leaving no ambiguity about what this tool does.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description says exactly when to choose an alternative: 'Use gvt_set_test_visibility instead if you only want to hide a test from public view.' It also clarifies the boundary with scheduled tests, telling agents that recurring runs continue unless removed via gvt_schedule_test delete mode.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
gvt_get_baselineGet baseline snapshotARead-onlyIdempotentInspect
Get the oldest baseline test snapshot for a single URL, including expired tests. For pre-computed baseline-vs-latest deltas without fetching full snapshots, use gvt_get_score_trend instead; for baselines across many URLs in one call, use gvt_list_baselines. Use with gvt_get_latest to compare baseline vs current when you need the full snapshot detail on both ends. The url must be passed exactly as the test was run — matching is exact string equality (scheme, host, path, trailing slash), not nearest-match. When unsure of the recorded form, discover it with gvt_list_tests first.
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | The URL to look up (must match the URL used when the test was run). |
Output Schema
| Name | Required | Description |
|---|---|---|
| tid | No | Test ID of the snapshot |
| url | No | The URL this snapshot analyzed |
| oldest | No | Null when no test exists for the URL |
| status | No | Snapshot status (completed, failed, etc.) |
| message | No | Present only in the no-test case |
| isPublic | No | Whether this test is publicly visible |
| testType | No | The type of the test run |
| createdAt | No | ISO 8601 instant this snapshot was created |
| expiresAt | No | ISO 8601 instant this snapshot will expire |
| updatedAt | No | ISO 8601 instant this snapshot was last updated |
| shareableTid | No | Public share ID, non-null when shareable |
| analysisSummary | No | Aggregate scores for this test |
| findingSentence | No | Pre-generated natural language verdict for this snapshot |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the readOnly/idempotent/non-destructive annotations, the description discloses important behavior: it returns the oldest baseline, includes expired tests, and uses exact string equality matching on scheme, host, path, and trailing slash. This materially affects correct invocation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Every sentence earns its place: purpose, alternative tools, companion usage, and exact matching caveat. It is front-loaded with the core action and remains tightly organized without redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given a single required parameter, high schema coverage, an output schema, and annotations covering safety, the description supplies all needed selection and invocation context. It also routes to appropriate sibling tools for related use cases.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema already documents the url parameter, so the baseline is 3. The description adds value by specifying exact string-equality semantics and trailing-slash sensitivity, going beyond the schema's 'must match the URL used when the test was run.'
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: 'Get the oldest baseline test snapshot for a single URL, including expired tests.' It clearly distinguishes this tool from related siblings by naming what it is not and how it differs.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives explicit alternatives: use gvt_get_score_trend for pre-computed deltas, gvt_list_baselines for many URLs, and gvt_get_latest for comparison. It also advises discovering the exact URL form with gvt_list_tests when uncertain.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
gvt_get_batch_statusGet batch statusARead-onlyIdempotentInspect
Retrieves the current progress, state, and test IDs (tids) of an ongoing or completed visibility testing batch. Poll this tool with the batchId returned by gvt_run_visibility_batch until status is completed or failed. A tid appears on each URL entry only after its test completes — then call gvt_get_batch_summary for the compact per-page scores and top shared issues, or gvt_get_test_results for a single test. Batch records are retained for the lifetime of the account (no automated pruning); an unknown, pruned, or foreign batchId returns 404. Read-only: consumes rate limit, never test credits. Duplicate URLs in a batch are collapsed, so counts always sum to totalUrls.
| Name | Required | Description | Default |
|---|---|---|---|
| batchId | Yes | The UUID of the batch to check, returned by gvt_run_visibility_batch. |
Output Schema
| Name | Required | Description |
|---|---|---|
| urls | No | One entry per submitted URL, in submission order |
| counts | No | Per-URL test state counts; the four fields always sum to totalUrls |
| status | No | Aggregate batch state; completed means every per-URL test reached a terminal state |
| batchId | No | The batch that was queried |
| totalUrls | No | URLs in the batch after duplicate collapsing |
| completedAt | No | ISO 8601 instant when the batch reached a terminal state, null while in flight |
| submittedAt | No | ISO 8601 instant of batch submission |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover readOnly/idempotent/non-destructive, but the description goes further: rate-limit consumption without test credits, 404 on unknown/pruned/foreign batchId, account-lifetime retention with no pruning, and deduplication so counts sum to totalUrls. These are non-obvious behaviors an agent needs.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads purpose, then polling loop, then routing, then edge-case behaviors. Every sentence contributes distinct, actionable information with no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Output schema exists, so return values need no explanation. For an async batch-polling tool the description still covers the full lifecycle: polling termination, follow-up routing, retention, error handling, and counting semantics.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and the single batchId param is fully documented there, so the description's restatement that it comes from gvt_run_visibility_batch adds little. Baseline 3 is appropriate when the schema carries the parameter semantics.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (retrieves) and resource (visibility testing batch) plus the exact payload (progress, state, tids). It is clearly distinguished from sibling readers like gvt_get_batch_summary and gvt_get_test_results, which it explicitly routes to.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives an explicit polling loop: call with the batchId from gvt_run_visibility_batch until status is completed or failed. It also names the follow-up tools and the condition (tid present after a test completes) that selects each, leaving nothing to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
gvt_get_batch_summaryGet batch summaryARead-onlyIdempotentInspect
Compact per-page results and pre-computed aggregates for a completed visibility testing batch: per-URL overall and category scores, findingSentence, average overall score, strongest and weakest categories, best and worst scoring pages, and the top shared issue types with their fixIds. One call replaces one gvt_get_test_results call per URL for audit and status-report workflows — a 10-URL batch drops from roughly 1.6 MB to a few KB of context. The batch must be completed (poll gvt_get_batch_status first); an in-flight batch returns its current status without per-page data. Read-only: consumes rate limit, never test credits. For element-level detail on a single page, call gvt_get_test_results with that page's tid.
| Name | Required | Description | Default |
|---|---|---|---|
| batchId | Yes | The UUID of the completed batch to summarize, returned by gvt_run_visibility_batch. |
Output Schema
| Name | Required | Description |
|---|---|---|
| pages | No | One compact entry per URL, in batch order. Failed tests appear with status failed and null scores; they are excluded from the aggregates. |
| counts | No | Per-URL test state counts; the four fields always sum to totalUrls |
| status | No | Aggregate batch state; pages and summary are present only when completed |
| batchId | No | The batch that was summarized |
| message | No | Polling guidance, present when the batch is not completed |
| summary | No | Aggregates computed over the completed pages |
| totalUrls | No | URLs in the batch after duplicate collapsing |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover the read-only/idempotent safety profile, and the description adds substantial non-obvious behavior: in-flight batches silently return status without per-page data, the call consumes rate limit but never test credits, and it quantifies context savings (~1.6 MB to a few KB for 10 URLs).
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the payload description before the routing and prerequisite notes, and every clause carries distinct information. The sentences are long and comma-chained, which costs some readability, but nothing is padding.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a single-parameter read tool with an existing output schema and full annotation coverage, the description covers purpose, prerequisites, cost, edge-case behavior, and sibling routing. An agent has everything needed to invoke it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Only one parameter with 100% schema description coverage, so the schema already documents batchId and its provenance. The description adds no additional syntax or format detail beyond what the schema states, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Compact per-page results and pre-computed aggregates for a completed visibility testing batch') and enumerates exactly what is returned. It clearly distinguishes itself from siblings gvt_get_test_results (element-level detail) and gvt_get_batch_status (polling).
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Names the workflows it serves (audit, status reports), the alternative it replaces (one gvt_get_test_results per URL), the prerequisite (batch must be completed, poll gvt_get_batch_status first), and the fallback for element-level detail. Explicit when-to-use and when-not-to-use.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
gvt_get_fixRead a fix recordARead-onlyIdempotentInspect
Read the canonical fix record for one GVT issue type: severity range, impact, scoring weight, plain-language description, and the recommended remediation. Issue IDs come from the fixId field on issues in gvt_get_test_results output. The same content is served as the gvt://knowledge/fixes/{issue_id} resource template for resource-capable clients. Copy fixId verbatim from the results output — unknown or mistyped IDs return not-found rather than a fuzzy match. Responses are static reference content and cacheable.
| Name | Required | Description | Default |
|---|---|---|---|
| issue_id | Yes | Stable issue identifier, e.g. og_title_missing or color-contrast. Matches the fixId field in gvt_get_test_results output. |
Output Schema
| Name | Required | Description |
|---|---|---|
| ttlMs | No | Suggested client cache lifetime in milliseconds |
| impact | No | Plain-language statement of what the issue costs |
| weight | No | Scoring weight |
| category | No | Analysis category the issue belongs to |
| issue_id | No | Stable issue identifier, e.g. og_title_missing |
| resourceUri | No | The gvt:// resource this record was read from |
| subcategory | No | Finer-grained grouping within the category |
| recommendation | No | The recommended remediation |
| severity_range | No | Severity levels this issue can take |
| human_description | No | Human-readable description of the issue |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already establish read-only, idempotent, non-destructive behavior. The description adds meaningful behavioral detail beyond annotations: responses are static reference content, cacheable, and lookups are exact-match only with no fuzzy fallback. This materially informs the agent's invocation and caching expectations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and front-loaded: the primary purpose and content list appear first, followed by ID provenance, resource-template equivalence, exact-match warning, and cacheability. Each sentence contributes distinct operational value with no filler or redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With one well-documented parameter, an output schema present, and annotations covering safety, the description is complete for selection and invocation. It explains where the ID comes from, how to handle it, and the response characteristics, leaving no critical gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, with issue_id already documented as a stable identifier matching the fixId field. The description adds extra operational semantics beyond the schema: 'Copy fixId verbatim from the results output' and the not-found behavior for unknown or mistyped IDs. This enriches the parameter's meaning beyond the schema baseline.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'Read the canonical fix record for one GVT issue type,' and enumerates the record's contents (severity range, impact, scoring weight, plain-language description, recommended remediation). This makes the tool's function unmistakable and distinct from generic knowledge or list tools among siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear usage context: issue IDs come from the fixId field in gvt_get_test_results outputcars, and IDs must be copied verbatim because unknown or mistyped IDs return not-found. It does not explicitly name alternative tools or state when not to use this tool, but the guidance is sufficient for an agent to invoke it correctly.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
gvt_get_knowledgeRead GVT knowledgeARead-onlyIdempotentInspect
Read GVT reference knowledge: how scoring works, score bands and category weights, what each analysis category checks, or GEO glossary definitions. Use it to interpret test results and scores correctly instead of guessing. The topic is matched exactly and case-sensitively (Methodology will not match methodology); unrecognized values fail fast with the valid list. The same content is served as MCP resources at gvt://knowledge/app/* for resource-capable clients.
| Name | Required | Description | Default |
|---|---|---|---|
| topic | No | Which knowledge topic to read: methodology (how GVT analyzes pages), scoring (score bands and category weights), categories (what each of the six categories checks), glossary (GEO term definitions), or all (complete knowledge base). | all |
Output Schema
| Name | Required | Description |
|---|---|---|
| uri | No | The gvt:// resource the content came from |
| topic | No | The topic that was served |
| ttlMs | No | Suggested client cache lifetime in milliseconds |
| knowledge | No | Reference content for the requested topic |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, and non-destructive. The description adds valuable behavioral detail: exact case-sensitive matching, failure behavior (fail fast with valid list), and the service as MCP resources. This provides extra context beyond annotations without contradicting them.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences with no wasted words. The main purpose is front-loaded, followed by important matching details and resource alternative. Every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The output schema exists, so return values are covered. The description covers purpose, usage, matching behavior, and failure handling. It's complete for a simple read-only knowledge tool with one optional parameter.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema description covers 100% of parameters, so baseline 3 is appropriate. The description mentions exact matching and case-sensitivity, which reinforces the enum behavior, but doesn't add additional semantics beyond what the schema already explains.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool reads GVT reference knowledge, listing specific content areas (scoring, bands, weights, categories, glossary). It distinguishes itself from sibling tools that run tests or get results, making selection unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear context: use it to interpret test results instead of guessing. It doesn't explicitly name alternative tools, but given the tool's unique read-only knowledge role among siblings (e.g., gvt_get_test_results), the usage context is clear enough. Lacks explicit when-not conditions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
gvt_get_latestGet latest snapshotARead-onlyIdempotentInspect
Get the most recent non-expired test snapshot for a single URL. For pre-computed baseline-vs-latest deltas without fetching full snapshots, use gvt_get_score_trend instead. Use with gvt_get_baseline to compare baseline vs current when you need the full snapshot detail on both ends. The url must be passed exactly as the test was run — matching is exact string equality, not nearest-match; an unrecorded URL returns 404. Considers only the authenticated caller's own non-expired tests.
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | The URL to look up (must match the URL used when the test was run). |
Output Schema
| Name | Required | Description |
|---|---|---|
| tid | No | Test ID of the snapshot |
| url | No | The URL this snapshot analyzed |
| latest | No | Null when no test exists for the URL |
| status | No | Snapshot status (completed, failed, etc.) |
| message | No | Present only in the no-test case |
| isPublic | No | Whether this test is publicly visible |
| testType | No | The type of the test run |
| createdAt | No | ISO 8601 instant this snapshot was created |
| expiresAt | No | ISO 8601 instant this snapshot will expire |
| updatedAt | No | ISO 8601 instant this snapshot was last updated |
| shareableTid | No | Public share ID, non-null when shareable |
| analysisSummary | No | Aggregate scores for this test |
| findingSentence | No | Pre-generated natural language verdict for this snapshot |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, and destructiveHint false. The description adds valuable specifics: exact string equality matching (not nearest-match), 404 on unrecorded URL, and scoping to the caller's own non-expired tests. While it doesn't detail response contents, the output schema exists to cover that. It goes beyond the annotations without contradicting them.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is concise—four sentences, each serving a purpose: primary function, alternative tool, combination usage, and matching constraint. Information is front-loaded, and there is no redundant phrasing.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a single-parameter read-only tool with a full output schema and safety annotations, the description covers all necessary operational details: what it returns (latest non-expired snapshot), how to invoke it correctly (exact URL match), what errors to expect (404), and scope (own tests). No critical gap remains.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% (the url parameter description matches the intent), and the description adds the exact-match requirement and 404 behavior, which the schema does not specify. It also clarifies the operation is for a single URL. This adds meaningful context beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the operation: 'Get the most recent non-expired test snapshot for a single URL.' It distinguishes itself from the sibling gvt_get_score_trend by noting that tool is for pre-computed deltas, and also references gvt_get_baseline for full snapshot comparison. This is a specific verb+resource with explicit sibling differentiation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly names when to use gvt_get_score_trend instead ('For pre-computed baseline-vs-latest deltas without fetching full snapshots') and when to combine with gvt_get_baseline ('when you need the full snapshot detail on both ends'). It also provides a critical constraint: URL must match exactly, with 404 on mismatch. This gives clear, actionable selection guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
gvt_get_promptRender a prompt workflowARead-onlyIdempotentInspect
Render a GVT prompt workflow with concrete argument values, returning the complete step-by-step instruction text ready to follow. Also returns the effective arguments that were bound (provided values plus defaults for omitted optionals). Prompt names and their arguments come from gvt_list_prompts; an unknown name or argument key errors rather than guessing. The rendered text references GVT tools by name — call them as instructed.
| Name | Required | Description | Default |
|---|---|---|---|
| name | Yes | Prompt name from gvt_list_prompts, e.g. gvt_setup_domain_monitoring. | |
| arguments | No | Argument values for the prompt, keyed by argument name (e.g. {"domain": "example.com", "frequency": "weekly"}). Omit optional arguments to use their defaults. |
Output Schema
| Name | Required | Description |
|---|---|---|
| name | No | The prompt name that was rendered |
| text | No | Fully interpolated instruction text referencing GVT tools by name |
| arguments | No | Effective arguments after binding defaults |
| description | No | The prompt catalog description |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare idempotentHint=true, readOnlyHint=true, and destructiveHint=false, so no contradiction. The description adds value beyond annotations by disclosing that it errors on unknown names/arguments rather than guessing, and that it returns effective arguments with defaults applied. It also states the output includes step-by-step instruction text, which is not in annotations. Minor gap: no mention of output format beyond being text, but with an output schema present, that is acceptable.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is three sentences: the first gives the core purpose and output, the second explains source and error behavior, and the third instructs how to use the output. It is tightly written with no filler, though it front-loads the purpose but leaves the error caveat until the second sentence. Slightly more wordy than ideal but efficient.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given two parameters with full schema coverage and an output schema present, the description is complete for invoking the tool. It covers argument binding default behavior and error handling. One minor gap: it doesn't explicitly say that the returned steps might require additional context from gvt_list_prompts, but that is strong enough. The output schema likely details the return structure, so this is sufficient.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%: the name parameter description is clear and provides an example, and arguments is described with a key-value example. The description adds value by clarifying that omitted optional arguments use defaults and that it returns effective arguments, but this is about behavior rather than parameter format. The schema already explains both parameters well, so description doesn't add significant new meaning.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool renders a GVT prompt workflow with concrete arguments and returns the instruction text plus bound arguments. It distinguishes itself from sibling tools like gvt_list_prompts by referencing it as the source of prompt names/arguments, and from other get_* tools by focusing on prompt rendering rather than retrieving data.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly states when to use: when you need a rendered prompt workflow. It references sibling gvt_list_prompts for obtaining valid prompt names/argumentsages, and warns that unknown names/arguments cause errors, guiding the agent to check that list first. It also instructs to call the GVT tools referenced in the rendered text, avoiding misuse.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
gvt_get_score_trendGet score trendARead-onlyIdempotentInspect
Compare the oldest baseline vs the latest non-expired test for a single URL or an entire domain, with pre-computed per-category deltas (latest minus baseline). Provide exactly one of url or domain. Domain mode matches the exact domain, www., and subdomains (alphabetical, capped by limit). Page through the whole domain with offset: offset=0 for the first page, offset=limit for the second, following hasMore/nextOffset. Returns scores and deltas only, not full snapshot detail; for complete snapshots of one URL use gvt_get_baseline and gvt_get_latest. This replaces paging the full test history and computing comparisons client-side, one call per site.
| Name | Required | Description | Default |
|---|---|---|---|
| url | No | Single URL scope (XOR with domain). | |
| limit | No | Cap on returned records in domain mode (max 100, default 25). | |
| domain | No | Domain scope — matches the domain, www.<domain>, and subdomains (XOR with url). | |
| offset | No | Number of matched URLs to skip in domain mode (row skip, not a page number). Use offset=0 for the first page, offset=limit for the second. Ignored in single-URL mode. |
Output Schema
| Name | Required | Description |
|---|---|---|
| url | No | The compared URL in url mode, else null |
| limit | No | Max records per page in domain mode |
| scope | No | Which mode the comparison ran in |
| ttlMs | No | Suggested client cache lifetime in milliseconds |
| domain | No | The compared domain in domain mode, else null |
| offset | No | Row offset applied to this page |
| trends | No | One row per matched URL, baseline vs latest with deltas |
| compare | No | Comparison window settings used |
| hasMore | No | True when more pages exist |
| returned | No | Records in this page |
| urlCount | No | Total matched URLs before pagination |
| truncated | No | Deprecated. Use `hasMore` instead. |
| nextOffset | No | Offset for the next page, null on the last page |
| windowSummary | No | Summary of window comparison results |
| includeSubdomains | No | Whether subdomains were included in the domain search |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the read-only/idempotent annotations, it discloses meaningful behavior: only non-expired tests are considered, deltas are pre-computed, results are limited to scores/deltas (not full snapshots), domain mode is alphabetical and capped by limit, and pagination follows hasMore/nextOffset. No contradictions with annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with the core comparison, then logically moves through scope, pagination, return scope, and alternatives. Every sentence has a purpose, though the pagination explanation partly duplicates the schema and could be tightened.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the two-mode behavior, pagination, and likely domain expansion, the description covers all operational essentials: what the call does, how to constrain scope, how to page, what is returned, and which sibling to use for full snapshots. The output schema and annotations handle the remaining detail.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3. The description reinforces the XOR relationship and pagination flow, but most of that detail already exists in the input schema (offset is a row skip, ignored in single-URL mode, default limit 25, max 100). It adds clarity but not much new semantic information.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific verb ('Compare'), a clear resource (oldest baseline vs latest non-expired test), and precise scope (single URL or domain). It also distinguishes itself from siblings by stating it returns scores and deltas only, and that full snapshots belong to gvt_get_baseline and gvt_get_latest.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives explicit usage rules: provide exactly one of url or domain, describes domain-mode matching and pagination, and tells the agent when to prefer siblings ('for complete snapshots of one URL use gvt_get_baseline and gvt_get_latest'). This is actionable, not merely implied.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
gvt_get_test_resultsGet test resultsARead-onlyIdempotentInspect
Get results for a visibility test by tid: overall scores, findingSentence, shareableTid, and per-category resultData. The tid is an opaque 8-character HashID — copy it verbatim from gvt_run_visibility_test, gvt_list_tests, or gvt_get_batch_status; partial or reformatted values will not match, and results are scoped to the caller's own tests. If status is "pending" or "running", wait and retry. findingSentence is a pre-generated natural-language verdict; use it as the primary summary. shareableTid is non-null when the result is public and linkable. detail selects the payload shape: "summary" (recommended for reporting workflows) returns analysisSummary plus the top issues per category without the semantic tree, per-element markup, or the nojs_rendered duplicate pass; "full" (the default) returns complete per-category resultData for all 7 analysis types (heading, semantic, accessibility, schema, javascript, nojs, social). For compact per-page scores across a whole completed batch, prefer gvt_get_batch_summary instead of one call per tid.
| Name | Required | Description | Default |
|---|---|---|---|
| tid | Yes | Opaque 8-character hash ID returned by gvt_run_visibility_test or gvt_list_tests. Not an integer. | |
| detail | No | Payload shape. summary: analysisSummary plus the top issues per category (deduplicated by type, capped at 10, sorted by severity then weight) — no semantic tree, per-element markup, or nojs_rendered pass. full: complete resultData for all analysis types (previous behavior, the default). | full |
Output Schema
| Name | Required | Description |
|---|---|---|
| tid | No | The test ID that was polled |
| url | No | The URL that was analyzed |
| detail | No | Present with value "summary" when detail=summary was requested; omitted from full responses so existing callers see no change |
| status | No | Pending and running mean poll again; completed means full results are present |
| results | No | Per-category result records. detail=full: each carries analysisType and complete resultData (semantic tree, per-element issues); issues carry fixId for gvt_get_fix. detail=summary: each carries analysisType, score, and the top issues deduplicated by type (capped at 10 per category); the nojs_rendered duplicate pass is omitted — analysisSummary.categoryIssueCounts keeps the true totals. |
| isPublic | No | Whether this test result is publicly accessible |
| testType | No | Type of test performed (e.g., URL or direct HTML paste) |
| createdAt | No | ISO 8601 instant of test creation |
| expiresAt | No | ISO 8601 instant when this test session and its results will expire and be deleted |
| updatedAt | No | ISO 8601 instant of the last update to this test session |
| shareableTid | No | Non-null when the result is publicly shareable. Public share identifier for this session. Equal to tid when the session is shareable; null when the session is private. Presence of a value is the indicator that share links may be generated. |
| analysisSummary | No | Aggregate scores for one test session |
| findingSentence | No | Pre-generated natural language verdict; use as the primary summary |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover read-only/idempotent/non-destructive, but the description adds real context beyond them: results are scoped to the caller's own tests, pending/running states require a wait-and-retry, shareableTid signals public linkability, and the two detail modes have different payload contents. It stops short of pagination/size limits or output-shape guarantees.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One dense paragraph, front-loaded with the returned fields before constraints and mode selection; every sentence carries information. It is longer than strictly needed (repeats some detail-enum content already in the schema), which keeps it from a 5.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, yet the description still characterizes key return fields and the pending/running lifecycle, leaving no ambiguity about when the call is useful or how to interpret its results. Complete for a 2-parameter read tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so baseline is 3; however, the description adds meaning beyond the schema by warning that tid must be copied verbatim (partial/reformatted values fail to match) and by recommending the 'summary' mode for reporting versus the 'full' default.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Get results for a visibility test by tid') and enumerates what is returned (overall scores, findingSentence, shareableTid, per-category resultData). It also distinguishes itself from siblings by naming gvt_get_batch_summary for batch-level scores and citing gvt_run_visibility_test/gvt_list_tests as tid sources.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly routes the agent away from one-call-per-tid to gvt_get_batch_summary for compact per-page batch scores, and gives the retry condition when status is pending/running. It also advises which detail mode suits reporting workflows, so alternatives and conditions are both covered.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
gvt_list_baselinesList baselines (batch)ARead-onlyIdempotentInspect
Get the oldest baseline test snapshots for a batch of URLs (up to 100) in a single call. Each result includes the URL and its oldest test object (null if no test exists for that URL). For a single URL use gvt_get_baseline; for pre-computed baseline-vs-latest deltas use gvt_get_score_trend. This bulk form is ideal for monthly/periodic batch comparisons: fetch baselines in bulk here, then call gvt_get_latest per URL. One call replaces up to 100 gvt_get_baseline calls; split larger URL sets into consecutive calls of at most 100. Order of results follows the submitted order.
| Name | Required | Description | Default |
|---|---|---|---|
| urls | Yes | Array of URLs to look up (max 100). |
Output Schema
| Name | Required | Description |
|---|---|---|
| found | No | URLs that had at least one test |
| ttlMs | No | Suggested client cache lifetime in milliseconds |
| results | No | One entry per requested URL, in request order |
| requested | No | Number of URLs requested |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description adds behavioral details beyond the annotations: it states that results include the URL and its oldest test object (or null if no test exists), that results follow the submitted order, and that the batch limit is 100. While the annotations already cover read-only, idempotency, and open-world hints, the description provides useful additional constraints about batch limits and null handling.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is concise and well-structured: it opens with the core functionality, then explains output structure, differentiators, use cases, batching guidance, and ordering – all in a few sentences. Every sentence adds value without redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's moderate complexity (one parameter, clear output structure, batch limits), the description covers everything needed: purpose, output shape, limitations, alternatives, and usage patterns. The presence of an output schema reduces the need to detail return values, and the description fills the remaining gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema already documents the 'urls' parameter with a description, but the tool description adds significant context: it explains that the URLs are used for baseline lookups, that the batch supports up to 100 items, and prescribes how to handle larger sets ('split larger URL sets into consecutive calls of at most 100'). This goes beyond the schema's basic description.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly identifies the tool's function: retrieving the oldest baseline test snapshots for a batch of URLs in a single call. It explicitly distinguishes it from similar tools like gvt_get_baseline (single URL) and gvt_get_score_trend (pre-computed deltas), making the purpose unmistakable.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides explicit guidance on when to use this tool: for bulk operations and monthly/periodic batch comparisons(via 'ideal for monthly/periodic batch comparisons'), when to use alternatives ('For a single URL use gvt_get_baseline'), and notes that it replaces up to 100 single calls, with clear batch size limits. It also suggests a follow-up step('then call gvt_get_latest per URL').
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
gvt_list_promptsList prompt workflowsARead-onlyIdempotentInspect
List the GVT prompt workflows: user-invocable recipes that chain GVT tools into complete tasks (check a page, review a monthly batch, analyze regressions, verify fixes, set up domain monitoring, report client status). Each entry lists its arguments. Use gvt_get_prompt to render a prompt with concrete argument values, then follow the rendered instructions to run the workflow. This catalog is static and costs nothing to call on every session start — it never queues tests or consumes credits.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
| ttlMs | No | Suggested client cache lifetime in milliseconds |
| prompts | No | The prompt workflow catalog |
| promptCount | No | Number of prompt workflows in the catalog |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already indicate readOnlyHint, idempotentHint, and destructiveHint are all true/false appropriately sums up safety. The description adds value by stating the catalog is static, costs nothing, never queues tests or consumes credits, which goes beyond what annotations provide. This is exactly the kind of behavioral transparency that helps an agent know side-effect-free it is.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, all packed with useful information: what it lists, what the entries contain, how to use the result, and the cost/static nature. No fluff. The most critical information (listing workflows and how to render them) is front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter, read-only catalog tool, the description covers everything an agent needs: what it returns, how to act on it, and the fact it's free. The output schema likely lists the structure, so return values are covered. The description is complete for a tool of this simplicity.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters, so the schema is complete. The description doesn't add param-specific semantics because there are none. It does explain what each entry in the returned list contains ('Each entry lists its arguments'), which adds context about the output beyond just the schema. Given 0 params, a baseline of 4 is appropriate, and the description meets it.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool lists GVT prompt workflows and explains what those are (user-invocable recipes that chain tools into tasks), with examples of the types of workflows. It distinguishes this from siblings like gvt_get_prompt, which renders a specific prompt. The verb 'list' and resource 'GVT prompt workflows' are specific and unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly says when to use this tool: to get a catalog of workflows, and then use gvt_get_prompt to render a specific one. It also states that it costs nothing and can be called on every session start, implying it's appropriate as an initial step. It does not explicitly say when not to use it, but the routing to gvt_get_prompt is clear, and the context of siblings makes the alternative obvious.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
gvt_list_schedulesList schedulesARead-onlyIdempotentInspect
Retrieves the paginated roster of the authenticated caller's test monitoring schedules, allowing filtering by domain, frequency, and status. This is the only way to discover existing schedules and their schids, which the pause/resume/cancel/change_frequency modes of gvt_schedule_test require. Completed one-off rows are hidden by default: batch runs create a throwaway once row per queued URL, so include them only with includeCompletedOnce true. The row's tid is deliberately not surfaced (it is often stale); use gvt_list_tests or gvt_get_latest for test result lookups. Read-only: consumes rate limit, never test credits.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | The maximum number of schedules requested for this page. | |
| domain | No | Filter to schedules matching this domain substring. | |
| offset | No | The number of skipped items from the beginning of the result set. | |
| status | No | Filter by the derived overall state of the schedule. | |
| frequency | No | Filter by repetition interval. | |
| sortDirection | No | Sort order by creation time; desc is newest first. | desc |
| includeSubdomains | No | Whether subdomains should be included in the domain filter match. | |
| includeCompletedOnce | No | Whether to include completed one-off schedules in the results. Hidden by default because batch runs create a throwaway once row per queued URL; include them only when investigating batch history. |
Output Schema
| Name | Required | Description |
|---|---|---|
| limit | No | Page size applied to this response |
| total | No | Total schedules matching the filters, across all pages |
| domain | No | The applied domain filter, null when unfiltered |
| offset | No | Offset applied to this response |
| status | No | The applied status filter, null when unfiltered |
| hasMore | No | True when more results exist beyond this page |
| frequency | No | The applied frequency filter, null when unfiltered |
| schedules | No | One entry per schedule, newest first by creation time |
| nextOffset | No | Offset for the next page, null when hasMore is false |
| includeSubdomains | No | The applied subdomain-inclusion setting |
| allowedFrequencies | No | Frequencies the caller may set when creating schedules |
| maxActiveSchedules | No | The caller's cap on active schedules; add mode enforces this limit |
| includeCompletedOnce | No | Whether completed one-off rows were included in this response |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare readOnly, idempotent, and non-destructive, and the description adds meaningful behavioral detail beyond those: pagination, hidden completed one-off rows due to batch-run throwaway entries, the deliberate omission of tid because it is often stale, and the distinction that it consumes rate limit but never test credits. This is rich, honest context.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Five sentences, each carrying a distinct piece of necessary information: function, unique purpose, hidden-row caveat, tid caveat with alternative routing, and read-only/rate-limit behavior. It is front-loaded with the core purpose and contains no filler or repetition of schema content.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For an 8-parameter, 0-required, read-only list tool with an output schema, the description covers purpose, alternatives, behavioral caveats, and parameter edge cases. The output schema handles return-value details. Nothing an agent needs to invoke this correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Input schema coverage is 100%, with each parameter already having a solid description, so the description does not need to restate them. The description does add useful context for includeCompletedOnce (why it is hidden by default and when to include it), but that is a marginal enhancement over an already complete schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Retrieves'), resource ('paginated roster of the authenticated caller's test monitoring schedules'), and capabilities (filtering by domain, frequency, status). It also distinguishes itself as the only way to discover schedules and schids, clearly separating it from sibling list tools like gvt_list_tests.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly identifies when to use this tool (discovering existing schedules and schids for gvt_schedule_test's pause/resume/cancel/change_frequency modes) and when not to (for test result lookups, pointing to gvt_list_tests or gvt_get_latest). It also clarifies the edge case around includeCompletedOnce, leaving no ambiguity about selection.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
gvt_list_sitemap_urlsList sitemap URLsARead-onlyIdempotentInspect
Discovers and ranks URLs from a domain's XML sitemaps to find the most significant pages (homepages, product pages, recently updated content) before analyzing them. Supports URL-pattern filtering and automatically flags robots.txt-blocked and already-tested pages. Legal/policy boilerplate (privacy, terms, cookies, disclaimers, refunds) is flagged legalPage:true with a -25 significance penalty; use the legalPages filter to drop it (exclude) or isolate it for a policy-coverage audit (only). Parameters group into three jobs: discovery (url), ranking and paging (limit, sort), and filtering (include and exclude URL patterns, excludeDisallowed, excludeTested, excludeScheduled, legalPages).
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | The target website URL or an explicit sitemap .xml URL. | |
| sort | No | How to sort the discovered URLs. 'significance' uses SEO heuristics. | significance |
| limit | No | Maximum number of URLs to return (default 10). | |
| exclude | No | Wildcard patterns to reject (e.g., '*archive*'). | |
| include | No | Wildcard patterns to require (e.g., '*blog*'). Only '*' wildcards are supported. | |
| legalPages | No | Tri-state filter for legal/policy boilerplate pages (privacy, terms, cookies, disclaimers, refund policies, etc). "include" keeps them (default), "exclude" drops them, and "only" returns just the legal/policy pages — useful for auditing a site's policy coverage. | include |
| excludeTested | No | If true, omits URLs the user has already tested. | |
| excludeScheduled | No | If true, omits URLs currently in the testing queue. | |
| excludeDisallowed | No | If true, silently drops URLs that are blocked by robots.txt. |
Output Schema
| Name | Required | Description |
|---|---|---|
| sort | No | The sort actually applied to the results |
| urls | No | Discovered URLs, ranked per the sort setting |
| limit | No | Effective result cap |
| errors | No | Per-sitemap fetch or parse failures; non-empty means discovery was partial |
| hasMore | No | True when the filtered set was cut short by the limit |
| matched | No | URLs remaining after include/exclude and legalPages filters |
| returned | No | URLs actually returned after the limit |
| truncated | No | True if the underlying sitemap parser hit its 5000-URL cap during fetching. |
| legalPages | No | The legal-page filter actually applied |
| bulkUrlLimit | No | How many of these URLs the caller's tier permits in one batch-analyze call. |
| requestedUrl | No | The URL or sitemap URL that was resolved |
| sitemapCount | No | Number of sitemaps parsed |
| robotsTxtFound | No | Whether a robots.txt was found and consulted |
| legalPagesFound | No | How many of the URLs discovered across all sitemaps were detected as legal/policy boilerplate, counted BEFORE any filtering or limiting is applied. |
| totalDiscovered | No | Raw URL count discovered before filtering |
| resolvedSitemaps | No | Sitemap URLs actually parsed |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, openWorldHint=true, idempotentHint=true, and destructiveHint=false. The description goes beyond these by revealing concrete behaviors: it 'automatically flags robots.txt-blocked and already-tested pages', applies a '-25 significance penalty' to legal pages, and groups parameters into three jobs (discovery, ranking/paging, filtering). This adds substantive behavioral context that the annotations do not convey.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is about 100 words—informative without being bloated. It front-loads the core purpose and then systematically covers flags and parameter grouping. While it could be tightened slightly, every sentence contributes useful information and no schema detail is needlessly repeated.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with 9 parameters, an output schema, and rich annotations, the description is remarkably complete. It covers the purpose, the parameter organization, the specific behavioral flags, and the legalPages tri-state use case. The existence of an output schema covers return-value details, so no major gaps remain for an agent to call the tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so each parameter is already documented. The description adds a valuable organizational layer by grouping the nine parameters into three conceptual jobs (discovery, ranking/paging, filtering) and explaining the purpose of the legalPages filter. This helps an agent understand parameter relationships beyond the flat schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('discovers and ranks URLs') and a clear resource ('from a domain's XML sitemaps'), and explains the goal ('to find the most significant pages before analyzing them'). This is distinct from all sibling tools, which concern tests, baselines, prompts, and schedules—no other tool targets sitemap discovery. The purpose is unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description positions the tool as a precursor to analysis ('before analyzing them') and gives explicit guidance on the legalPages filter for either dropping legal boilerplate or isolating it for a policy-coverage audit. It doesn't explicitly name alternative tools or state when not to use it, but the sibling set is clearly different, so the use case is well implied.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
gvt_list_testsList testsARead-onlyIdempotentInspect
List visibility tests for the authenticated user, newest first. Supports filtering by domain (optionally including subdomains), by URL (exact or contains match), by creation date range, by status, and by test type. This is the tid and history discovery surface: use it to find the exact recorded URL string and tids that gvt_get_baseline, gvt_get_latest, and gvt_get_test_results require, and use gvt_get_score_trend instead when you only need baseline-vs-latest deltas — it skips the paging entirely.
| Name | Required | Description | Default |
|---|---|---|---|
| url | No | URL to filter tests by; matched per urlMode (contains by default). | |
| limit | No | Max results to return (default 10, max 100) | |
| domain | No | Domain to filter tests by, e.g. example.com. Subdomains are included unless includeSubdomains is false. | |
| offset | No | Number of records to skip before returning results (not a page number). Use offset=0 for the first page, offset=limit for the second page, etc. | |
| status | No | Test status to filter by: pending, processing, completed, or failed. | |
| urlMode | No | How the url filter matches: "exact" for full-URL equality, "contains" for substring match. | contains |
| testType | No | Test type to filter by: url, html, or html_paste. | |
| createdTo | No | Exclusive creation-date upper bound. YYYY-MM-DD (UTC midnight) or full ISO 8601 instant. | |
| createdFrom | No | Inclusive creation-date lower bound. YYYY-MM-DD (UTC midnight) or full ISO 8601 instant. | |
| sortDirection | No | Sort direction by creation date: desc (newest first, default) or asc (oldest first). | desc |
| includeSubdomains | No | If true (the default), includes tests for subdomains of the given domain. Set false to match the exact domain only. |
Output Schema
| Name | Required | Description |
|---|---|---|
| limit | No | Page size actually applied |
| tests | No | The matching test sessions on this page |
| total | No | Total matching tests before pagination |
| ttlMs | No | Suggested client cache lifetime in milliseconds |
| domain | No | The domain filter actually applied, if any |
| offset | No | Row offset applied to this page |
| hasMore | No | True when more pages exist |
| createdTo | No | Effective upper creation-date bound, if any |
| nextOffset | No | Offset for the next page, null on the last page |
| createdFrom | No | Effective lower creation-date bound, if any |
| includeSubdomains | No | Whether subdomains were included in the domain filter |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare read-only, idempotent, non-destructive behavior, and the description does not contradict them. It adds contextual value by framing the tool as a discovery surface and stating the default ordering (newest first), which supplements the safety profile without duplicating it.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is structured with the primary purpose front-loaded, a concise list of filter capabilities, and a valuable routing note at the end. Every sentence earns its place, and the length is appropriate for an 11-parameter tool.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description fully complements the rich schema and output schema. It explains the tool's role in the workflow, tells the agent exactly how to use it to obtain tids and URL strings for sibling tools, and provides an alternative for simpler needs. Nothing an agent needs to call or route correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 100% schema description coverage, the schema already documents all 11 parameters including defaults and enums. The description mentions filter options (domain, URL, date range, status, test type) but adds no new parametric detail beyond what the schema covers; it is a baseline 3.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool lists visibility tests for the authenticated user, newest first. It further positions it as the 'tid and history discovery surface' and distinguishes it from gvt_get_score_trend, making its purpose specific and unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly tells when to use this tool (to find tids and URL strings needed by related tools) and when not to (use gvt_get_score_trend for deltas). It names the alternative and the condition that selects it, leaving no ambiguity.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
gvt_run_visibility_batchRun batch testsAInspect
Queue up to 500 URLs for batch GEO visibility testing. Returns a batch ID for tracking progress. Optionally add URLs to recurring schedules (daily, weekly, monthly) during the same call. Poll gvt_get_batch_status with the batchId until status is completed or failed; then get the compact per-page score summary with gvt_get_batch_summary, or individual results with gvt_get_test_results. This tool queues paid test runs and consumes credits; for score interpretation, methodology, or reference knowledge, use gvt_get_knowledge instead — it never runs tests. A webhook URL can be provided to receive a completion notification; when set, the response returns webhookSecret, an HMAC key for verifying that the notification genuinely came from GVT. Duplicate URLs in the submitted array are collapsed, so totalUrls may be smaller than the array you sent.
| Name | Required | Description | Default |
|---|---|---|---|
| urls | Yes | Array of URL objects to test (max 500) | |
| webhookUrl | No | Optional HTTPS URL to receive a batch completion webhook |
Output Schema
| Name | Required | Description |
|---|---|---|
| status | No | Always queued at submission time |
| batchId | No | UUID for tracking the batch |
| totalUrls | No | URLs accepted into the batch |
| webhookSecret | No | HMAC secret for verifying the webhook, null when no webhook |
| webhookConfigured | No | True when a completion webhook URL was supplied |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations cover the safety profile (readOnlyHint=false, idempotentHint=false, openWorldHint=true), and the description adds substantial context beyond them: it queues paid runs that consume credits, operates asynchronously with a batch ID, supports an optional webhook that returns an HMAC webhookSecret for verification, and collapses duplicate URLs so totalUrls may be smaller than the input array. This is exactly the extra behavioral detail the annotations cannot convey.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with purpose, then workflow, then cost, then webhook and dedupe caveats. Every sentence carries information, though the paragraph is dense and could be split for scannability. No filler or repetition.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With an output schema present, return values needn't be described, and the description still covers the async workflow, credit cost, webhook verification, and dedupe caveat. Nothing an agent needs to queue and track a batch correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the description adds real meaning beyond the schema: webhookUrl's effect (completion notification + returned webhookSecret for HMAC verification), the dedupe behavior affecting totalUrls, and the recurring-schedule option embedded in each URL object. Only the exact addToSchedule semantics remain schema-only.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (queue) and resource (up to 500 URLs for batch GEO visibility testing) with explicit scope. It cleanly distinguishes itself from the single-URL sibling and from the read-side batch tools it hands off to. An agent can identify the tool's role without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicit workflow routing: poll gvt_get_batch_status until completed/failed, then use gvt_get_batch_summary or gvt_get_test_results. It also carves out the when-not case by directing interpretation/methodology needs to gvt_get_knowledge. Alternatives and conditions are named rather than inferred.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
gvt_run_visibility_testRun a visibility testAInspect
Run a GEO visibility test on a URL. Analyzes 7 categories: JavaScript dependencies (js), no-JS rendering (nojs), semantic HTML5, heading structure, schema.org markup, social tags, and accessibility. This is an async operation — it returns a tid immediately. Pass that tid to gvt_get_test_results to retrieve scores and issues once the test completes (poll until status is not "pending" or "running").
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | The URL to analyze | |
| waitForJs | No | Wait for JavaScript rendering before analysis. Set to false to skip JS rendering (faster, but may miss dynamically injected content). |
Output Schema
| Name | Required | Description |
|---|---|---|
| tid | No | 8-character test ID to poll with |
| next | No | The suggested next call |
| status | No | Always pending at queue time |
| message | No | Human-readable acknowledgement with polling guidance |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Pipelines the async nature explicitly: it says 'returns a tid immediately' and that results are retrieved separately by passing tid to gvt_get_test_results, with polling instructions. This behavior isn't implied by annotations (readOnlyHint=false, no idempotentHint), so the description carries the burden and does so thoroughly. No contradiction with annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Concise but information-dense. Two sentences: first defines the tool's function and scope, second addresses the async behavior and next steps. No filler; every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given an output schema exists (though not detailed here), the description need not explain return values. It covers the core: what the tool does (analyzes 7 categories), the async contract, and how to proceed (poll until status changes). For a tool of moderate complexity with a well-covered schema, this is complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, both parameters (url and waitForJs) already have clear descriptions in the schema. The description adds marginal value beyond the schema—it doesn't mention parameters at all, but the schema is sufficient. Baseline 3 applies because the description adds no extra parameter context but doesn't need to.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States the specific action (run a GEO visibility test on a URL), lists the 7 categories it analyzes (js, nojs, semantic HTML5, heading structure, schema.org markup, social tags, accessibility), and differentiates itself from siblings like gvt_run_visibility_batch (batch) and gvt_schedule_test (scheduled). The purpose is unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly instructs when to use this tool: to run a single visibility test on a URLichelses immediate asynchronous result. It contrasts with gvt_get_test_results (retrieve results by tid) and implies not for batch (run_visibility_batch). It provides clear workflow guidance: pass the tid to gvt_get_test_results and poll until status changes.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
gvt_schedule_testManage schedulesADestructiveInspect
Manage recurring GVT test schedules: add, delete, reset, pause, resume, cancel, and change_frequency. Discover existing schedules and their schids via gvt_list_schedules; control modes require a schid. For one-off tests use gvt_run_visibility_test; for queued batches without recurrence use gvt_run_visibility_batch. Modes: "add" upserts schedules for the provided urls without duplicating; "delete" permanently removes schedule entries by schid or url (test results are never deleted by any mode); "reset" is destructive — deletes all matching schedules first (domain-scoped when provided, otherwise ALL account-wide), then upserts the provided urls; "pause" halts reversibly; "resume" restarts a paused schedule and clears auto-pause reasons and failure counts; "cancel" permanently ends the schedule (schid unusable); "change_frequency" updates the recurrence interval. Management modes (add/delete/reset) require urls; control modes require schid. Subscription-tier enforcement applies.
| Name | Required | Description | Default |
|---|---|---|---|
| mode | No | Schedule operation mode. add/delete/reset operate on URLs (require urls array). pause/resume/cancel/change_frequency operate on a single schedule (require schid). | add |
| urls | No | Array of URL/schedule objects. Required for add, delete, and reset modes. | |
| schid | No | Schedule ID. Required for pause, resume, cancel, and change_frequency modes. | |
| domain | No | Domain to scope a reset operation (reset mode only). e.g. "example.com". Matches https://, http://, with/without www., with/without trailing path. Omit to reset ALL schedules. | |
| frequency | No | New frequency for the schedule. Required for change_frequency mode. |
Output Schema
| Name | Required | Description |
|---|---|---|
| success | No | True when the requested mode completed |
| scheduled | No | Schedules now in effect (add/reset modes) |
| deletedByUrl | No | Per-URL breakdown of deleted schedules (delete mode). Present on delete mode responses, including schid-based deletes; omitted for reset mode. |
| deletedCount | No | Schedule entries removed (delete/reset modes) |
| upsertedCount | No | Schedules created or updated (add/reset modes) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description goes well beyond the annotations by detailing consequences: delete permanently removes entries but never deletes test results; reset is destructive and deletes all matching schedules before upserting; pause is reversible; resume clears auto-pause reasons and failure counts; cancel makes the schid unusable. The destructiveHint annotation is confirmed and enriched with mode-specific scope and side effects.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense but every sentence earns its place: it starts with the primary purpose, then immediately routes to sibling tools, then defines each mode with its behavioral consequence. The mode definitions are compact yet complete, and the requirement summary sentence is a useful checklist.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's multi-mode complexity and zero required fields in the schema, the description fully compensates by specifying mode-specific argument requirements, destructive behavior, domain scoping, and the relationship to schedule discovery. The presence of an output schema means return-value details are already covered elsewhere, so nothing essential is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Even though the schema covers all parameters, the description adds crucial operational meaning: 'add' upserts without duplicating, 'delete' can target schid or url, 'reset' is domain-scoped or account-wide, and 'change_frequency' requires a new frequency. This clarifies how the optional-looking schema fields (all non-required) actually combine per mode, which the schema alone does not convey.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'Manage recurring GVT test schedules,' then enumerates the seven supported modes (add, delete, reset, pause, resume, cancel, change_frequency). It differentiates from sibling tools by explicitly naming gvt_list_schedules, gvt_run_visibility_test, and gvt_run_visibility_batch as alternatives for discovery and non-recurring runs.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives explicit routing guidance: use gvt_list_schedules to discover schids, use gvt_run_visibility_test for one-off tests, and use gvt_run_visibility_batch for queued non-recurring batches. It also states mode-specific requirements, such as 'Management modes (add/delete/reset) require urls; control modes require schid,' so an agent knows exactly when to invoke this tool versus siblings.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
gvt_set_test_visibilitySet test visibilityAIdempotentInspect
Toggle the public visibility of a test session. Set is_public to true to make the test shareable and receive its public share URL, or false to make it private. The response confirms the new state and, when making a test public, includes the share_url the recipient can visit without any authentication. is_public is the only lever: true both publishes and yields the link in this same response; false revokes access to any previously issued share_url immediately. This is the only way to control who can view a test result without requiring MCP authentication.
| Name | Required | Description | Default |
|---|---|---|---|
| tid | Yes | The 8-character test ID returned by gvt_run_visibility_test. | |
| is_public | Yes | true to make the test publicly shareable, false to make it private. |
Output Schema
| Name | Required | Description |
|---|---|---|
| tid | No | The test that was modified |
| success | No | True when the visibility change was applied |
| is_public | No | The new public visibility state |
| share_url | No | Public share URL, present when is_public is true; the recipient needs no authentication. Null when sharing failed or the test is private |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description goes well beyond the annotations, detailing response behavior (confirmation, share_url), authentication-free recipient access, immediate revocation of prior share URLs, and that is_public is the only control lever. This aligns with idempotentHint and destructiveHint and adds meaningful non-obvious behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with the core action, and every subsequent sentence contributes either behavioral nuance or usage context. There is no filler or repetition of what annotations already provide.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With only two required parameters, a complete schema, and an output schema, the description covers all operational essentials: state change, response contents, share URL behavior, no-auth access, and revocation. No critical information is missing for correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3. The description adds extra meaning by explaining the consequences of true vs false (publishing and returning share_url vs revoking access) and reinforcing that is_public is the only lever. tid's source is already fully documented in the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific action and resource: 'Toggle the public visibility of a test session.' It then explains the two states and their outcomes, and closes by positioning this as the only way to control who can view a test result, distinguishing it from the sibling list.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It clearly states when to use the tool: when you want to share or privatize a test session ('This is the only way to control who can view a test result...'). It does not explicitly name sibling alternatives, but the 'only way' framing gives a strong selection signal.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
2 tool updates
- Added
gvt_get_batch_summary - Changed
gvt_get_test_results6 fields changed- added
Input schema / properties / detailAdded value: +{ + "default": "full", + "description": "Payload shape. summary: analysisSummary plus the top issues per category (deduplicated by type, capped at 10, sorted by severity then weight) — no semantic tree, per-element markup, or nojs_rendered pass. full: complete resultData for all analysis types (previous behavior, the default).", + "enum": [ + "summary", + "full" + ], + "type": "string" +} - changed
Output schema / descriptionPrevious value: -"Full results for one test. While status is pending or running, results are incomplete — wait and retry."New value: +"Results for one test. While status is pending or running, results are incomplete — wait and retry. detail=summary returns the compact shape; detail=full (the default) returns complete per-category resultData." - added
Output schema / properties / detailAdded value: +{ + "description": "Present with value \"summary\" when detail=summary was requested; omitted from full responses so existing callers see no change", + "enum": [ + "summary" + ], + "type": "string" +} - changed
Output schema / properties / results / descriptionPrevious value: -"Per-category result records. Each carries analysisType and resultData; issues carry fixId for gvt_get_fix."New value: +"Per-category result records. detail=full: each carries analysisType and complete resultData (semantic tree, per-element issues); issues carry fixId for gvt_get_fix. detail=summary: each carries analysisType, score, and the top issues deduplicated by type (capped at 10 per category); the nojs_rendered duplicate pass is omitted — analysisSummary.categoryIssueCounts keeps the true totals." - added
Output schema / properties / results / items / properties / issuesAdded value: +{ + "description": "Issue records; in summary mode each carries type, severity, impact, weight, element, description, and fixId (remediation comes from gvt_get_fix, not the payload)", + "type": "array" +} - added
Output schema / properties / results / items / properties / scoreAdded value: +{ + "description": "Category score (0-100), present in summary mode when the category reports one", + "type": "number" +}
1 tool update
- Added
gvt_list_schedules
9 tool updates
- Added
gvt_get_baseline - Removed
gvt_get_oldest - Removed
gvt_get_oldest_list - Removed
gvt_get_sitemap_urls - Added
gvt_list_baselines - Added
gvt_list_sitemap_urls - Changed
gvt_run_visibility_batch1 field changed- changed
Output schema / descriptionPrevious value: -"Batch acknowledgement. Poll gvt_get_test_results per URL once the batch completes."New value: +"Batch acknowledgement. Poll gvt_get_batch_status with the batchId until status is completed or failed, then retrieve results via gvt_get_test_results."
- Changed
gvt_set_test_visibility1 field changed- added
Output schema / properties / share_urlAdded value: +{ + "description": "Public share URL, present when is_public is true; the recipient needs no authentication. Null when sharing failed or the test is private", + "type": [ + "string", + "null" + ] +}
- Removed
gvt_share_test
1 tool update
- Added
gvt_get_batch_status
6 tool updates
- Changed
gvt_get_latest19 fields changed- changed
Output schema / properties / analysisSummary / descriptionPrevious value: -"Aggregate scores for one test session"New value: +"Aggregate scores for this test" - changed
Output schema / properties / analysisSummary / properties / categoryIssueCounts / properties / accessibility / descriptionPrevious value: -"Accessibility score (0-100)"New value: +"Number of open accessibility issues" - changed
Output schema / properties / analysisSummary / properties / categoryIssueCounts / properties / accessibility / typePrevious value: -"number"New value: +"integer" - changed
Output schema / properties / analysisSummary / properties / categoryIssueCounts / properties / headings / descriptionPrevious value: -"Heading structure score (0-100)"New value: +"Number of open heading structure issues" - changed
Output schema / properties / analysisSummary / properties / categoryIssueCounts / properties / headings / typePrevious value: -"number"New value: +"integer" - changed
Output schema / properties / analysisSummary / properties / categoryIssueCounts / properties / javascript / descriptionPrevious value: -"JavaScript dependency score (0-100)"New value: +"Number of open JavaScript dependency issues" - changed
Output schema / properties / analysisSummary / properties / categoryIssueCounts / properties / javascript / typePrevious value: -"number"New value: +"integer" - changed
Output schema / properties / analysisSummary / properties / categoryIssueCounts / properties / nojs / descriptionPrevious value: -"No-JS rendering score (0-100)"New value: +"Number of open No-JS rendering issues" - changed
Output schema / properties / analysisSummary / properties / categoryIssueCounts / properties / nojs / typePrevious value: -"number"New value: +"integer" - changed
Output schema / properties / analysisSummary / properties / categoryIssueCounts / properties / schema / descriptionPrevious value: -"schema.org markup score (0-100)"New value: +"Number of open schema.org markup issues" - changed
Output schema / properties / analysisSummary / properties / categoryIssueCounts / properties / schema / typePrevious value: -"number"New value: +"integer" - changed
Output schema / properties / analysisSummary / properties / categoryIssueCounts / properties / semantic / descriptionPrevious value: -"Semantic HTML5 score (0-100)"New value: +"Number of open semantic HTML5 issues" - changed
Output schema / properties / analysisSummary / properties / categoryIssueCounts / properties / semantic / typePrevious value: -"number"New value: +"integer" - changed
Output schema / properties / analysisSummary / properties / categoryIssueCounts / properties / social / descriptionPrevious value: -"Social tags score (0-100)"New value: +"Number of open social tag issues" - changed
Output schema / properties / analysisSummary / properties / categoryIssueCounts / properties / social / typePrevious value: -"number"New value: +"integer" - added
Output schema / properties / expiresAtAdded value: +{ + "description": "ISO 8601 instant this snapshot will expire", + "type": [ + "string", + "null" + ] +} - added
Output schema / properties / isPublicAdded value: +{ + "description": "Whether this test is publicly visible", + "type": "boolean" +} - added
Output schema / properties / testTypeAdded value: +{ + "description": "The type of the test run", + "enum": [ + "url", + "html_paste" + ], + "type": "string" +} - added
Output schema / properties / updatedAtAdded value: +{ + "description": "ISO 8601 instant this snapshot was last updated", + "type": "string" +}
- Changed
gvt_get_oldest19 fields changed- changed
Output schema / properties / analysisSummary / descriptionPrevious value: -"Aggregate scores for one test session"New value: +"Aggregate scores for this test" - changed
Output schema / properties / analysisSummary / properties / categoryIssueCounts / properties / accessibility / descriptionPrevious value: -"Accessibility score (0-100)"New value: +"Number of open accessibility issues" - changed
Output schema / properties / analysisSummary / properties / categoryIssueCounts / properties / accessibility / typePrevious value: -"number"New value: +"integer" - changed
Output schema / properties / analysisSummary / properties / categoryIssueCounts / properties / headings / descriptionPrevious value: -"Heading structure score (0-100)"New value: +"Number of open heading structure issues" - changed
Output schema / properties / analysisSummary / properties / categoryIssueCounts / properties / headings / typePrevious value: -"number"New value: +"integer" - changed
Output schema / properties / analysisSummary / properties / categoryIssueCounts / properties / javascript / descriptionPrevious value: -"JavaScript dependency score (0-100)"New value: +"Number of open JavaScript dependency issues" - changed
Output schema / properties / analysisSummary / properties / categoryIssueCounts / properties / javascript / typePrevious value: -"number"New value: +"integer" - changed
Output schema / properties / analysisSummary / properties / categoryIssueCounts / properties / nojs / descriptionPrevious value: -"No-JS rendering score (0-100)"New value: +"Number of open No-JS rendering issues" - changed
Output schema / properties / analysisSummary / properties / categoryIssueCounts / properties / nojs / typePrevious value: -"number"New value: +"integer" - changed
Output schema / properties / analysisSummary / properties / categoryIssueCounts / properties / schema / descriptionPrevious value: -"schema.org markup score (0-100)"New value: +"Number of open schema.org markup issues" - changed
Output schema / properties / analysisSummary / properties / categoryIssueCounts / properties / schema / typePrevious value: -"number"New value: +"integer" - changed
Output schema / properties / analysisSummary / properties / categoryIssueCounts / properties / semantic / descriptionPrevious value: -"Semantic HTML5 score (0-100)"New value: +"Number of open semantic HTML5 issues" - changed
Output schema / properties / analysisSummary / properties / categoryIssueCounts / properties / semantic / typePrevious value: -"number"New value: +"integer" - changed
Output schema / properties / analysisSummary / properties / categoryIssueCounts / properties / social / descriptionPrevious value: -"Social tags score (0-100)"New value: +"Number of open social tag issues" - changed
Output schema / properties / analysisSummary / properties / categoryIssueCounts / properties / social / typePrevious value: -"number"New value: +"integer" - added
Output schema / properties / expiresAtAdded value: +{ + "description": "ISO 8601 instant this snapshot will expire", + "type": [ + "string", + "null" + ] +} - added
Output schema / properties / isPublicAdded value: +{ + "description": "Whether this test is publicly visible", + "type": "boolean" +} - added
Output schema / properties / testTypeAdded value: +{ + "description": "The type of the test run", + "enum": [ + "url", + "html_paste" + ], + "type": "string" +} - added
Output schema / properties / updatedAtAdded value: +{ + "description": "ISO 8601 instant this snapshot was last updated", + "type": "string" +}
- Changed
gvt_get_score_trend4 fields changed- added
Output schema / properties / compareAdded value: +{ + "description": "Comparison window settings used", + "type": [ + "object", + "null" + ] +} - added
Output schema / properties / includeSubdomainsAdded value: +{ + "description": "Whether subdomains were included in the domain search", + "type": "boolean" +} - added
Output schema / properties / truncatedAdded value: +{ + "deprecated": true, + "description": "Deprecated. Use `hasMore` instead.", + "type": "boolean" +} - added
Output schema / properties / windowSummaryAdded value: +{ + "description": "Summary of window comparison results", + "type": [ + "object", + "null" + ] +}
- Changed
gvt_get_sitemap_urls3 fields changed- changed
Output schema / properties / bulkUrlLimit / descriptionPrevious value: -"How many of these URLs the caller tier permits in one batch call"New value: +"How many of these URLs the caller's tier permits in one batch-analyze call." - changed
Output schema / properties / legalPagesFound / descriptionPrevious value: -"Legal/policy pages detected before filtering"New value: +"How many of the URLs discovered across all sitemaps were detected as legal/policy boilerplate, counted BEFORE any filtering or limiting is applied." - changed
Output schema / properties / truncated / descriptionPrevious value: -"True when the sitemap parser hit its 5000-URL cap"New value: +"True if the underlying sitemap parser hit its 5000-URL cap during fetching."
- Changed
gvt_get_test_results19 fields changed- changed
Output schema / properties / analysisSummary / properties / categoryIssueCounts / properties / accessibility / descriptionPrevious value: -"Accessibility score (0-100)"New value: +"Number of open accessibility issues" - changed
Output schema / properties / analysisSummary / properties / categoryIssueCounts / properties / accessibility / typePrevious value: -"number"New value: +"integer" - changed
Output schema / properties / analysisSummary / properties / categoryIssueCounts / properties / headings / descriptionPrevious value: -"Heading structure score (0-100)"New value: +"Number of open heading structure issues" - changed
Output schema / properties / analysisSummary / properties / categoryIssueCounts / properties / headings / typePrevious value: -"number"New value: +"integer" - changed
Output schema / properties / analysisSummary / properties / categoryIssueCounts / properties / javascript / descriptionPrevious value: -"JavaScript dependency score (0-100)"New value: +"Number of open JavaScript dependency issues" - changed
Output schema / properties / analysisSummary / properties / categoryIssueCounts / properties / javascript / typePrevious value: -"number"New value: +"integer" - changed
Output schema / properties / analysisSummary / properties / categoryIssueCounts / properties / nojs / descriptionPrevious value: -"No-JS rendering score (0-100)"New value: +"Number of open No-JS rendering issues" - changed
Output schema / properties / analysisSummary / properties / categoryIssueCounts / properties / nojs / typePrevious value: -"number"New value: +"integer" - changed
Output schema / properties / analysisSummary / properties / categoryIssueCounts / properties / schema / descriptionPrevious value: -"schema.org markup score (0-100)"New value: +"Number of open schema.org markup issues" - changed
Output schema / properties / analysisSummary / properties / categoryIssueCounts / properties / schema / typePrevious value: -"number"New value: +"integer" - changed
Output schema / properties / analysisSummary / properties / categoryIssueCounts / properties / semantic / descriptionPrevious value: -"Semantic HTML5 score (0-100)"New value: +"Number of open semantic HTML5 issues" - changed
Output schema / properties / analysisSummary / properties / categoryIssueCounts / properties / semantic / typePrevious value: -"number"New value: +"integer" - changed
Output schema / properties / analysisSummary / properties / categoryIssueCounts / properties / social / descriptionPrevious value: -"Social tags score (0-100)"New value: +"Number of open social tag issues" - changed
Output schema / properties / analysisSummary / properties / categoryIssueCounts / properties / social / typePrevious value: -"number"New value: +"integer" - added
Output schema / properties / expiresAtAdded value: +{ + "description": "ISO 8601 instant when this test session and its results will expire and be deleted", + "type": [ + "string", + "null" + ] +} - added
Output schema / properties / isPublicAdded value: +{ + "description": "Whether this test result is publicly accessible", + "type": [ + "boolean", + "null" + ] +} - changed
Output schema / properties / shareableTid / descriptionPrevious value: -"Non-null when the result is publicly shareable"New value: +"Non-null when the result is publicly shareable. Public share identifier for this session. Equal to tid when the session is shareable; null when the session is private. Presence of a value is the indicator that share links may be generated." - added
Output schema / properties / testTypeAdded value: +{ + "description": "Type of test performed (e.g., URL or direct HTML paste)", + "enum": [ + "url", + "html_paste" + ], + "type": "string" +} - added
Output schema / properties / updatedAtAdded value: +{ + "description": "ISO 8601 instant of the last update to this test session", + "type": "string" +}
- Changed
gvt_schedule_test1 field changed- changed
Output schema / properties / deletedByUrl / descriptionPrevious value: -"Per-URL breakdown of deleted schedules (delete mode)"New value: +"Per-URL breakdown of deleted schedules (delete mode). Present on delete mode responses, including schid-based deletes; omitted for reset mode."
17 tool updates
- Changed
gvt_delete_test2 fields changed- added
Output schema / properties / message / descriptionAdded value: +"Human-readable confirmation" - added
Output schema / properties / success / descriptionAdded value: +"True when the test was deleted"
- Changed
gvt_get_fix9 fields changed- added
Output schema / properties / category / descriptionAdded value: +"Analysis category the issue belongs to" - added
Output schema / properties / human_description / descriptionAdded value: +"Human-readable description of the issue" - added
Output schema / properties / impact / descriptionAdded value: +"Plain-language statement of what the issue costs" - added
Output schema / properties / issue_id / descriptionAdded value: +"Stable issue identifier, e.g. og_title_missing" - added
Output schema / properties / recommendation / descriptionAdded value: +"The recommended remediation" - added
Output schema / properties / resourceUri / descriptionAdded value: +"The gvt:// resource this record was read from" - added
Output schema / properties / severity_range / descriptionAdded value: +"Severity levels this issue can take" - added
Output schema / properties / subcategory / descriptionAdded value: +"Finer-grained grouping within the category" - added
Output schema / properties / ttlMs / descriptionAdded value: +"Suggested client cache lifetime in milliseconds"
- Changed
gvt_get_knowledge2 fields changed- added
Output schema / properties / topic / descriptionAdded value: +"The topic that was served" - added
Output schema / properties / ttlMs / descriptionAdded value: +"Suggested client cache lifetime in milliseconds"
- Changed
gvt_get_latest10 fields changed- added
Output schema / properties / analysisSummary / properties / categoryIssueCounts / descriptionAdded value: +"Open issue count per analysis category" - added
Output schema / properties / analysisSummary / properties / categoryScores / descriptionAdded value: +"Score per analysis category (0-100)" - added
Output schema / properties / analysisSummary / properties / jsDependencySource / descriptionAdded value: +"How the JS dependency score was derived" - added
Output schema / properties / createdAt / descriptionAdded value: +"ISO 8601 instant this snapshot was created" - added
Output schema / properties / findingSentence / descriptionAdded value: +"Pre-generated natural language verdict for this snapshot" - added
Output schema / properties / message / descriptionAdded value: +"Present only in the no-test case" - added
Output schema / properties / shareableTid / descriptionAdded value: +"Public share ID, non-null when shareable" - added
Output schema / properties / status / descriptionAdded value: +"Snapshot status (completed, failed, etc.)" - added
Output schema / properties / tid / descriptionAdded value: +"Test ID of the snapshot" - added
Output schema / properties / url / descriptionAdded value: +"The URL this snapshot analyzed"
- Changed
gvt_get_oldest10 fields changed- added
Output schema / properties / analysisSummary / properties / categoryIssueCounts / descriptionAdded value: +"Open issue count per analysis category" - added
Output schema / properties / analysisSummary / properties / categoryScores / descriptionAdded value: +"Score per analysis category (0-100)" - added
Output schema / properties / analysisSummary / properties / jsDependencySource / descriptionAdded value: +"How the JS dependency score was derived" - added
Output schema / properties / createdAt / descriptionAdded value: +"ISO 8601 instant this snapshot was created" - added
Output schema / properties / findingSentence / descriptionAdded value: +"Pre-generated natural language verdict for this snapshot" - added
Output schema / properties / message / descriptionAdded value: +"Present only in the no-test case" - added
Output schema / properties / shareableTid / descriptionAdded value: +"Public share ID, non-null when shareable" - added
Output schema / properties / status / descriptionAdded value: +"Snapshot status (completed, failed, etc.)" - added
Output schema / properties / tid / descriptionAdded value: +"Test ID of the snapshot" - added
Output schema / properties / url / descriptionAdded value: +"The URL this snapshot analyzed"
- Changed
gvt_get_oldest_list9 fields changed- added
Output schema / properties / found / descriptionAdded value: +"URLs that had at least one test" - added
Output schema / properties / requested / descriptionAdded value: +"Number of URLs requested" - added
Output schema / properties / results / descriptionAdded value: +"One entry per requested URL, in request order" - added
Output schema / properties / results / items / properties / oldest / properties / categoryScores / descriptionAdded value: +"Score per analysis category (0-100)" - added
Output schema / properties / results / items / properties / oldest / properties / createdAt / descriptionAdded value: +"ISO 8601 instant of this snapshot" - added
Output schema / properties / results / items / properties / oldest / properties / overallScore / descriptionAdded value: +"Overall GEO visibility score (0-100), null when the test failed" - added
Output schema / properties / results / items / properties / oldest / properties / tid / descriptionAdded value: +"Test ID of this snapshot" - added
Output schema / properties / results / items / properties / url / descriptionAdded value: +"The URL this record describes" - added
Output schema / properties / ttlMs / descriptionAdded value: +"Suggested client cache lifetime in milliseconds"
- Changed
gvt_get_prompt2 fields changed- added
Output schema / properties / description / descriptionAdded value: +"The prompt catalog description" - added
Output schema / properties / name / descriptionAdded value: +"The prompt name that was rendered"
- Changed
gvt_get_score_trend23 fields changed- added
Output schema / properties / domain / descriptionAdded value: +"The compared domain in domain mode, else null" - added
Output schema / properties / hasMore / descriptionAdded value: +"True when more pages exist" - added
Output schema / properties / limit / descriptionAdded value: +"Max records per page in domain mode" - added
Output schema / properties / nextOffset / descriptionAdded value: +"Offset for the next page, null on the last page" - added
Output schema / properties / offset / descriptionAdded value: +"Row offset applied to this page" - added
Output schema / properties / returned / descriptionAdded value: +"Records in this page" - added
Output schema / properties / scope / descriptionAdded value: +"Which mode the comparison ran in" - added
Output schema / properties / trends / descriptionAdded value: +"One row per matched URL, baseline vs latest with deltas" - added
Output schema / properties / trends / items / properties / baseline / properties / categoryScores / descriptionAdded value: +"Score per analysis category (0-100)" - added
Output schema / properties / trends / items / properties / baseline / properties / createdAt / descriptionAdded value: +"ISO 8601 instant of this snapshot" - added
Output schema / properties / trends / items / properties / baseline / properties / overallScore / descriptionAdded value: +"Overall GEO visibility score (0-100), null when the test failed" - added
Output schema / properties / trends / items / properties / baseline / properties / tid / descriptionAdded value: +"Test ID of this snapshot" - added
Output schema / properties / trends / items / properties / deltas / properties / categories / descriptionAdded value: +"Per-category score change, latest minus baseline" - added
Output schema / properties / trends / items / properties / deltas / properties / overall / descriptionAdded value: +"Overall score change, latest minus baseline" - added
Output schema / properties / trends / items / properties / latest / properties / categoryScores / descriptionAdded value: +"Score per analysis category (0-100)" - added
Output schema / properties / trends / items / properties / latest / properties / createdAt / descriptionAdded value: +"ISO 8601 instant of this snapshot" - added
Output schema / properties / trends / items / properties / latest / properties / overallScore / descriptionAdded value: +"Overall GEO visibility score (0-100), null when the test failed" - added
Output schema / properties / trends / items / properties / latest / properties / tid / descriptionAdded value: +"Test ID of this snapshot" - added
Output schema / properties / trends / items / properties / testCount / descriptionAdded value: +"Total tests ever run for this URL" - added
Output schema / properties / trends / items / properties / url / descriptionAdded value: +"The URL this trend row describes" - added
Output schema / properties / ttlMs / descriptionAdded value: +"Suggested client cache lifetime in milliseconds" - added
Output schema / properties / url / descriptionAdded value: +"The compared URL in url mode, else null" - added
Output schema / properties / urlCount / descriptionAdded value: +"Total matched URLs before pagination"
- Changed
gvt_get_sitemap_urls23 fields changed- added
Output schema / properties / errors / descriptionAdded value: +"Per-sitemap fetch or parse failures; non-empty means discovery was partial" - added
Output schema / properties / hasMore / descriptionAdded value: +"True when the filtered set was cut short by the limit" - added
Output schema / properties / legalPages / descriptionAdded value: +"The legal-page filter actually applied" - added
Output schema / properties / limit / descriptionAdded value: +"Effective result cap" - added
Output schema / properties / matched / descriptionAdded value: +"URLs remaining after include/exclude and legalPages filters" - added
Output schema / properties / requestedUrl / descriptionAdded value: +"The URL or sitemap URL that was resolved" - added
Output schema / properties / resolvedSitemaps / descriptionAdded value: +"Sitemap URLs actually parsed" - added
Output schema / properties / returned / descriptionAdded value: +"URLs actually returned after the limit" - added
Output schema / properties / robotsTxtFound / descriptionAdded value: +"Whether a robots.txt was found and consulted" - added
Output schema / properties / sitemapCount / descriptionAdded value: +"Number of sitemaps parsed" - added
Output schema / properties / sort / descriptionAdded value: +"The sort actually applied to the results" - added
Output schema / properties / totalDiscovered / descriptionAdded value: +"Raw URL count discovered before filtering" - added
Output schema / properties / urls / descriptionAdded value: +"Discovered URLs, ranked per the sort setting" - added
Output schema / properties / urls / items / properties / alreadyScheduled / descriptionAdded value: +"True when this URL is already in the testing queue" - added
Output schema / properties / urls / items / properties / alreadyTested / descriptionAdded value: +"True when the caller has an existing test for this URL" - added
Output schema / properties / urls / items / properties / changefreq / descriptionAdded value: +"Sitemap change frequency hint, when present" - added
Output schema / properties / urls / items / properties / disallowedByRobots / descriptionAdded value: +"True when robots.txt disallows crawling this URL" - added
Output schema / properties / urls / items / properties / lastmod / descriptionAdded value: +"Last-modified timestamp from the sitemap, when present" - added
Output schema / properties / urls / items / properties / legalPage / descriptionAdded value: +"True when detected as legal/policy boilerplate" - added
Output schema / properties / urls / items / properties / pathDepth / descriptionAdded value: +"Path depth from the site root (1 = homepage)" - added
Output schema / properties / urls / items / properties / priority / descriptionAdded value: +"Sitemap priority hint, when present" - added
Output schema / properties / urls / items / properties / significanceScore / descriptionAdded value: +"SEO significance score used for the default ranking" - added
Output schema / properties / urls / items / properties / url / descriptionAdded value: +"The discovered URL"
- Changed
gvt_get_test_results8 fields changed- added
Output schema / properties / analysisSummary / properties / categoryIssueCounts / descriptionAdded value: +"Open issue count per analysis category" - added
Output schema / properties / analysisSummary / properties / categoryScores / descriptionAdded value: +"Score per analysis category (0-100)" - added
Output schema / properties / analysisSummary / properties / jsDependencySource / descriptionAdded value: +"How the JS dependency score was derived" - added
Output schema / properties / createdAt / descriptionAdded value: +"ISO 8601 instant of test creation" - added
Output schema / properties / results / items / properties / analysisType / descriptionAdded value: +"Which analysis category this record holds" - added
Output schema / properties / status / descriptionAdded value: +"Pending and running mean poll again; completed means full results are present" - added
Output schema / properties / tid / descriptionAdded value: +"The test ID that was polled" - added
Output schema / properties / url / descriptionAdded value: +"The URL that was analyzed"
- Changed
gvt_list_prompts10 fields changed- added
Output schema / properties / promptCount / descriptionAdded value: +"Number of prompt workflows in the catalog" - added
Output schema / properties / prompts / descriptionAdded value: +"The prompt workflow catalog" - added
Output schema / properties / prompts / items / properties / arguments / descriptionAdded value: +"Arguments the prompt accepts" - added
Output schema / properties / prompts / items / properties / arguments / items / properties / description / descriptionAdded value: +"What the argument controls" - added
Output schema / properties / prompts / items / properties / arguments / items / properties / name / descriptionAdded value: +"Argument name" - added
Output schema / properties / prompts / items / properties / arguments / items / properties / required / descriptionAdded value: +"True when the argument must be supplied" - added
Output schema / properties / prompts / items / properties / description / descriptionAdded value: +"What the prompt workflow does" - added
Output schema / properties / prompts / items / properties / name / descriptionAdded value: +"Prompt name for gvt_get_prompt" - added
Output schema / properties / prompts / items / properties / title / descriptionAdded value: +"Human-readable display name" - added
Output schema / properties / ttlMs / descriptionAdded value: +"Suggested client cache lifetime in milliseconds"
- Changed
gvt_list_tests11 fields changed- added
Output schema / properties / createdFrom / descriptionAdded value: +"Effective lower creation-date bound, if any" - added
Output schema / properties / createdTo / descriptionAdded value: +"Effective upper creation-date bound, if any" - added
Output schema / properties / domain / descriptionAdded value: +"The domain filter actually applied, if any" - added
Output schema / properties / hasMore / descriptionAdded value: +"True when more pages exist" - added
Output schema / properties / includeSubdomains / descriptionAdded value: +"Whether subdomains were included in the domain filter" - added
Output schema / properties / limit / descriptionAdded value: +"Page size actually applied" - added
Output schema / properties / nextOffset / descriptionAdded value: +"Offset for the next page, null on the last page" - added
Output schema / properties / offset / descriptionAdded value: +"Row offset applied to this page" - added
Output schema / properties / tests / descriptionAdded value: +"The matching test sessions on this page" - added
Output schema / properties / total / descriptionAdded value: +"Total matching tests before pagination" - added
Output schema / properties / ttlMs / descriptionAdded value: +"Suggested client cache lifetime in milliseconds"
- Changed
gvt_run_visibility_batch4 fields changed- added
Output schema / properties / status / descriptionAdded value: +"Always queued at submission time" - added
Output schema / properties / totalUrls / descriptionAdded value: +"URLs accepted into the batch" - added
Output schema / properties / webhookConfigured / descriptionAdded value: +"True when a completion webhook URL was supplied" - added
Output schema / properties / webhookSecret / descriptionAdded value: +"HMAC secret for verifying the webhook, null when no webhook"
- Changed
gvt_run_visibility_test5 fields changed- added
Output schema / properties / message / descriptionAdded value: +"Human-readable acknowledgement with polling guidance" - added
Output schema / properties / next / properties / arguments / descriptionAdded value: +"Arguments for the next call" - added
Output schema / properties / next / properties / arguments / properties / tid / descriptionAdded value: +"The tid to poll with" - added
Output schema / properties / next / properties / tool / descriptionAdded value: +"The tool to call next" - added
Output schema / properties / status / descriptionAdded value: +"Always pending at queue time"
- Changed
gvt_schedule_test11 fields changed- added
Output schema / properties / deletedByUrl / descriptionAdded value: +"Per-URL breakdown of deleted schedules (delete mode)" - added
Output schema / properties / deletedByUrl / items / properties / count / descriptionAdded value: +"Number of schedule entries deleted for this URL" - added
Output schema / properties / deletedByUrl / items / properties / frequencies / descriptionAdded value: +"Frequencies that were removed for this URL" - added
Output schema / properties / deletedByUrl / items / properties / url / descriptionAdded value: +"The URL whose schedules were deleted" - added
Output schema / properties / deletedCount / descriptionAdded value: +"Schedule entries removed (delete/reset modes)" - added
Output schema / properties / scheduled / descriptionAdded value: +"Schedules now in effect (add/reset modes)" - added
Output schema / properties / scheduled / items / properties / schid / descriptionAdded value: +"Schedule ID for later pause/resume/cancel calls" - added
Output schema / properties / scheduled / items / properties / url / descriptionAdded value: +"The URL placed on schedule" - added
Output schema / properties / scheduled / items / properties / wpPostId / descriptionAdded value: +"Associated WordPress post ID, when provided" - added
Output schema / properties / success / descriptionAdded value: +"True when the requested mode completed" - added
Output schema / properties / upsertedCount / descriptionAdded value: +"Schedules created or updated (add/reset modes)"
- Changed
gvt_set_test_visibility3 fields changed- added
Output schema / properties / is_public / descriptionAdded value: +"The new public visibility state" - added
Output schema / properties / success / descriptionAdded value: +"True when the visibility change was applied" - added
Output schema / properties / tid / descriptionAdded value: +"The test that was modified"
- Changed
gvt_share_test3 fields changed- added
Output schema / properties / is_public / descriptionAdded value: +"The test public visibility state after the call" - added
Output schema / properties / success / descriptionAdded value: +"True when sharing was enabled" - added
Output schema / properties / tid / descriptionAdded value: +"The test that was shared"
17 tool updates
- Changed
gvt_delete_test1 field changed- changed
Output schema / (root)Previous value: -nullNew value: +{ + "properties": { + "message": { + "type": "string" + }, + "success": { + "type": "boolean" + } + }, + "type": "object" +}
- Changed
gvt_get_fix1 field changed- changed
Output schema / (root)Previous value: -nullNew value: +{ + "description": "Canonical fix record for one issue type", + "properties": { + "category": { + "type": "string" + }, + "human_description": { + "type": "string" + }, + "impact": { + "type": "string" + }, + "issue_id": { + "type": "string" + }, + "recommendation": { + "type": "string" + }, + "resourceUri": { + "type": "string" + }, + "severity_range": { + "items": { + "type": "string" + }, + "type": "array" + }, + "subcategory": { + "type": "string" + }, + "ttlMs": { + "type": "integer" + }, + "weight": { + "description": "Scoring weight", + "type": "number" + } + }, + "type": "object" +}
- Changed
gvt_get_knowledge1 field changed- changed
Output schema / (root)Previous value: -nullNew value: +{ + "properties": { + "knowledge": { + "description": "Reference content for the requested topic", + "type": "object" + }, + "topic": { + "enum": [ + "all", + "methodology", + "scoring", + "categories", + "glossary" + ], + "type": "string" + }, + "ttlMs": { + "type": "integer" + }, + "uri": { + "description": "The gvt:// resource the content came from", + "type": "string" + } + }, + "type": "object" +}
- Changed
gvt_get_latest1 field changed- changed
Output schema / (root)Previous value: -nullNew value: +{ + "description": "Most recent non-expired snapshot for one URL. On success the snapshot fields sit at the top level; when no test exists the response is { url, latest: null, message }.", + "properties": { + "analysisSummary": { + "description": "Aggregate scores for one test session", + "properties": { + "categoryIssueCounts": { + "properties": { + "accessibility": { + "description": "Accessibility score (0-100)", + "type": "number" + }, + "headings": { + "description": "Heading structure score (0-100)", + "type": "number" + }, + "javascript": { + "description": "JavaScript dependency score (0-100)", + "type": "number" + }, + "nojs": { + "description": "No-JS rendering score (0-100)", + "type": "number" + }, + "schema": { + "description": "schema.org markup score (0-100)", + "type": "number" + }, + "semantic": { + "description": "Semantic HTML5 score (0-100)", + "type": "number" + }, + "social": { + "description": "Social tags score (0-100)", + "type": "number" + } + }, + "type": "object" + }, + "categoryScores": { + "properties": { + "accessibility": { + "description": "Accessibility score (0-100)", + "type": "number" + }, + "headings": { + "description": "Heading structure score (0-100)", + "type": "number" + }, + "javascript": { + "description": "JavaScript dependency score (0-100)", + "type": "number" + }, + "nojs": { + "description": "No-JS rendering score (0-100)", + "type": "number" + }, + "schema": { + "description": "schema.org markup score (0-100)", + "type": "number" + }, + "semantic": { + "description": "Semantic HTML5 score (0-100)", + "type": "number" + }, + "social": { + "description": "Social tags score (0-100)", + "type": "number" + } + }, + "type": "object" + }, + "jsDependencySource": { + "enum": [ + "measured", + "estimated", + "unavailable" + ], + "type": "string" + }, + "overallScore": { + "description": "Overall GEO visibility score (0-100)", + "type": "number" + }, + "timestamp": { + "description": "ISO 8601 instant of the analysis", + "type": "string" + }, + "totalIssues": { + "description": "Total open issues across all categories", + "type": "integer" + } + }, + "type": "object" + }, + "createdAt": { + "type": "string" + }, + "findingSentence": { + "type": [ + "string", + "null" + ] + }, + "latest": { + "description": "Null when no test exists for the URL", + "type": [ + "object", + "null" + ] + }, + "message": { + "type": "string" + }, + "shareableTid": { + "type": [ + "string", + "null" + ] + }, + "status": { + "type": "string" + }, + "tid": { + "type": "string" + }, + "url": { + "type": "string" + } + }, + "type": "object" +}
- Changed
gvt_get_oldest1 field changed- changed
Output schema / (root)Previous value: -nullNew value: +{ + "description": "Oldest baseline snapshot for one URL, including expired tests. On success the snapshot fields sit at the top level; when no test exists the response is { url, oldest: null, message }.", + "properties": { + "analysisSummary": { + "description": "Aggregate scores for one test session", + "properties": { + "categoryIssueCounts": { + "properties": { + "accessibility": { + "description": "Accessibility score (0-100)", + "type": "number" + }, + "headings": { + "description": "Heading structure score (0-100)", + "type": "number" + }, + "javascript": { + "description": "JavaScript dependency score (0-100)", + "type": "number" + }, + "nojs": { + "description": "No-JS rendering score (0-100)", + "type": "number" + }, + "schema": { + "description": "schema.org markup score (0-100)", + "type": "number" + }, + "semantic": { + "description": "Semantic HTML5 score (0-100)", + "type": "number" + }, + "social": { + "description": "Social tags score (0-100)", + "type": "number" + } + }, + "type": "object" + }, + "categoryScores": { + "properties": { + "accessibility": { + "description": "Accessibility score (0-100)", + "type": "number" + }, + "headings": { + "description": "Heading structure score (0-100)", + "type": "number" + }, + "javascript": { + "description": "JavaScript dependency score (0-100)", + "type": "number" + }, + "nojs": { + "description": "No-JS rendering score (0-100)", + "type": "number" + }, + "schema": { + "description": "schema.org markup score (0-100)", + "type": "number" + }, + "semantic": { + "description": "Semantic HTML5 score (0-100)", + "type": "number" + }, + "social": { + "description": "Social tags score (0-100)", + "type": "number" + } + }, + "type": "object" + }, + "jsDependencySource": { + "enum": [ + "measured", + "estimated", + "unavailable" + ], + "type": "string" + }, + "overallScore": { + "description": "Overall GEO visibility score (0-100)", + "type": "number" + }, + "timestamp": { + "description": "ISO 8601 instant of the analysis", + "type": "string" + }, + "totalIssues": { + "description": "Total open issues across all categories", + "type": "integer" + } + }, + "type": "object" + }, + "createdAt": { + "type": "string" + }, + "findingSentence": { + "type": [ + "string", + "null" + ] + }, + "message": { + "type": "string" + }, + "oldest": { + "description": "Null when no test exists for the URL", + "type": [ + "object", + "null" + ] + }, + "shareableTid": { + "type": [ + "string", + "null" + ] + }, + "status": { + "type": "string" + }, + "tid": { + "type": "string" + }, + "url": { + "type": "string" + } + }, + "type": "object" +}
- Changed
gvt_get_oldest_list1 field changed- changed
Output schema / (root)Previous value: -nullNew value: +{ + "description": "Oldest baseline snapshots for a batch of URLs (up to 100)", + "properties": { + "found": { + "type": "integer" + }, + "requested": { + "type": "integer" + }, + "results": { + "items": { + "properties": { + "oldest": { + "description": "One test snapshot: tid, createdAt, overallScore, categoryScores", + "properties": { + "categoryScores": { + "properties": { + "accessibility": { + "description": "Accessibility score (0-100)", + "type": "number" + }, + "headings": { + "description": "Heading structure score (0-100)", + "type": "number" + }, + "javascript": { + "description": "JavaScript dependency score (0-100)", + "type": "number" + }, + "nojs": { + "description": "No-JS rendering score (0-100)", + "type": "number" + }, + "schema": { + "description": "schema.org markup score (0-100)", + "type": "number" + }, + "semantic": { + "description": "Semantic HTML5 score (0-100)", + "type": "number" + }, + "social": { + "description": "Social tags score (0-100)", + "type": "number" + } + }, + "type": "object" + }, + "createdAt": { + "type": "string" + }, + "overallScore": { + "type": [ + "number", + "null" + ] + }, + "tid": { + "type": "string" + } + }, + "type": [ + "object", + "null" + ] + }, + "url": { + "type": "string" + } + }, + "type": "object" + }, + "type": "array" + }, + "ttlMs": { + "type": "integer" + } + }, + "type": "object" +}
- Changed
gvt_get_prompt1 field changed- changed
Output schema / (root)Previous value: -nullNew value: +{ + "properties": { + "arguments": { + "description": "Effective arguments after binding defaults", + "type": "object" + }, + "description": { + "type": "string" + }, + "name": { + "type": "string" + }, + "text": { + "description": "Fully interpolated instruction text referencing GVT tools by name", + "type": "string" + } + }, + "type": "object" +}
- Changed
gvt_get_score_trend1 field changed- changed
Output schema / (root)Previous value: -nullNew value: +{ + "description": "Baseline vs latest comparison with pre-computed deltas, scoped to one URL or a whole domain", + "properties": { + "domain": { + "type": [ + "string", + "null" + ] + }, + "hasMore": { + "type": "boolean" + }, + "limit": { + "type": "integer" + }, + "nextOffset": { + "type": [ + "integer", + "null" + ] + }, + "offset": { + "type": "integer" + }, + "returned": { + "type": "integer" + }, + "scope": { + "enum": [ + "url", + "domain" + ], + "type": "string" + }, + "trends": { + "items": { + "properties": { + "baseline": { + "description": "One test snapshot: tid, createdAt, overallScore, categoryScores", + "properties": { + "categoryScores": { + "properties": { + "accessibility": { + "description": "Accessibility score (0-100)", + "type": "number" + }, + "headings": { + "description": "Heading structure score (0-100)", + "type": "number" + }, + "javascript": { + "description": "JavaScript dependency score (0-100)", + "type": "number" + }, + "nojs": { + "description": "No-JS rendering score (0-100)", + "type": "number" + }, + "schema": { + "description": "schema.org markup score (0-100)", + "type": "number" + }, + "semantic": { + "description": "Semantic HTML5 score (0-100)", + "type": "number" + }, + "social": { + "description": "Social tags score (0-100)", + "type": "number" + } + }, + "type": "object" + }, + "createdAt": { + "type": "string" + }, + "overallScore": { + "type": [ + "number", + "null" + ] + }, + "tid": { + "type": "string" + } + }, + "type": [ + "object", + "null" + ] + }, + "deltas": { + "description": "Latest minus baseline", + "properties": { + "categories": { + "properties": { + "accessibility": { + "description": "Accessibility score (0-100)", + "type": "number" + }, + "headings": { + "description": "Heading structure score (0-100)", + "type": "number" + }, + "javascript": { + "description": "JavaScript dependency score (0-100)", + "type": "number" + }, + "nojs": { + "description": "No-JS rendering score (0-100)", + "type": "number" + }, + "schema": { + "description": "schema.org markup score (0-100)", + "type": "number" + }, + "semantic": { + "description": "Semantic HTML5 score (0-100)", + "type": "number" + }, + "social": { + "description": "Social tags score (0-100)", + "type": "number" + } + }, + "type": "object" + }, + "overall": { + "type": [ + "number", + "null" + ] + } + }, + "type": [ + "object", + "null" + ] + }, + "isSameTest": { + "description": "True when baseline and latest are the same test (no movement yet)", + "type": "boolean" + }, + "latest": { + "description": "One test snapshot: tid, createdAt, overallScore, categoryScores", + "properties": { + "categoryScores": { + "properties": { + "accessibility": { + "description": "Accessibility score (0-100)", + "type": "number" + }, + "headings": { + "description": "Heading structure score (0-100)", + "type": "number" + }, + "javascript": { + "description": "JavaScript dependency score (0-100)", + "type": "number" + }, + "nojs": { + "description": "No-JS rendering score (0-100)", + "type": "number" + }, + "schema": { + "description": "schema.org markup score (0-100)", + "type": "number" + }, + "semantic": { + "description": "Semantic HTML5 score (0-100)", + "type": "number" + }, + "social": { + "description": "Social tags score (0-100)", + "type": "number" + } + }, + "type": "object" + }, + "createdAt": { + "type": "string" + }, + "overallScore": { + "type": [ + "number", + "null" + ] + }, + "tid": { + "type": "string" + } + }, + "type": [ + "object", + "null" + ] + }, + "testCount": { + "type": "integer" + }, + "url": { + "type": "string" + } + }, + "type": "object" + }, + "type": "array" + }, + "ttlMs": { + "type": "integer" + }, + "url": { + "type": [ + "string", + "null" + ] + }, + "urlCount": { + "type": "integer" + } + }, + "type": "object" +}
- Changed
gvt_get_sitemap_urls1 field changed- changed
Output schema / (root)Previous value: -nullNew value: +{ + "description": "Ranked URL discovery from a domain XML sitemaps", + "properties": { + "bulkUrlLimit": { + "description": "How many of these URLs the caller tier permits in one batch call", + "type": "integer" + }, + "errors": { + "items": { + "type": "object" + }, + "type": "array" + }, + "hasMore": { + "type": "boolean" + }, + "legalPages": { + "enum": [ + "include", + "exclude", + "only" + ], + "type": "string" + }, + "legalPagesFound": { + "description": "Legal/policy pages detected before filtering", + "type": "integer" + }, + "limit": { + "type": "integer" + }, + "matched": { + "type": "integer" + }, + "requestedUrl": { + "type": "string" + }, + "resolvedSitemaps": { + "items": { + "type": "string" + }, + "type": "array" + }, + "returned": { + "type": "integer" + }, + "robotsTxtFound": { + "type": "boolean" + }, + "sitemapCount": { + "type": "integer" + }, + "sort": { + "enum": [ + "significance", + "lastmod", + "alphabetical", + "document" + ], + "type": "string" + }, + "totalDiscovered": { + "type": "integer" + }, + "truncated": { + "description": "True when the sitemap parser hit its 5000-URL cap", + "type": "boolean" + }, + "urls": { + "items": { + "properties": { + "alreadyScheduled": { + "type": "boolean" + }, + "alreadyTested": { + "type": "boolean" + }, + "changefreq": { + "type": [ + "string", + "null" + ] + }, + "disallowedByRobots": { + "type": "boolean" + }, + "lastmod": { + "type": [ + "string", + "null" + ] + }, + "legalPage": { + "type": "boolean" + }, + "pathDepth": { + "type": "integer" + }, + "priority": { + "type": [ + "number", + "null" + ] + }, + "significanceScore": { + "type": "number" + }, + "url": { + "type": "string" + } + }, + "type": "object" + }, + "type": "array" + } + }, + "type": "object" +}
- Changed
gvt_get_test_results1 field changed- changed
Output schema / (root)Previous value: -nullNew value: +{ + "description": "Full results for one test. While status is pending or running, results are incomplete — wait and retry.", + "properties": { + "analysisSummary": { + "description": "Aggregate scores for one test session", + "properties": { + "categoryIssueCounts": { + "properties": { + "accessibility": { + "description": "Accessibility score (0-100)", + "type": "number" + }, + "headings": { + "description": "Heading structure score (0-100)", + "type": "number" + }, + "javascript": { + "description": "JavaScript dependency score (0-100)", + "type": "number" + }, + "nojs": { + "description": "No-JS rendering score (0-100)", + "type": "number" + }, + "schema": { + "description": "schema.org markup score (0-100)", + "type": "number" + }, + "semantic": { + "description": "Semantic HTML5 score (0-100)", + "type": "number" + }, + "social": { + "description": "Social tags score (0-100)", + "type": "number" + } + }, + "type": "object" + }, + "categoryScores": { + "properties": { + "accessibility": { + "description": "Accessibility score (0-100)", + "type": "number" + }, + "headings": { + "description": "Heading structure score (0-100)", + "type": "number" + }, + "javascript": { + "description": "JavaScript dependency score (0-100)", + "type": "number" + }, + "nojs": { + "description": "No-JS rendering score (0-100)", + "type": "number" + }, + "schema": { + "description": "schema.org markup score (0-100)", + "type": "number" + }, + "semantic": { + "description": "Semantic HTML5 score (0-100)", + "type": "number" + }, + "social": { + "description": "Social tags score (0-100)", + "type": "number" + } + }, + "type": "object" + }, + "jsDependencySource": { + "enum": [ + "measured", + "estimated", + "unavailable" + ], + "type": "string" + }, + "overallScore": { + "description": "Overall GEO visibility score (0-100)", + "type": "number" + }, + "timestamp": { + "description": "ISO 8601 instant of the analysis", + "type": "string" + }, + "totalIssues": { + "description": "Total open issues across all categories", + "type": "integer" + } + }, + "type": "object" + }, + "createdAt": { + "type": "string" + }, + "findingSentence": { + "description": "Pre-generated natural language verdict; use as the primary summary", + "type": [ + "string", + "null" + ] + }, + "results": { + "description": "Per-category result records. Each carries analysisType and resultData; issues carry fixId for gvt_get_fix.", + "items": { + "properties": { + "analysisType": { + "enum": [ + "heading", + "semantic", + "accessibility", + "schema", + "javascript", + "nojs", + "social" + ], + "type": "string" + } + }, + "type": "object" + }, + "type": "array" + }, + "shareableTid": { + "description": "Non-null when the result is publicly shareable", + "type": [ + "string", + "null" + ] + }, + "status": { + "enum": [ + "pending", + "processing", + "completed", + "failed" + ], + "type": "string" + }, + "tid": { + "type": "string" + }, + "url": { + "type": [ + "string", + "null" + ] + } + }, + "type": "object" +}
- Changed
gvt_list_prompts1 field changed- changed
Output schema / (root)Previous value: -nullNew value: +{ + "properties": { + "promptCount": { + "type": "integer" + }, + "prompts": { + "items": { + "properties": { + "arguments": { + "items": { + "properties": { + "description": { + "type": "string" + }, + "name": { + "type": "string" + }, + "required": { + "type": "boolean" + } + }, + "type": "object" + }, + "type": "array" + }, + "description": { + "type": "string" + }, + "name": { + "type": "string" + }, + "title": { + "type": "string" + } + }, + "type": "object" + }, + "type": "array" + }, + "ttlMs": { + "type": "integer" + } + }, + "type": "object" +}
- Changed
gvt_list_tests1 field changed- changed
Output schema / (root)Previous value: -nullNew value: +{ + "description": "Filtered, paginated test history, newest first", + "properties": { + "createdFrom": { + "type": [ + "string", + "null" + ] + }, + "createdTo": { + "type": [ + "string", + "null" + ] + }, + "domain": { + "type": [ + "string", + "null" + ] + }, + "hasMore": { + "type": "boolean" + }, + "includeSubdomains": { + "type": "boolean" + }, + "limit": { + "type": "integer" + }, + "nextOffset": { + "type": [ + "integer", + "null" + ] + }, + "offset": { + "type": "integer" + }, + "tests": { + "items": { + "description": "Test session summaries (tid, url, status, analysisSummary, createdAt)", + "type": "object" + }, + "type": "array" + }, + "total": { + "type": "integer" + }, + "ttlMs": { + "type": "integer" + } + }, + "type": "object" +}
- Changed
gvt_run_visibility_batch1 field changed- changed
Output schema / (root)Previous value: -nullNew value: +{ + "description": "Batch acknowledgement. Poll gvt_get_test_results per URL once the batch completes.", + "properties": { + "batchId": { + "description": "UUID for tracking the batch", + "type": "string" + }, + "status": { + "enum": [ + "queued" + ], + "type": "string" + }, + "totalUrls": { + "type": "integer" + }, + "webhookConfigured": { + "type": "boolean" + }, + "webhookSecret": { + "type": [ + "string", + "null" + ] + } + }, + "type": "object" +}
- Changed
gvt_run_visibility_test1 field changed- changed
Output schema / (root)Previous value: -nullNew value: +{ + "description": "Async acknowledgement. Poll gvt_get_test_results with the returned tid until status is not pending or running.", + "properties": { + "message": { + "type": "string" + }, + "next": { + "description": "The suggested next call", + "properties": { + "arguments": { + "properties": { + "tid": { + "type": "string" + } + }, + "type": "object" + }, + "tool": { + "enum": [ + "gvt_get_test_results" + ], + "type": "string" + } + }, + "type": "object" + }, + "status": { + "enum": [ + "pending" + ], + "type": "string" + }, + "tid": { + "description": "8-character test ID to poll with", + "type": "string" + } + }, + "type": "object" +}
- Changed
gvt_schedule_test1 field changed- changed
Output schema / (root)Previous value: -nullNew value: +{ + "description": "Result varies by mode: add/reset return upsertedCount and scheduled; delete returns deletedCount and deletedByUrl", + "properties": { + "deletedByUrl": { + "items": { + "properties": { + "count": { + "type": "integer" + }, + "frequencies": { + "items": { + "enum": [ + "daily", + "weekly", + "monthly", + "once" + ], + "type": "string" + }, + "type": "array" + }, + "url": { + "type": "string" + } + }, + "type": "object" + }, + "type": "array" + }, + "deletedCount": { + "type": "integer" + }, + "scheduled": { + "items": { + "properties": { + "schid": { + "type": "string" + }, + "url": { + "type": "string" + }, + "wpPostId": { + "type": [ + "string", + "null" + ] + } + }, + "type": "object" + }, + "type": "array" + }, + "success": { + "type": "boolean" + }, + "upsertedCount": { + "type": "integer" + } + }, + "type": "object" +}
- Changed
gvt_set_test_visibility1 field changed- changed
Output schema / (root)Previous value: -nullNew value: +{ + "properties": { + "is_public": { + "type": "boolean" + }, + "success": { + "type": "boolean" + }, + "tid": { + "type": "string" + } + }, + "type": "object" +}
- Changed
gvt_share_test1 field changed- changed
Output schema / (root)Previous value: -nullNew value: +{ + "properties": { + "is_public": { + "type": "boolean" + }, + "share_url": { + "description": "Public URL, or null when sharing failed", + "type": [ + "string", + "null" + ] + }, + "success": { + "type": "boolean" + }, + "tid": { + "type": "string" + } + }, + "type": "object" +}
2 tool updates
- Removed
gvt_manage_test - Added
gvt_set_test_visibility
2 tool updates
- Removed
gvt_batch_run_tests - Added
gvt_run_visibility_batch
1 tool update
- Changed
gvt_get_sitemap_urls2 fields changed- changed
Input schema / properties / limit / defaultPrevious value: -25New value: +10 - changed
Input schema / properties / limit / descriptionPrevious value: -"Maximum number of URLs to return."New value: +"Maximum number of URLs to return (default 10)."
17 tool updates
- First observed
gvt_batch_run_tests - First observed
gvt_delete_test - First observed
gvt_get_fix - First observed
gvt_get_knowledge - First observed
gvt_get_latest - First observed
gvt_get_oldest - First observed
gvt_get_oldest_list - First observed
gvt_get_prompt - First observed
gvt_get_score_trend - First observed
gvt_get_sitemap_urls - First observed
gvt_get_test_results - First observed
gvt_list_prompts - First observed
gvt_list_tests - First observed
gvt_manage_test - First observed
gvt_run_visibility_test - First observed
gvt_schedule_test - First observed
gvt_share_test
Related MCP Connectors
AI visibility reports, competitor insights, readiness audits, and GEO content workflows.
- CiteHawkOAuthcom.citehawk
AI-search visibility (GEO) analytics: brand scores, competitor rankings, recommendations, evidence.
Free SEO, GEO, and AEO audits: analyze any page or domain, AI-crawler access, agent readiness.
Measure how AI engines cite your brand. Cross-engine GEO visibility, as agent tools.
Related MCP Servers
- -licenseNot gradedqualityCmaintenanceEnables AI assistants to perform comprehensive SEO and GEO measurements, including site audits, keyword research, ranking tracking, and brand visibility analysis across search engines and generative AI platforms.-
- AlicenseAqualityAmaintenanceProvides AI-visibility scoring and site auditing capabilities for websites, enabling agents to check how sites appear in AI engines like ChatGPT and Perplexity, run full SEO/security audits, and monitor changes over time.15156 npmMIT
- AlicenseAqualityDmaintenanceEnables users to scan any website for AI search visibility, producing AEO, GEO, agent readiness, and mention-readiness scores along with AI identity and business profile insights. Paid tools extend this to competitive comparisons, detailed audits, and generated fixes.41MIT

Seonix SEO MCPofficial
AlicenseAqualityCmaintenanceLets any AI agent audit any website for SEO, GEO/AEO, and speed problems, reporting issues and recommendations without modifying the site.4MIT
Glama MCP Gateway
Add one secure layer between your agents and this server.