Citation Intelligence MCP
This is a self-hosted Citation Intelligence platform for tracking, analyzing, and improving how AI search engines cite your content. It covers citation lookup, competitive analysis, site auditing, and trend monitoring across Perplexity, Claude, ChatGPT, Gemini, Google AI Mode, and Bing.
Citation Checking & Discovery
Check which URLs are cited by AI engines for any query, including cross-engine consensus and individual engine results
Verify if your domain is cited across a cluster of queries with per-engine breakdowns and citation rate summaries
Extract the actual snippets explaining why a URL was cited, not just that it was
Determine where your URL appears in an AI's answer (early, middle, or late)
Signals & Monitoring
Check if a Google AI Overview exists for a query and which URLs it cites
Identify Wikipedia articles referencing your domain (a key LLM training signal, no API key needed)
Join Google Search Console data with AI citation status to find queries where you rank in Google but are invisible to AI
Score the recency of cited pages to surface opportunities to publish fresher content
Trend Tracking & History
Track citation rate changes over time using saved query panels and snapshot comparisons
Diff citation history between two time windows to see which queries were gained or lost
List all queries a domain has been cited for from local cache (no API spend)
Competitive Analysis
Identify the top-cited competitor domains for any query across multiple engines
Compare your URL's citation signals side-by-side against competitors to pinpoint gaps
Run end-to-end competitive snapshots for a query
Site Auditing
Predict citation likelihood (0–100) for any URL from public signals (schema.org, llms.txt, Wikipedia, GitHub, HTTPS, content depth) — no LLM required
Bulk audit all URLs in a sitemap for citation likelihood, sorted worst-first
Cross-reference sitemap URLs with the citation cache to see which pages have already been cited
Validate schema.org implementation, flag missing or malformed fields, and generate ready-to-paste JSON-LD fixes
Verify AI crawler access (GPTBot, ClaudeBot, PerplexityBot, etc.) via robots.txt and live user-agent checks
Generate an
llms.txtfile from a sitemap to guide AI systems on your site's structure
Privacy: Runs locally, uses your own API keys, caches results locally, and sends no data to third parties.
Performs web rank checks using the Brave Search API.
Uses GitHub as a signal for citation likelihood prediction.
Retrieves Google AI Overview citations via SerpAPI.
Joins Google Search Console performance data with AI citation status for gap analysis.
Integrates with OpenAI's Responses API to retrieve citations from ChatGPT.
Lists Wikipedia articles that mention a given domain.
Citation Intelligence MCP
A free, self-hosted MCP server that tells your agent what LLMs cite - across Perplexity, Google AI Overviews, ChatGPT, Claude, Gemini, and Bing.
What this is
An MCP server for agents and developers who need to know which URLs get cited by AI search engines for any query. Install once, query from any MCP-compatible client (Claude Desktop, Cursor, Claude Code, Continue, Cline, n8n, LangGraph). Self-hosted, no account, no centralized backend. Bring your own API keys; nothing is stored on a remote server.
Related MCP server: openresearch-mcp
Who this is for
Install this if you're:
Building an agent that does research and want it to cite sources LLMs already trust
A solo dev or indie hacker checking whether your SaaS is showing up in AI search
A content creator confirming your articles are being cited by ChatGPT, Claude, or Perplexity
An SEO or GEO practitioner who wants programmatic citation data without a $295-$499/mo dashboard
Running an editorial pipeline and want citation-deficit-driven topic selection
Comparing competitor visibility across AI engines for any niche
Do NOT install this if you want:
A polished marketing dashboard with charts and team seats - try Profound, AthenaHQ, or Otterly.AI
A hosted service with SLAs - this is self-hosted by design
Citation tracking for academic papers - try citecheck
350M+ pre-modeled prompts - that's Ahrefs Brand Radar
Why this exists
The AI citation tracking market is dominated by VC-funded dashboards starting at $295/mo. None ships MCP-first. If you're an agent or developer who wants citation data piped directly into your workflow - not into a SaaS login - there isn't a tool for you. This is that tool.
Tools
Tools are grouped into seven namespaces: citations_*, domain_*, signals_*, panel_*, report_*, competitors_*, audit_*. The prefix is the question category; the suffix is the action. Wire names use underscores (not dots) so Anthropic-API-based MCP clients (Claude Desktop, Claude Code) can forward the tool list without HTTP 400.
Start with citations_provenance or domain_am_i_cited. Single-engine results (citations_check with a pinned engine) are directional; multi-engine consensus is the honest signal. A URL cited by 4 of 5 engines is a very different finding than one cited by 1.
citations_* — query-level: who cites what, with what evidence
Tool | Purpose |
| Recommended first tool. Fan a query across engines; per-URL cross-engine consensus matrix. Returns |
| URLs cited by Perplexity / Claude / ChatGPT / Gemini / Google AI Mode for a query; or web rank via bing_serp / brave_serp |
| Extract the cited snippet from |
| Citation likelihood from public signals - no LLM fired |
| Time-series report of citation rate + per-query gained/lost deltas |
| Recency score (halflife=365d) for the pages an engine cites |
domain_* — domain-level: am I cited, what for
Tool | Purpose |
| Domain citation check. With |
| Queries the domain has been cited for, from local cache |
| Diff of |
signals_* — external signals: AI Overview, Wikipedia, GSC, answer-box position
Tool | Purpose |
| Google AI Overview presence + cited sources |
| List Wikipedia articles referencing a domain (zero keys) |
| Join Google Search Console performance with AI citation status |
| Bin each citation's first mention in |
panel_* — saved query panels (editorial watchlists)
Tool | Purpose |
| Save / load / list named query panels (editorial watchlists) |
| Run a panel through |
report_* — turnkey reporting artifacts
Tool | Purpose |
| One-call AI visibility report over a query set (or panel): citation rate (mention frequency), share of voice vs competitors, average rank, and brand sentiment. Returns structured data + a Markdown artifact for a public page. |
competitors_* — competitive landscape per query
Tool | Purpose |
| Top cited domains per query, aggregated across engines |
| End-to-end competitive snapshot: your URL vs top cited competitors |
| Side-by-side |
audit_* — fixable on-page / on-site checks
Tool | Purpose |
| Deep schema.org validation - required fields per |
| Repair-oriented schema.org diagnostics + suggested patches |
| Verify GPTBot / ClaudeBot / PerplexityBot / CCBot / Google-Extended etc. can fetch a URL |
| Bulk |
| Cross-reference sitemap URLs with cached citations (inverse of |
| Generate an |
Prompts
Server-side prompt templates the client can offer end users (call via the MCP prompt list):
audit_citation_readiness(url)- chainscitations_predict+audit_schemaaudit_competitor_snapshot(query, your_url?)- chainscompetitors_canonical_set+competitors_competeaudit_crawler_checkup(url)- runsaudit_crawler_accessand writes a remediation listaudit_gap_analysis(domain, days?)- drivessignals_gsc_gapand suggests next movesaudit_sitemap_coverage(sitemap_url)- runsaudit_sitemap_mapand recommends priorities
Resources
Cache views the client can read or subscribe to (no tool call required):
citation://cache/summary- entry counts by type/engine, unique queries/URLs, oldest/newestcitation://panels- saved panels + per-panel snapshot countscitation://docs/llms-txt- llms.txt primer (markdown)citation://docs/ai-crawlers- AI crawlers cheatsheet (markdown)citation://domain/{domain}/cited-for- dynamic template: citations for{domain}
What this actually measures
Every response includes a surface field that tells you exactly how the data was collected. Understanding this is important before drawing conclusions.
Surface | Engines | What it means |
|
| Proxied through a real consumer-facing AI search product. Closest to what your users see. |
|
| API call to a search-enabled LLM. May differ from consumer product behavior — different model versions, no UI-level ranking logic, no personalization. Use as a directional proxy, not as ground truth. |
|
| Traditional web search rank (not LLM citation). Measures whether a URL appears in SERP results, not whether an LLM cites it. |
|
| Offline signal computed from public data. No live LLM query. |
Per-engine notes
perplexity (consumer_scrape) — Sonar Pro via the Perplexity API with a consumer-equivalent system prompt. Reasonably close to Perplexity.ai. Citations come from search_results in the response; the citations fallback contains URL-only entries without title.
claude (api_proxy) — Claude Sonnet via the Anthropic Messages API with web_search tool enabled. The consumer Claude.ai product uses different routing and ranking logic. Citation behavior can differ, especially for recent/time-sensitive queries.
openai (api_proxy) — gpt-4o + the web_search_preview tool via the OpenAI Responses API. Replaces the deprecated gpt-4o-search-preview alias OpenAI retired; base gpt-4o plus the tool is the supported path.
gemini (api_proxy) — Gemini 2.5 Pro via the Generative Language API with google_search grounding. Consumer Gemini uses the same grounding index but different re-ranking. Results are directional.
google_ai_mode (consumer_scrape) — Google AI Mode results via SerpAPI. Closest to what users see in Google Search. Requires SERPAPI_KEY.
bing_serp / brave_serp (web_rank) — Traditional SERP rank. Does NOT measure LLM citations. Use citations_check with these engines to compare organic web rank against LLM citation rank. domain_am_i_cited refuses these engines — it only measures LLM behavior.
The proxy nature of api_proxy engines is a feature, not a bug: it lets you run citation checks without consuming expensive consumer-product quota. Just don't report API-proxy numbers as "ChatGPT cites you" without the caveat.
Every tool response includes an interpretation_note field that summarizes the fidelity in one sentence. Full per-engine fidelity ratings: docs/surface-fidelity.md.
Quick start
npx -y @automatelab/citation-intelligenceRequires Node 20 or later.
Claude Desktop
Add to %APPDATA%\Claude\claude_desktop_config.json (Windows) or ~/Library/Application Support/Claude/claude_desktop_config.json (macOS):
{
"mcpServers": {
"citation-intelligence": {
"command": "npx",
"args": ["-y", "@automatelab/citation-intelligence"],
"env": {
"PERPLEXITY_API_KEY": "pplx-...",
"SERPAPI_KEY": "...",
"ANTHROPIC_API_KEY": "sk-ant-...",
"OPENAI_API_KEY": "sk-...",
"GEMINI_API_KEY": "..."
}
}
}
}Set only the keys you have. Any MCP client that supports stdio transport works - same command / args pattern.
How it stays free
No central backend. The server runs on your machine. Nothing is uploaded.
Free tier first. SerpAPI gives 100 free Google AI Overview lookups/month. Bing Web Search has a free tier. Perplexity offers free Sonar access on signup.
Bring your own paid keys if you want the premium engines (Claude, ChatGPT, Gemini). Keys pass through to the vendor and never touch any third party.
Local cache at
~/.config/citation-intelligence/cache.json. Repeated queries hit cache, not API. Default TTL: 7 days.citations_predictruns with zero keys - it scores citation likelihood from public signals (Wikipedia, schema.org, llms.txt, GitHub) without firing any LLM.
Privacy
All API calls go from your machine directly to the vendor (Anthropic, OpenAI, Google, Perplexity, Bing, SerpAPI).
No proxy. No analytics. No telemetry by default.
API keys are read from environment variables on the MCP process - never logged, never persisted.
Cache file lives at
~/.config/citation-intelligence/cache.json. Delete it any time.
Environment variables
Var | Purpose | Free tier? |
|
| Yes |
|
| 100/month free |
|
| Paid only |
|
| Paid only |
|
| Yes |
|
| Yes |
|
| Yes (2000/month) |
| Cache TTL for | n/a |
| Cache TTL for | n/a |
| Override config dir (default | n/a |
Example: am I cited?
You: For the queries "best AI citation tracker", "MCP for AI search", "self-hosted GEO tool",
is automatelab.tech cited?
(agent invokes `domain_am_i_cited`)
Result:
{
"domain": "automatelab.tech",
"engine": "perplexity",
"results": [
{ "query": "best AI citation tracker", "cited": true, "rank": 4 },
{ "query": "MCP for AI search", "cited": true, "rank": 1 },
{ "query": "self-hosted GEO tool", "cited": false, "matching_urls": [] }
],
"summary": {
"queries_total": 3,
"queries_cited": 2,
"citation_rate": 0.67,
"average_rank": 2.5
}
}Example: predict citation likelihood (no key required)
You: How likely is https://example.com/blog/post to be cited by AI?
(agent invokes `citations_predict`)
Result:
{
"url": "https://example.com/blog/post",
"score": 62,
"grade": "C",
"signals": {
"wikipedia_linked": false,
"github_referenced": false,
"reddit_referenced": true,
"llms_txt_present": true,
"https": true,
"has_article_schema": true,
"has_faq_schema": false,
"has_breadcrumb_schema": true,
"canonical_clean": true,
"word_count": 1850,
"reading_time_minutes": 8,
"h2_count": 7,
"h2_question_count": 1,
"authority_link_count": 2,
"external_link_count": 6,
"internal_link_count": 11,
"last_modified_days_ago": 42,
"has_open_graph": true
},
"fixes": [
{ "signal": "has_faq_schema", "suggestion": "Page already has question-style H2s. Wrap them in FAQPage JSON-LD - high-leverage win.", "estimated_lift": "high" },
{ "signal": "h2_question_count", "suggestion": "Reframe at least 2 H2s as questions users actually ask...", "estimated_lift": "medium" }
]
}The Wikipedia signal is measured (it correlates with citation) but no "go get a Wikipedia article" suggestion is emitted - the advice would be non-actionable. Scoring is split across six buckets - domain authority, structured data, content depth, link graph, freshness, metadata - so a thin page and a deep page on the same domain get meaningfully different scores.
Workflow recipes
Concrete patterns that compose the 26 tools into something useful. Costs assume ChatGPT or Perplexity at ~$0.01-0.03/query.
1. Weekly citation tracker
The single highest-ROI pattern. Pick 20-30 queries from your editorial backlog, snapshot weekly, watch the rate trend.
# One-time setup
panel_track name="editorial-watchlist" domain="example.com" action="save"
queries=["best widget tutorial", "how to set up X", ...]
# Weekly cron (5 min, ~$0.20-0.60 per run)
panel_run name="editorial-watchlist"
# Anytime
citations_trend panel="editorial-watchlist"citations_trend returns per-query deltas: which queries flipped from cited: false to cited: true since the first snapshot. That's your real editorial-impact metric.
2. Pre-publish gate
Before publishing a post, find out who owns the citation slot and whether the slot is worth competing for.
# 1. Is there an AI Overview to compete for?
signals_ai_overview query="<target query>"
# 2. Who is cited today?
citations_check query="<target query>"
# 3. After publish + 14 days: did the post break in?
domain_am_i_cited domain="example.com" queries=["<target query>"]If citations_check returns 5+ strong incumbents on a low-volume query, pick a different angle. If ai_overview_present: false, the query has no AI surface - reconsider.
3. Bulk site audit
Catch site-wide structural issues across every page in one pass. Zero API spend.
audit_sitemap sitemap_url="https://example.com/sitemap.xml" limit=200Returns worst_first sorted by citation-likelihood score. Surfaces missing schema, conflicting canonicals, missing /llms.txt, broken HTTPS.
4. Competitor signal gap
You're not cited; they are. Why?
# 1. Find the top-cited URLs for your target query
citations_check query="<query>"
# 2. Compare your URL to theirs signal-by-signal
competitors_compare urls=[
"https://example.com/your-post",
"https://competitor-1.com/their-post",
"https://competitor-2.com/their-post"
]diverging_signals is the list of where you're losing. Usually obvious once you see it - they have FAQ schema, GitHub references, Wikipedia links - you don't.
5. Google-rank vs AI-citation gap
The closest editorial wins are queries where you already rank in Google's top 10 but are invisible to AI. Requires a GCP service account with webmasters.readonly scope.
signals_gsc_gap
domain="example.com"
queries=["...editorial watchlist..."]
start_date="2026-04-01"
end_date="2026-05-01"closest_wins returns queries with position <= 10 and ai_cited: false, sorted by impressions desc. Push citation signals on those specific URLs first.
6. Wikipedia mention monitor
Wikipedia is the top-correlation signal but the advice "get on Wikipedia" is useless. So instead: watch when it happens organically.
signals_wikipedia domain="example.com" limit=50Returns Wikipedia article URLs that already link to the domain. Re-run quarterly; the diff is your "we got a Wikipedia citation" alert.
Schema.org
{
"@context": "https://schema.org",
"@type": "SoftwareApplication",
"name": "Citation Intelligence MCP",
"applicationCategory": "DeveloperApplication",
"operatingSystem": "Cross-platform",
"description": "Self-hosted MCP server for querying AI citation data from Perplexity, Claude, ChatGPT, Gemini, Bing, and Google AI Overviews.",
"offers": { "@type": "Offer", "price": "0" },
"url": "https://github.com/AutomateLab-tech/citation-intelligence"
}Contributing
Bug reports, feature ideas, and PRs welcome. See CONTRIBUTING.md.
Security
Report a vulnerability via SECURITY.md.
License
MIT - see LICENSE.
Built by automatelab.tech
Available Tools
26 toolsaudit_crawler_accessARead-onlyIdempotent
Verify that major AI crawlers (GPTBot, OAI-SearchBot, ClaudeBot, PerplexityBot, CCBot, Google-Extended, Applebot-Extended, Bytespider, Meta-ExternalAgent, plus real-time fetch UAs) can fetch a URL. Parses robots.txt and does a live GET with each bot's User-Agent. Surfaces robots.txt blocks AND UA-based gating that breaks AI citation.
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | Page URL to test for AI crawler access. | |
| bots | No | Override the default bot list. Each entry is a User-Agent token (e.g. 'GPTBot', 'ClaudeBot'). | |
| fetch_with_ua | No | If true, do a live GET as each bot's User-Agent and report status. Disable to only parse robots.txt (no extra requests). |
Output Schema
| Name | Required | Description |
|---|---|---|
| url | Yes | The URL that was audited. |
| bots | Yes | Per-bot access verdict combining robots.txt + live UA test. |
| note | No | |
| summary | Yes | |
| fetched_at | Yes | UTC ISO-8601 timestamp. |
| robots_url | Yes | robots.txt URL that was parsed. |
| robots_error | Yes | Error message if robots.txt fetch failed. |
| robots_status | Yes | HTTP status of the robots.txt fetch. |
| robots_present | Yes | Whether a non-empty robots.txt was found. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, openWorldHint, and destructiveHint=false. The description adds valuable behavioral context beyond that: it parses robots.txt and makes live GET requests using each bot's User-Agent, and it surfaces both robots.txt blocks and UA-based gating. It does not mention timeouts, rate limits, or network cost, but the added detail is substantial.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tightly packed sentences with no filler. The core action is front-loaded, and the long crawler list is justified because those names are operationally relevant to the tool.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so the description need not explain return values. Annotations cover the safety profile and the description covers the core behavior, making it nearly complete. It could still be clearer about when this tool is preferable to sibling audit tools, but no critical invocation detail is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents url, bots, and fetch_with_ua. The description names the default bot list, which lightly reinforces the bots parameter, but adds no syntax or behavioral meaning beyond what the schema already provides.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: verify AI crawler access to a URL. It names the crawler set and explains that it parses robots.txt and performs live GETs, clearly distinguishing it from broader audit siblings like audit_sitemap or audit_llms_txt.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description makes the use case inferable — diagnosing whether AI crawlers can fetch a page — but does not explicitly say when to choose this over sibling audit tools, nor does it state exclusions or prerequisites. Usage is implied rather than spelled out.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
audit_llms_txtARead-onlyIdempotent
Generate an llms.txt file (https://llmstxt.org spec) from a sitemap. Parses sitemap.xml + nested indexes, groups URLs by top-level path, and emits a Markdown document with H1+description+sectioned link lists. Set fetch_titles=true to pull per URL (slower, richer output).
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | Max URLs to include. Truncated after sitemap parse, before title fetch. | |
| site_title | Yes | Site title - top H1 in the generated llms.txt file. | |
| sitemap_url | Yes | URL of sitemap.xml (or sitemap index). Nested sitemaps are followed. | |
| fetch_titles | No | If true, fetch each URL to extract <title> for richer links. Slower (one HEAD-ish GET per URL). Default false uses the URL path as the link text. | |
| site_description | No | One-paragraph site description placed under the H1. Optional but strongly recommended. |
Output Schema
| Name | Required | Description |
|---|---|---|
| note | No | |
| content | Yes | Generated llms.txt file content; save to /llms.txt at site root. |
| sections | Yes | Number of top-level sections in the generated file. |
| fetched_at | Yes | UTC ISO-8601 timestamp. |
| sitemap_url | Yes | Sitemap URL that was processed. |
| urls_included | Yes | URLs included after applying the limit. |
| titles_fetched | Yes | Number of pages fetched to extract <title>. |
| total_urls_in_sitemap | Yes | Total URLs found in the sitemap. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover the safety profile (readOnly, idempotent, non-destructive, openWorld). The description adds useful procedural context: sitemap parsing, nested index following, and the slower/richer tradeoff of fetch_titles. However it doesn't state limits, pagination, or failure modes beyond what the schema already says.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences, front-loading the artifact and spec, then the pipeline, then the key parameter tradeoff. No filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Annotations carry safety, an output schema exists, and schema descriptions are complete. The description fills the remaining gaps (pipeline behavior, fetch_titles tradeoff). Minor gap: no mention of rate limits or failure behavior when the sitemap is unreachable.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so baseline is 3. The description adds meaning by explaining the fetch_titles tradeoff and the output structure (H1+description+sectioned links), which helps the agent reason about parameter selection beyond the schema's per-parameter notes.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (generate) and artifact (llms.txt file) and names the spec URL. Distinct from every sibling (audit_*, citations_*, competitors_*, etc.), none of which produce an llms.txt.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explains what it does and hints at the fetch_titles tradeoff, but provides no explicit when-to-use vs when-not-to-use guidance or alternative selection against siblings.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
audit_schemaARead-onlyIdempotent
Deep schema.org validation for a URL. Parses every JSON-LD block and microdata node, checks required fields per @type (Article needs headline+author+datePublished, FAQPage needs mainEntity, HowTo needs step, etc.), and flags missing fields and malformed JSON-LD. Returns issues list and a valid/invalid verdict. Use to fix structured-data bugs that predict_citation flags but can't explain.
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | URL whose JSON-LD and microdata to validate against schema.org expected fields. |
Output Schema
| Name | Required | Description |
|---|---|---|
| url | Yes | URL that was audited. |
| note | No | |
| issues | Yes | Validation issues found. |
| summary | Yes | |
| fetched_at | Yes | UTC ISO-8601 timestamp. |
| json_ld_blocks | Yes | Total JSON-LD blocks found. |
| json_ld_parse_errors | Yes | Number of JSON-LD blocks that failed to parse. |
| schema_types_present | Yes | @type values found across all JSON-LD blocks. |
| microdata_types_present | Yes | Schema types found via microdata (itemtype). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly, idempotent and non-destructive behavior, so safety is covered. The description adds real value beyond that: it discloses the parsing method (every JSON-LD block and microdata node), the validation rules applied per @type, and that malformed JSON-LD is flagged. It does not discuss cost, rate limits, or failure modes.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two dense sentences, front-loaded with the verb and resource, and the inline @type examples earn their place by showing the depth of validation. Slightly run-on, but there is no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists and the description still summarizes the return shape (issues list plus valid/invalid verdict), so return values are covered. Annotations cover the safety profile. The main omission is routing logic versus the very similar audit_structured_data sibling.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Only one parameter and schema description coverage is 100%, so the schema already defines the url fully. The description adds only that both JSON-LD and microdata at that URL are validated, which is scope rather than parameter syntax. Baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (validate) and resource (schema.org structured data) and enumerates exactly what is parsed and checked (JSON-LD blocks, microdata nodes, required fields per @type). However, it never distinguishes itself from the near-identical sibling audit_structured_data, so an agent cannot tell the two apart from the description alone.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives a concrete trigger: use it to diagnose structured-data bugs that predict_citation flags but cannot explain. That is a clear when-to-use context. It does not state when to prefer audit_structured_data instead, nor any preconditions (e.g., page must be reachable).
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
audit_sitemapARead-onlyIdempotent
Fetch a sitemap.xml (or sitemap index) and run predict_citation on every URL. Returns results sorted worst-score-first. Surfaces systemic issues across a whole site in one pass. Zero engine keys needed.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | Max URLs to score. Sitemap is sliced after parsing. | |
| concurrency | No | Parallel predict_citation calls. Higher is faster but more rate-limit risk. | |
| sitemap_url | Yes | URL of sitemap.xml (or a sitemap index). Nested sitemaps are followed. |
Output Schema
| Name | Required | Description |
|---|---|---|
| errors | Yes | URLs whose audit threw an error. |
| audited | Yes | Number of URLs that were scored. |
| fetched_at | Yes | UTC ISO-8601 timestamp. |
| total_urls | Yes | Total URLs found in the sitemap. |
| sitemap_url | Yes | The sitemap URL that was audited. |
| worst_first | Yes | Up to 20 lowest-scoring URLs, worst first. |
| average_score | Yes | Mean predict_citation score across scored URLs. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover safety (readOnly, idempotent, non-destructive, open-world), so the bar is lower, and the description adds genuinely useful behavior: results sorted worst-score-first and 'zero engine keys needed'. That keyless-run detail is valuable operational context not in the annotations. It stops short of describing pagination/truncation behavior, which the limit param implies.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four short sentences, front-loaded with the action and ending with a differentiating benefit. No filler; each sentence adds either mechanism, output shape, value, or constraint.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With an output schema present and rich annotations, the description needn't explain return values and doesn't try. It supplies the ordering and keyless-run context an agent needs, though it omits guidance on concurrency/rate-limit tradeoffs and the sibling overlap already noted.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so limit, concurrency and sitemap_url are fully documented by the schema; baseline 3 applies. The description adds no parameter-specific meaning, and its claim of running on 'every URL' sits slightly in tension with the limit/concurrency slicing controls.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource chain: fetch a sitemap, run predict_citation on every URL, and it names the output ordering. This clearly separates it from generic audit siblings like audit_schema or audit_llms_txt. However, it does not differentiate itself from the close sibling audit_sitemap_map, which an agent would need help distinguishing.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied rather than explicit – 'Surfaces systemic issues across a whole site in one pass' conveys the scenario (whole-site diagnostics) but nowhere states when to prefer this over audit_sitemap_map or predict_citation directly. No exclusions or prerequisites are given.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
audit_sitemap_mapARead-onlyIdempotent
Cross-reference a sitemap with the citation cache. For each sitemap URL, reports whether it appears in cached citations (and how many queries/engines cited it). Inverse of audit_sitemap: not 'how citable is each URL', but 'has each URL actually been cited yet'. Cache must be primed via check_citations or run_panel first.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | Max sitemap URLs to consider. | |
| since | No | ISO date floor; only count citations recorded on or after this date. | |
| domain | No | Domain to look up citations for. If omitted, inferred from the sitemap host. | |
| sitemap_url | Yes | URL of sitemap.xml (or a sitemap index). Nested sitemaps are followed. |
Output Schema
| Name | Required | Description |
|---|---|---|
| note | No | |
| since | No | |
| domain | Yes | Domain whose citation cache was queried. |
| mapped | Yes | URLs found in the citation cache. |
| unmapped | Yes | URLs not yet seen in the citation cache. |
| fetched_at | Yes | UTC ISO-8601 timestamp. |
| total_urls | Yes | Total sitemap URLs considered. |
| mapped_urls | Yes | Sitemap URLs present in the cache, sorted by citation count desc. |
| sitemap_url | Yes | The sitemap that was processed. |
| coverage_pct | Yes | Percentage of sitemap URLs that have been cited (0-100). |
| unmapped_urls | Yes | Sitemap URLs not yet cited (up to 200). |
| citations_in_cache | Yes | Total citation cache entries for this domain. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare read-only, idempotent, non-destructive behavior, so the safety profile is covered. The description adds the important operational dependency (prime the cache first) and the per-URL reporting content (whether/ how many queries and engines cited it), which the annotations do not convey.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight sentences, front-loaded with the core operation, followed by the inverse framing and the prerequisite. No filler; every sentence carries unique information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists so return values need not be explained, and the description covers the operation, the differentiation from siblings, and the cache-priming prerequisite. Nothing an agent needs to call it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so limit, since, domain and sitemap_url are all documented in the schema. The description adds no syntax or format detail beyond what the schema provides, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (cross-reference) and resources (sitemap vs. citation cache), then explicitly frames itself as the inverse of audit_sitemap. An agent can distinguish it from the sibling without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Names the sibling it contrasts with (audit_sitemap) and states the precondition to use it at all: the cache must be primed via check_citations or run_panel first. This is explicit alternative and prerequisite guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
audit_structured_dataARead-onlyIdempotent
Suggest missing JSON-LD additions for a URL. Fetches the page, detects existing schema types, and returns ready-to-paste templates for types that are missing but signalled by page content (BlogPosting from og:type=article or bylines, FAQPage from Q&A pairs, HowTo from numbered steps, BreadcrumbList from nested paths, Organization on homepages). Templates are pre-filled from page metadata where possible; fields marked FILL: require manual completion.
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | URL to inspect for missing JSON-LD. The page is fetched and its content signals are used to suggest schema types. |
Output Schema
| Name | Required | Description |
|---|---|---|
| url | Yes | URL that was inspected. |
| note | No | |
| summary | Yes | |
| fetched_at | Yes | UTC ISO-8601 timestamp. |
| suggestions | Yes | Schema additions suggested for this page. |
| signals_detected | Yes | Content signals detected (e.g. og:type=article). |
| schema_types_present | Yes | @type values already present on the page. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover readOnly, idempotent, openWorld and non-destructive. The description adds genuinely useful behavior beyond that: it fetches the page, detects existing types, returns ready-to-paste templates, pre-fills from metadata, and marks incomplete fields with FILL:. That return-shape and convention context is valuable.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with a one-line purpose, followed by useful specifics. The parenthetical signal-to-type mapping is dense but each example earns its place by clarifying detection behavior.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With an output schema present and annotations carrying the safety profile, the description need only cover purpose and behavior, which it does. The only residual gap is the unstated relationship to the sibling audit_schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% and there is a single url parameter already documented in the schema, so the description adds little parameter-level meaning beyond noting the page is fetched. Baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (suggest) and resource (missing JSON-LD additions) with clear scope, and enumerates exactly which types it can propose from which signals. It does not explicitly distinguish itself from the sibling audit_schema, which likely inspects existing schema markup, so an agent must infer the boundary.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies when the tool is useful (page is missing structured data but content signals suggest types), but gives no explicit when-to-use/when-not or naming of alternatives like audit_schema. Usage is left to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
citations_checkAIdempotent
Return URLs cited by an AI engine (Perplexity, Claude, ChatGPT, Gemini, or Bing) for a query. Use this when an agent or user wants to see what sources an AI search engine grounds answers on. Requires at least one engine API key; auto-picks the first available.
| Name | Required | Description | Default |
|---|---|---|---|
| query | Yes | The search query to test (what would a user ask an AI?) | |
| engine | No | Engine to query. • perplexity / google_ai_mode — consumer_scrape: closest to real product behavior. • claude / openai / gemini — api_proxy: API-tier call, may differ from consumer product. • bing_serp / brave_serp — web_rank: traditional SERP rank, NOT LLM citation. 'auto' prefers SerpAPI (google_ai_mode) → Perplexity → LLM adapters → web_rank. | auto |
| max_results | No | Maximum citations to return. | |
| perplexity_model | No | Perplexity model override (e.g. 'sonar', 'sonar-pro', 'sonar-reasoning'). Only used when engine='perplexity'. Defaults to 'sonar-pro'. |
Output Schema
| Name | Required | Description |
|---|---|---|
| query | Yes | The query that was executed. |
| cached | Yes | Whether the result was served from the local cache. |
| engine | Yes | Engine used for this response. |
| surface | Yes | Engine surface type: consumer_scrape, api_proxy, or web_rank. |
| citations | Yes | Cited URLs ordered by rank. |
| fetched_at | Yes | UTC ISO-8601 timestamp of the fetch. |
| raw_answer | No | Raw answer text from the engine, if available. Absent for web_rank engines (bing_serp, brave_serp) that return ranked URLs without a synthesized answer. |
| interpretation_note | Yes | Guidance on how to interpret results from this engine. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations cover the safety profile (openWorldHint, idempotentHint, destructiveHint=false), so the bar is lower. The description still adds genuine operational context beyond the structured fields: an engine API key is required and the tool auto-picks the first available engine, which affects setup and failure modes.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences with zero padding; the output is described first and the usage condition and prerequisite follow. Every clause earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With an output schema present and full parameter coverage, the description need not explain return values. It covers purpose, usage trigger, and the API-key prerequisite, leaving only sibling disambiguation as a gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and the enum, defaults, max_results bounds, and perplexity_model override are all fully documented in the schema. The description adds no parameter-level detail beyond naming the engines, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: 'Return URLs cited by an AI engine ... for a query', and enumerates the engines covered. It does not distinguish itself from the many citations_* siblings (evidence, provenance, freshness, predict), so an agent cannot tell from the description alone which citation tool to pick.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives a clear use condition: 'Use this when an agent or user wants to see what sources an AI search engine grounds answers on.' There is no when-not guidance and no explicit routing to alternative sibling tools such as citations_evidence or citations_provenance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
citations_evidenceARead-onlyIdempotent
Extract the cited snippet from the AI engine's raw answer for each citation. Calls check_citations, then for each returned URL finds the first mention in raw_answer and returns a context window plus the nearest quoted span or containing sentence. Use to see why an engine cited a URL, not just that it did. Returns 'not found' for engines without raw_answer (Bing, Brave).
| Name | Required | Description | Default |
|---|---|---|---|
| query | Yes | Search query whose AI answer to extract citation evidence from. | |
| engine | No | AI engine to query. web_rank engines (bing_serp, brave_serp) lack raw_answer and return no evidence. | auto |
| max_results | No | Max citations to extract evidence for. | |
| context_chars | No | Half-width of the snippet window around each citation mention (chars). Total snippet is up to 2x this. |
Output Schema
| Name | Required | Description |
|---|---|---|
| note | No | |
| query | Yes | The query whose answer was analyzed. |
| engine | Yes | Engine used. |
| evidence | Yes | Per-citation evidence extracted from the raw answer. |
| fetched_at | Yes | UTC ISO-8601 timestamp. |
| evidence_found | Yes | Citations whose URL was located in the raw answer. |
| has_raw_answer | Yes | Whether the engine returned a raw answer. |
| citations_total | Yes | |
| raw_answer_chars | Yes | Length of the engine's raw answer. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover safety (readOnly, idempotent, openWorld). The description adds valuable behavioral detail beyond annotations: the internal call to check_citations, the window/quote extraction logic, and the engine-specific 'not found' behavior for Bing/Brave. Doesn't mention rate limits or performance, but adds meaningful context.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three dense sentences covering purpose, mechanism, and limitations. Front-loaded with the core action. Slightly packing mechanism + edge case into one sentence, but no wasted words.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given output schema exists (so return structure needn't be described), the description covers purpose, internal dependency, extraction logic, and key limitation (engine coverage). Complete for an agent to invoke correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% and each parameter has a clear description. The description reinforces the engine limitation ('engines without raw_answer') which aligns with the schema's enum description, adding no new semantic detail beyond what the schema provides. Baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Specific verb+resource (extract cited snippet from raw answer for each citation). Explicitly distinguishes from the sibling citations_check by stating it 'Calls check_citations, then...' and clarifies the why-vs-that distinction.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives a clear usage framing ('Use to see why an engine cited a URL, not just that it did') and notes an important limitation (returns 'not found' for Bing/Brave). However it doesn't explicitly contrast against the other citations_* siblings like citations_provenance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
citations_freshnessARead-onlyIdempotent
Score how recent the pages cited for a query are. Calls check_citations, then collects dateModified for each cited URL, returns a 0-100 recency_score (halflife=365d) plus per-URL freshness bucket (fresh/current/stale/ancient/unknown). Surfaces queries where AI cites old content - opportunity to ship fresher.
| Name | Required | Description | Default |
|---|---|---|---|
| query | Yes | Search query whose cited URLs to score for freshness. | |
| engine | No | AI engine to query for the citation set. | auto |
| max_results | No | How many cited URLs to inspect. |
Output Schema
| Name | Required | Description |
|---|---|---|
| note | No | |
| query | Yes | The query whose citations were scored. |
| engine | Yes | Engine used. |
| buckets | Yes | Freshness bucket distribution. |
| per_url | Yes | Per-URL freshness details. |
| fetched_at | Yes | UTC ISO-8601 timestamp. |
| recency_score | Yes | 0-100 average recency weight across cited URLs (halflife=365d). |
| average_days_old | Yes | Mean age in days across URLs with a detectable dateModified. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover the safety profile (readOnly, idempotent, openWorld, non-destructive), so the bar is lower, yet the description adds real value: it discloses the multi-step behavior (calls check_citations, then collects dateModified per URL), the scoring model (0-100, halflife=365d), and the bucket taxonomy (fresh/current/stale/ancient/unknown). It does not mention latency or cost of querying an AI engine, which is the main remaining gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three compact sentences, front-loaded with the action and scoring outcome, followed by the workflow and the motivation for using it. No filler and no repetition of schema fields.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return-value enumeration is not required, and the description still summarizes the headline outputs (score plus per-URL buckets). Combined with the annotations covering side-effect semantics and 100% schema coverage for inputs, an agent has everything needed to invoke this correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and the engine enum plus max_results bounds are fully documented in the schema, so the baseline is 3. The description adds no parameter-level detail beyond what the schema already carries (e.g., no guidance on picking an engine or tuning max_results).
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Names a specific verb (Score) and resource (recency of pages cited for a query), and distinguishes itself from neighbors: it delegates citation fetching to check_citations and its output is a freshness score, not a citation list. An agent can separate it from citations_check, citations_trend, and citations_provenance without opening a schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives a clear context for use ('Surfaces queries where AI cites old content - opportunity to ship fresher'), which tells the agent when this tool is the right pick. It does not name alternative tools or state exclusions, so it stops short of explicit when-not guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
citations_predictARead-onlyIdempotent
Score citation likelihood for a URL from public signals (Wikipedia link presence, schema.org markup, /llms.txt, GitHub and Reddit references, canonical hygiene, HTTPS). No LLM fired - all heuristic. Returns 0-100 score, grade, signal breakdown, and ranked fixes.
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | URL to score for citation likelihood. Must be absolute http(s). |
Output Schema
| Name | Required | Description |
|---|---|---|
| url | Yes | URL that was scored. |
| fixes | Yes | Ranked list of concrete improvements to raise the score. |
| grade | Yes | Letter grade (A-F) derived from the score. |
| score | Yes | 0-100 citation likelihood score. |
| signals | Yes | Per-signal boolean/numeric values used to compute the score. |
| fetched_at | Yes | UTC ISO-8601 timestamp. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover safety (readOnly, idempotent, non-destructive, openWorld), but the description adds genuinely new behavioral context: the entirely heuristic, no-LLM execution model and the shape of what comes back (0-100 score, grade, signal breakdown, ranked fixes). It still does not state latency or rate limits.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three compact sentences, no filler, with the core purpose front-loaded and the heuristic/return details following. Every clause conveys distinct information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a single-parameter, fully documented tool with an output schema, the description supplies the method and return shape needed to call it correctly; return-value detail is not required since an output schema exists. The only real gap is sibling routing, which is a usage concern rather than missing operational context.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Only one parameter with 100% schema description coverage, so the schema already carries the semantics (absolute http(s) URL). The description adds no constraint or format detail beyond that, making the baseline 3 appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (score) and resource (citation likelihood for a URL) and enumerates the public signals inspected, so the agent knows exactly what capability this is. It does not, however, distinguish itself from close siblings like citations_check or citations_evidence, so the boundary must be inferred.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No explicit when-to-use or when-not-to-use guidance, and the many citations_* siblings are never referenced. Usage is only implied by the phrase 'No LLM fired - all heuristic', which suggests a fast, cheap predictive pre-check rather than a full audit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
citations_provenanceARead-onlyIdempotent
Fan a query out across multiple AI engines and report per-URL cross-engine consensus. Returns each unique cited URL with the list of engines that cited it, plus a consensus_urls list (URLs cited by ALL engines). High engine_count = strong cross-engine citation signal; engine_count=1 = engine-specific.
| Name | Required | Description | Default |
|---|---|---|---|
| query | Yes | Search query to fan out across multiple engines. | |
| engines | No | Engines to query. If omitted, uses all LLM engines with a configured API key (perplexity, claude, openai, gemini, google_ai_mode). Include bing_serp/brave_serp only when you explicitly want web_rank comparison. | |
| max_results | No | Max citations per engine. |
Output Schema
| Name | Required | Description |
|---|---|---|
| note | No | |
| query | Yes | The query that was fanned across engines. |
| engines | Yes | Per-engine run summary. |
| per_url | Yes | All unique cited URLs sorted by cross-engine consensus (engine_count desc). |
| summary | Yes | |
| fetched_at | Yes | UTC ISO-8601 timestamp. |
| consensus_urls | Yes | URLs cited by ALL succeeding engines (requires >=2 engines). |
| engines_queried | Yes | |
| engines_succeeded | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, openWorldHint, and destructiveHint=false, so the safety profile is covered. The description adds real context beyond that: it names the engines queried by default, the return shape, and how to interpret engine_count. It does not discuss latency, rate limits, or cost of fanning out across engines, which would fully round it out.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, front-loaded with the core action and output shape, with the engine_count interpretation placed last as a useful reading aid. No filler sentences, though the trailing interpretation sentence slightly overlaps with output-schema territory.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only, idempotent fan-out tool with annotations covering safety, 100% schema coverage, and an output schema already present, the description is essentially complete. It even voluntarily explains return values despite the output schema existing, which covers the remaining ambiguity.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all three parameters (query, engines, max_results) are already documented in the schema with strong detail (enum list, maxItems, defaults, defaults-on-omission). The description adds nothing parameter-specific, so the baseline 3 for high coverage is correct.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific action (fan a query out across multiple AI engines) and resource (per-URL cross-engine citation consensus), and goes further to name the exact outputs (per-URL engine lists, consensus_urls). This is specific enough to distinguish it from sibling tools like citations_evidence or citations_check that lack the multi-engine consensus angle.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied rather than stated: the description explains engine_count semantics (1 = engine-specific, high = strong signal), which hints at when the tool is valuable, but it never names an alternative or states an explicit when/when-not condition against siblings like citations_evidence. No exclusions or routing guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
citations_trendARead-onlyIdempotent
Report citation rate over time for a panel from stored snapshots. Read-only; cache-only — makes no API calls to any AI engine and costs no API quota. Reads snapshot files from /snapshots//. Returns: snapshots[] (one entry per panel_run invocation, each with timestamp and citation_rate), plus per-query deltas (gained/lost/unchanged) comparing first vs last snapshot. Returns an empty series when no snapshots exist yet. No auth required. No rate limits. Use panel_run to accumulate snapshots first; use since to restrict the time window.
| Name | Required | Description | Default |
|---|---|---|---|
| panel | Yes | Panel name to report on. | |
| since | No | ISO date floor, e.g. '2026-01-01'. Only include snapshots on or after. |
Output Schema
| Name | Required | Description |
|---|---|---|
| panel | Yes | Panel name. |
| domain | No | Domain tracked by the panel. |
| series | Yes | Time-series of citation rates, one entry per snapshot. |
| snapshots | Yes | Number of snapshots available. |
| query_deltas | Yes | Per-query changes between first and last snapshot. |
| last_taken_at | No | Timestamp of the newest snapshot. |
| first_taken_at | No | Timestamp of the oldest snapshot. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Goes well beyond the annotations by disclosing that it is cache-only, makes no AI-engine API calls, consumes no API quota, requires no auth, and has no rate limits. It also clarifies the snapshot file location and the empty-series behavior when no snapshots exist, all of which materially affect calling decisions.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads purpose in the first clause, then layers behavioral and return details. Slight redundancy between 'cache-only... costs no API quota' and 'No auth required. No rate limits,' and the return-value enumeration partly overlaps the existing output schema, but overall it is dense and earns its length.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Complete for a read-only reporting tool: purpose, prerequisites, cache semantics, return shape, and edge case (empty series) are all covered. With an output schema present, the added interpretation of the deltas is a bonus rather than a required element.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so both parameters are already documented. The description adds only marginal meaning ('use since to restrict the time window'), restating what the schema says, so the baseline of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (report) and resource (citation rate over time), scoped clearly to 'a panel from stored snapshots'. This distinguishes it from siblings like citations_check (point-in-time) and citations_freshness, letting an agent differentiate without opening a schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly names the prerequisite sibling ('Use panel_run to accumulate snapshots first') and explains the role of the since parameter for narrowing the window. It gives clear usage context but does not state when-not to use it versus alternatives like citations_check or panel_track.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
competitors_canonical_setARead-onlyIdempotent
Fan a query across engines and aggregate citations by registered domain (not URL). Returns top competitor domains ranked by cross-engine consensus, with per-engine breakdown and top URLs per domain. Use to identify the canonical competitor set for a query - the domains every engine treats as authoritative.
| Name | Required | Description | Default |
|---|---|---|---|
| query | Yes | Search query to fan out across engines. | |
| top_n | No | Max competitor domains to return. | |
| engines | No | Engines to query. If omitted, uses all LLM engines with a configured API key (google_ai_mode, perplexity, claude, openai, gemini). Include bing_serp/brave_serp only for web_rank comparison. | |
| max_results | No | Max citations per engine. | |
| exclude_domains | No | Domains to filter out (e.g. your own brand, Wikipedia, Reddit). Suffix-match. |
Output Schema
| Name | Required | Description |
|---|---|---|
| note | No | |
| query | Yes | The query that was fanned across engines. |
| top_n | Yes | Maximum domains returned. |
| domains | Yes | Competitor domains ranked by cross-engine consensus. |
| engines | Yes | Per-engine run summary. |
| fetched_at | Yes | UTC ISO-8601 timestamp. |
| engines_queried | Yes | |
| excluded_domains | Yes | Registered domains that were filtered out. |
| engines_succeeded | Yes | |
| total_unique_domains | Yes | Total unique competitor domains found before top_n truncation. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, openWorldHint and destructiveHint=false, so the safety profile is covered. The description adds meaningful behavioral context beyond that: results are aggregated by registered domain, ranked by cross-engine consensus, and include per-engine breakdown plus top URLs. It does not mention latency or cost of fanning out across up to seven engines.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight sentences, front-loaded with the core action and its output shape, then the intended use case. No redundant restatement of the name or title and no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return-value explanation is not required, and the description still summarizes the return shape (consensus ranking, per-engine breakdown, top URLs). Parameter docs are fully covered by the schema. The only gap is practical guidance around multi-engine fan-out cost or latency for a tool with a 7-engine open-world dependency.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all five parameters (query, top_n, engines, max_results, exclude_domains) are already documented in the schema itself. The description reinforces the domain-level grouping concept but adds no syntax or default detail beyond the schema, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb+resource: 'Fan a query across engines and aggregate citations by registered domain (not URL).' The parenthetical explicitly disambiguates the aggregation granularity, and the response shape (ranked domains with per-engine breakdown) is stated so an agent can distinguish this from siblings like competitors_compare and competitors_compete.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
'Use to identify the canonical competitor set for a query - the domains every engine treats as authoritative' gives clear context for when to reach for this tool. However, it never names or excludes the closely related competitors_compare/competitors_compete siblings, leaving the agent to infer the boundary between them.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
competitors_compareARead-onlyIdempotent
Run predict_citation on 2-10 URLs and return a side-by-side signal table plus a list of signals where the URLs diverge. Use to compare your URL to top-cited competitors for the same query.
| Name | Required | Description | Default |
|---|---|---|---|
| urls | Yes | URLs to compare side-by-side. 2-10 URLs. One is typically yours and the rest are cited competitors. |
Output Schema
| Name | Required | Description |
|---|---|---|
| rows | Yes | Per-URL predict_citation rows (or { url, error } on failure). |
| fetched_at | Yes | UTC ISO-8601 timestamp. |
| diverging_signals | Yes | Signals where at least one URL differs from the others. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, openWorldHint, and destructiveHint=false, so the safety profile is fully covered by structured data. The description adds that the tool batches predict_citation across URLs and emits a divergence list, which is useful framing, but it discloses no auth, rate-limit, or cost characteristics of the batch call.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, zero filler, with the core operation front-loaded and the usage rationale following. Every clause earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be detailed, and the annotations carry the safety profile. The description covers purpose, usage, and output shape, which is sufficient for a single required parameter with full schema coverage, though sibling disambiguation is absent.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and the single parameter is fully documented in the schema, including the 2-10 bound and the 'one is typically yours, rest are competitors' convention. The description merely restates the 2-10 range, adding no syntax or format detail beyond the schema. Baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: runs predict_citation over 2-10 URLs and returns a side-by-side signal table plus divergent signals. This is clear and distinguishable from the single-URL citations_predict. It does not, however, explicitly differentiate itself from the sibling competitors_compete or competitors_canonical_set, so it stops short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The second sentence gives concrete when-to-use guidance: comparing your URL against top-cited competitors for the same query. It conveys the intended scenario well but names no alternative tool or exclusion condition, so there is no routing help beyond context.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
competitors_competeARead-onlyIdempotent
End-to-end competitive snapshot for a single query. Calls check_citations to get the cited URLs, then runs compare_domains on your_url vs the top cited competitors. Returns your score, the average competitor score, and the gap.
| Name | Required | Description | Default |
|---|---|---|---|
| query | Yes | Search query to test (what would a user ask an AI?). | |
| engine | No | AI engine to query for the citation set. 'auto' picks the first available key. | auto |
| your_url | Yes | Your URL to benchmark against the cited competitors. | |
| max_competitors | No | How many cited URLs to compare against your_url. Capped at 9 (compare_domains accepts max 10 URLs total including yours). |
Output Schema
| Name | Required | Description |
|---|---|---|
| query | Yes | The query that was tested. |
| engine | Yes | Engine used for the citation fetch. |
| your_url | Yes | Your URL that was benchmarked. |
| score_gap | Yes | your_score minus average_competitor_score. |
| comparison | Yes | Full compare_domains result. |
| fetched_at | Yes | UTC ISO-8601 timestamp. |
| your_score | Yes | predict_citation score for your URL (null on error). |
| competitors | Yes | Competitor URLs that were compared. |
| your_in_citations | Yes | Whether your URL appeared in the engine's citation list. |
| average_competitor_score | Yes | Mean score across competitor URLs. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, destructiveHint=false and openWorldHint, so the safety profile is covered. The description adds genuine behavioral value by disclosing that this triggers two sequential sub-calls and summarizing the returned aggregate (your score, competitor average, gap), which implies latency/cost characteristics.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight sentences: the deliverable first, the call sequence second, the return shape third. No filler and nothing buried.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With a full parameter schema, rich annotations, and an output schema, the description only needs to explain the composite behavior and it does. Minor gap: it doesn't hint at engine/failure behavior when a sub-call like check_citations has no key available, but that is largely covered by the 'auto' schema note.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and every parameter (query, engine, your_url, max_competitors) is documented in the schema, including the enum and the max-9 cap nuance. The description adds no parameter-level detail beyond what the schema already provides, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific deliverable ('end-to-end competitive snapshot for a single query') and names the exact sub-tools it orchestrates (check_citations, compare_domains), which distinguishes it from the plain single-purpose siblings like competitors_compare and citations_check.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The pipeline description makes the usage context clear: use it when you want a one-shot competitive snapshot rather than stitching the two underlying calls yourself. It doesn't state an explicit when-not or name the sibling to use instead, so it falls short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
domain_am_i_citedBIdempotent
Check whether a domain is cited by an AI engine across a cluster of queries. Returns per-query presence, rank, and a citation-rate summary. Use to measure visibility for a brand, product, or content site in AI search.
| Name | Required | Description | Default |
|---|---|---|---|
| domain | Yes | Domain to check, e.g. 'automatelab.tech' (without protocol). | |
| engine | No | LLM engine to check for citations. 'auto' runs all available LLM engines and returns per-engine breakdown + cross-engine consensus. Pin to a specific engine to reduce cost. 'bing_serp' and 'brave_serp' measure web rank, not LLM citations — use check_citations for those. | auto |
| queries | Yes | Queries to test the domain against. 1-20 queries per call. |
Output Schema
| Name | Required | Description |
|---|---|---|
| mode | Yes | Whether one or multiple engines were queried. |
| domain | Yes | The domain that was checked. |
| engine | No | Engine used (single_engine mode only). |
| engines | No | Per-engine summary rows (multi_engine mode). |
| results | No | Per-query results (single_engine mode). |
| summary | No | Aggregate summary (single_engine mode). |
| surface | No | Engine surface type (single_engine mode only). |
| consensus | No | Cross-engine consensus stats (multi_engine mode). |
| fetched_at | Yes | UTC ISO-8601 timestamp. |
| per_engine | No | Full per-engine detail (multi_engine mode). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare openWorldHint, idempotentHint, and destructiveHint=false, covering the safety profile. The description adds the return shape (per-query presence, rank, citation-rate summary), which is useful, but says nothing about cost, engine selection behavior, or why readOnlyHint is false for a 'check' operation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences with zero filler, front-loaded with the core action and return format. Efficient, though it spends a sentence on return values that the output schema already covers.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With an output schema present, return-value disclosure isn't strictly needed, and annotations carry the safety profile. The remaining gap is sibling differentiation among citations_check, domain_cited_for, and citations_* tools, which the description does not resolve.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema fully documents domain, engine, and queries, including the cost tradeoff of pinning an engine. The description adds no parameter detail beyond what the schema provides, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('check') and resource ('whether a domain is cited by an AI engine across a cluster of queries'), which distinguishes it from audit_* tools. However, it doesn't differentiate from close siblings like citations_check or domain_cited_for, so the agent can't tell them apart from the description alone.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The final sentence gives a use case ('measure visibility for a brand, product, or content site'), which implies when to reach for it. But there are no explicit when-not conditions and no named alternatives despite several look-alike siblings, leaving routing to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
domain_cited_forARead-onlyIdempotent
List queries that the given domain has been cited for, served from the local cache. Build up a corpus by calling check_citations or am_i_cited first; cited_for queries it without spending API budget.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | Maximum results. | |
| since | No | ISO date floor, e.g. '2026-01-01'. Only return entries fetched on or after this date. | |
| domain | Yes | Domain to look up, e.g. 'automatelab.tech'. | |
| engine | No | Filter by engine. Omit to include all. |
Output Schema
| Name | Required | Description |
|---|---|---|
| since | No | ISO date floor applied, if any. |
| total | Yes | Total entries returned. |
| domain | Yes | Domain that was looked up. |
| source | Yes | Always 'local_cache' — no external calls made. |
| results | Yes | Cache entries where this domain was cited. |
| engine_filter | No | Engine filter applied, if any. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations cover readOnly, idempotent, openWorld=false, and non-destructive. The description adds key behavioral context beyond annotations: it is served from local cache and does not spend API budget. This is valuable context not in annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences, front-loaded with what the tool does, followed by the prerequisite and cost-saving trait. No wasted words.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given an output schema exists and annotations cover safety, the description provides the key missing context: the tool reads from local cached corpus and is free. It doesn't mention pagination or return format, but those are covered by the output schema and parameter descriptions.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema documents all four parameters (limit, since, domain, engine). The description doesn't add parameter-level semantics beyond what the schema provides, so baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('List queries that the given domain has been cited for') with clear scope. It also names the sibling tools (check_citations, am_i_cited) that build the corpus, distinguishing it from related tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly says to call check_citations or am_i_cited first to build a corpus, and notes it 'queries it without spending API budget.' Clear usage context and prerequisite, though it doesn't explicitly state when not to use this tool.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
domain_cited_for_diffARead-onlyIdempotent
Diff cited_for between two time windows for a domain. Returns queries gained (cited now, not before baseline_until) and queries lost (cited before, not since current_since). Cache-only, no API spend. Use to track citation drift over time after publishing or migrating content.
| Name | Required | Description | Default |
|---|---|---|---|
| domain | Yes | Domain to diff, e.g. 'automatelab.tech'. | |
| engine | No | Filter by engine. Omit to include all. | |
| current_since | No | ISO date floor for the 'current' window. Defaults to baseline_until. | |
| baseline_until | Yes | ISO date (or ISO datetime). Baseline window = all cache entries fetched on or before this timestamp. |
Output Schema
| Name | Required | Description |
|---|---|---|
| lost | Yes | Queries lost (were cited, no longer are). |
| counts | Yes | |
| domain | Yes | Domain that was diffed. |
| gained | Yes | Queries gained (newly cited) since the baseline. |
| source | Yes | |
| fetched_at | Yes | UTC ISO-8601 timestamp. |
| current_since | Yes | Lower bound of the current window. |
| engine_filter | No | Engine filter applied, if any. |
| baseline_until | Yes | Upper bound of the baseline window. |
| unchanged_queries | Yes | Queries cited in both windows. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly, idempotent, non-destructive, closed-world, so the safety profile is covered. The description adds genuinely useful non-annotation context: 'Cache-only, no API spend' and the precise definition of what 'gained' and 'lost' mean relative to the windows.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, front-loaded with the core operation, then output semantics, then the cost/behavior note and usage hint. Every sentence earns its place with no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With an output schema present, return-value detail is not needed, and the description still supplies window semantics, cost behavior, and usage context. Nothing an agent needs to call it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so baseline is 3. The description goes beyond the schema by explaining how baseline_until and current_since define the two windows and what membership in each means, adding real semantic value over the field descriptions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (diff), resource (cited_for citations), and scope (two time windows for a domain), and names the gained/lost output semantics. An agent can distinguish it from the sibling domain_cited_for point-in-time lookup without opening either schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives an explicit context for use ('track citation drift over time after publishing or migrating content'), which implicitly routes the agent away from domain_cited_for. No explicit when-not condition or named alternative, so it falls short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
panel_runAIdempotent
Run a saved panel through am_i_cited and append a timestamped snapshot. Side effects: makes external API calls to the configured AI engine (costs API quota); writes one snapshot file to disk at /snapshots//.json. Requires at least one engine API key (same as am_i_cited). Returns per-query citation presence and a citation_rate summary for the run. Use panel_track to create a panel first; use citations_trend to read the accumulated trend after multiple runs.
| Name | Required | Description | Default |
|---|---|---|---|
| name | Yes | Panel name previously saved via track_queries. | |
| domain | No | Override the panel's default domain for this run. | |
| engine | No | AI engine to query. Use bing_serp/brave_serp for web_rank comparison only — am_i_cited will refuse them. | auto |
Output Schema
| Name | Required | Description |
|---|---|---|
| saved_to | Yes | Absolute file path of the snapshot that was written. |
| snapshot | Yes | The snapshot that was appended. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Well beyond the annotations, it discloses external API calls and quota cost, a concrete disk write path (<config>/snapshots/<panel>/<iso>.json), and the credential prerequisite (at least one engine API key, same as am_i_cited). The annotations already establish the write/openWorld/non-destructive profile, so this is additive rather than redundant. Minor nuance: appending a fresh timestamped file each run sits in mild tension with idempotentHint=true, but the tool's panel state is unchanged, so it is not a true contradiction.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four sentences, each carrying distinct payload (action, side effects, credentials, return shape, routing), with the core purpose front-loaded before the caveats. No filler or restated title text.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With an output schema present, the description needn't detail return structure, yet it still summarizes 'per-query citation presence and a citation_rate summary.' Combined with annotations and a fully documented 3-param schema, an agent has everything needed to call this correctly and cheaply.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so name, domain override, and the engine enum (including the serp restriction) are already documented in the schema; the description adds nothing parameter-specific. This is the expected baseline when the schema does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The opening sentence gives a specific verb and resource ('Run a saved panel through am_i_cited and append a timestamped snapshot'), immediately separating it from audit_* tools and from panel_track, which only creates panels. An agent can identify the operation and its scope without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explicitly routes the agent: 'Use panel_track to create a panel first; use citations_trend to read the accumulated trend after multiple runs,' naming both prerequisites and the downstream alternative. The engine note also states when serp engines are appropriate and that am_i_cited will refuse them.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
panel_trackAIdempotent
Save, load, or list named query panels. A panel is a persisted set of queries you want to monitor over time (e.g. editorial-watchlist). Use action=save with queries[] to create, action=load to read, action=list to enumerate. Panels live under /panels/.json.
| Name | Required | Description | Default |
|---|---|---|---|
| name | Yes | Panel name, e.g. 'editorial-watchlist'. Used to save and recall the query set. | |
| action | No | 'save' writes the panel, 'load' returns an existing panel, 'list' enumerates all panels. | save |
| domain | No | Default domain to track for this panel, e.g. 'automatelab.tech'. | |
| queries | No | Queries to save under this panel. Omit to read the existing panel. |
Output Schema
| Name | Required | Description |
|---|---|---|
| error | No | Error message when the panel was not found. |
| panel | No | The panel object that was saved or loaded. |
| saved | No | True when action=save succeeded. |
| panels | No | All panel names (action=list). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already carry a rich profile (readOnlyHint=false, idempotentHint=true, destructiveHint=false, openWorldHint=false), lowering the burden. The description adds the persistence path (<config>/panels/<name>.json), but never states whether action=save overwrites an existing panel — a meaningful omission for a write operation despite idempotentHint=true.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four tightly packed sentences: purpose, concept definition, per-action usage, and storage location. The scope is front-loaded and nothing is wasted.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return-value detail is unnecessary, and the description covers purpose, actions, and storage. The only real gap is the relationship to panel_run, which is absent but not strictly required to call the tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents name, action, domain, and queries, including the enum values. The description's action guidance is reinforcing rather than additive, so the baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description gives three specific verbs (save, load, list) applied to a clearly named resource (named query panels) and even defines what a panel is with a concrete example. It does not, however, distinguish itself from the sibling panel_run, which an agent could easily confuse for the execute-side of this feature.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It maps each action value to its purpose ('action=save with queries[] to create, action=load to read, action=list to enumerate'), which is genuine when-to-use guidance. It stops short of naming exclusions or pointing to panel_run as the alternative for running a panel.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
report_visibilityAIdempotent
Turnkey AI visibility report for a domain across a query set. Composes check_citations over every query (or a saved panel) and returns the metrics AI-visibility trackers sell as a dashboard, in one call: mention frequency (citation_rate), share_of_voice vs competitors, average rank when cited, and brand sentiment from the answer text. Side effects: one check_citations call per query (costs API quota for uncached queries; cached queries are free). Returns structured summary + top_domains + per_query, plus a rendered Markdown report (include_markdown=true) suitable for a public page. Provide queries[] or a panel name. Same engine selection as check_citations.
| Name | Required | Description | Default |
|---|---|---|---|
| panel | No | Name of a saved panel (see panel_track) to pull queries from. Provide this OR `queries`. | |
| domain | Yes | The domain you are measuring visibility for (e.g. automatelab.tech). | |
| engine | No | AI engine to query. 'auto' picks the first configured key. Same selection as check_citations. | auto |
| queries | No | Queries to run. Provide this OR `panel`. Each is sent to the AI engine via check_citations. | |
| brand_terms | No | Brand name variants to detect in answer text for sentiment (defaults to the domain's second-level label). | |
| competitors | No | Optional competitor domains to surface explicitly in the share-of-voice table. | |
| max_results | No | Max citations to pull per query. | |
| include_markdown | No | If true (default), include a rendered Markdown report under `markdown`. |
Output Schema
| Name | Required | Description |
|---|---|---|
| error | No | |
| summary | Yes | |
| markdown | No | Rendered Markdown report. Present when include_markdown=true. |
| per_query | Yes | |
| top_domains | Yes | Share-of-voice table, most-cited domains first. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Discloses concrete side effects beyond the annotations: one check_citations call per query, API quota cost for uncached queries with cached queries free. It also states the return shape (summary + top_domains + per_query + optional markdown), which the annotations alone don't convey.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the purpose before the mechanics, and every sentence carries information (metrics list, side effects, inputs, return shape). It is dense and somewhat long but with negligible filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a high-complexity aggregation tool with an output schema already present, the description covers inputs, cost semantics, and the composed-call behavior needed to invoke it correctly. Nothing essential is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the description adds real meaning: queries[] and panel are mutually exclusive alternatives, brand_terms defaults to the domain's second-level label, and engine selection mirrors check_citations. This goes beyond restating the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Turnkey AI visibility report for a domain across a query set') and enumerates exactly what it returns (mention frequency, share_of_voice, average rank, sentiment). It also clarifies its relationship to check_citations by describing itself as composing that tool over each query, which distinguishes it from the raw citation siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The 'one call' / 'dashboard' framing implies when to reach for it over individual checks, and it specifies the panel-vs-queries requirement. However, it never explicitly states when NOT to use it or names a sibling alternative (e.g. citations_check, competitors_compare) as the correct choice in a given scenario.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
signals_ai_overviewAIdempotent
Check whether Google shows an AI Overview for a query, and which URLs it cites. Uses SerpAPI (free tier: 100/month). Set SERPAPI_KEY.
| Name | Required | Description | Default |
|---|---|---|---|
| hl | No | Language code, default 'en'. | en |
| query | Yes | Search query to check for Google AI Overview. | |
| location | No | Location string, e.g. 'United States'. Affects AI Overview eligibility. |
Output Schema
| Name | Required | Description |
|---|---|---|
| query | Yes | The query checked. |
| cached | Yes | Whether the result was served from local cache. |
| sources | Yes | URLs cited in the AI Overview. |
| fetched_at | Yes | UTC ISO-8601 timestamp. |
| ai_overview_text | Yes | AI Overview text, if present. |
| ai_overview_present | Yes | Whether Google returned an AI Overview for this query. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Goes beyond annotations by disclosing the external dependency (SerpAPI), the required credential (SERPAPI_KEY), and a quota limit (free tier 100/month) that materially affects call frequency. Annotations already cover idempotency and openness; the quota and auth disclosures are exactly the added value the description should provide.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, front-loaded with the core purpose, followed immediately by the dependency and setup requirement. Zero filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be described, and the auth/quota prerequisites are stated. The only minor gap is that location's effect on AI Overview eligibility is left to the schema rather than reinforced in the description.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all three parameters (query, hl, location) are already documented with defaults and semantics. The description adds no parameter detail beyond the schema, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: check AI Overview presence and cited URLs for a query. This is clearly distinct from siblings like signals_answer_box or audit_schema, though it does not explicitly name which sibling to prefer. A near-complete purpose statement.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No when-to-use guidance and no alternative named. The only operational note is the SerpAPI dependency and free-tier quota, which is useful context but not usage routing. An agent must infer when this beats signals_answer_box or citations_check.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
signals_answer_boxARead-onlyIdempotent
Locate where each cited URL appears in the AI's raw answer text. Calls check_citations, finds the first mention of each citation's URL (or hostname) in raw_answer, and bins by char position into early/middle/late thirds. Surfaces whether your URL is cited up-front or buried near the end. Returns 'unknown' for engines without raw_answer (Bing, Brave).
| Name | Required | Description | Default |
|---|---|---|---|
| query | Yes | Search query whose AI answer to measure citation positions on. | |
| engine | No | AI engine to query. web_rank engines (bing_serp, brave_serp) lack raw_answer and will return position 'unknown'. | auto |
| max_results | No | Max citations to locate. |
Output Schema
| Name | Required | Description |
|---|---|---|
| note | No | |
| query | Yes | The query that was tested. |
| engine | Yes | Engine used. |
| buckets | Yes | Count of citations per position bucket. |
| positions | Yes | Per-citation position in the AI answer. |
| fetched_at | Yes | UTC ISO-8601 timestamp. |
| answer_chars | Yes | Total length of the engine's raw answer in characters. |
| citations_total | Yes | Total citations returned by the engine. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover readOnly/idempotent/non-destructive, and the description adds real behavioral context: the early/middle/late binning and the fact that engines lacking raw_answer (Bing, Brave) return 'unknown'. It stops short of describing the output shape or edge cases like multiple matches, but adds meaningful value beyond annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four tight sentences, front-loaded with the core purpose and followed by mechanism, value framing, and the key limitation. Every sentence earns its place with no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists so return values need no explanation, and the description covers purpose, mechanism, and the engine limitation thoroughly. Slightly short on routing guidance versus sibling citation tools, but otherwise complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the description adds engine-specific semantics ('engines without raw_answer... return position unknown') that reinforce the engine param's enum behavior in a way the schema only partially conveys.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Locate where each cited URL appears in the AI's raw answer text') and describes the mechanism (bins char position into early/middle/late thirds). It is clearly distinguishable from siblings like citations_check and signals_ai_overview.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides implied context (measuring citation position) and notes which engines yield 'unknown', but never states when to choose this over citations_check or signals_ai_overview, nor any exclusions beyond the engine caveat.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
signals_bing_gapARead-onlyIdempotent
Join Bing Webmaster Tools query stats with am_i_cited per query. Surfaces queries where the domain ranks well in Bing but is not cited in AI - the closest editorial wins. Bing's index backs Copilot/ChatGPT/Perplexity grounding, so a Bing rank gap is an LLM-citation gap. Requires BING_WEBMASTER_API_KEY (Bing Webmaster Tools -> Settings -> API Access).
| Name | Required | Description | Default |
|---|---|---|---|
| domain | Yes | Domain to analyze, e.g. 'automatelab.tech'. Used for the citation check. | |
| engine | No | AI engine for the citation check. | auto |
| queries | Yes | Queries to cross-reference. 1-20 per call. | |
| site_url | No | Verified Bing Webmaster site URL. Defaults to 'https://<domain>/'. Bing uses the https origin WITH a trailing slash, NOT the sc-domain: form GSC uses. |
Output Schema
| Name | Required | Description |
|---|---|---|
| rows | Yes | Per-query Bing rank + AI citation cross-reference. |
| domain | Yes | Domain analyzed. |
| engine | No | Engine used for the citation check. |
| site_url | Yes | Bing Webmaster siteUrl used (https origin with trailing slash). |
| closest_wins | Yes | Queries where domain ranks in Bing top-10 but is not AI-cited (the editorial gap). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly/openWorld/idempotent/non-destructive, and the description adds meaningful context beyond that: the required BING_WEBMASTER_API_KEY, where to obtain it, and a rationale linking Bing's index to Copilot/ChatGPT/Perplexity grounding. It does not discuss rate limits or result volume, and the output schema covers returns.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three front-loaded sentences with zero filler: purpose, the insight it yields, and the prerequisite. Every sentence carries information an agent needs to select and invoke the tool.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists so return values need not be explained, annotations cover the safety profile, and the schema fully documents all four parameters. The description supplies the remaining piece — the API-key prerequisite and its sourcing — making the definition complete for calling correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents domain, engine, queries, and site_url (including the trailing-slash/not-sc-domain subtlety). The description adds no parameter-level syntax or constraints beyond what the schema provides, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a concrete operation (joining Bing Webmaster query stats with am_i_cited per query) and the output it surfaces (queries ranking in Bing but not cited by AI). The 'Bing rank gap is an LLM-citation gap' framing makes it differentiable from the adjacent signals_gsc_gap sibling.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear context for use — finding editorial wins where Bing rank is strong but AI citation is absent — and implies this is a diagnostic gap-finding tool. It does not explicitly name alternatives like signals_gsc_gap or citations_check or state when not to use it, so it stops short of full when/when-not guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
signals_gsc_gapARead-onlyIdempotent
Join Google Search Console performance with am_i_cited per query. Surfaces queries where the domain ranks well in Google but is not cited in AI - the closest editorial wins. Requires GCP service account creds (credentials_path or GOOGLE_APPLICATION_CREDENTIALS env).
| Name | Required | Description | Default |
|---|---|---|---|
| domain | Yes | Domain to analyze, e.g. 'automatelab.tech'. Used both for the GSC site URL and the citation check. | |
| engine | No | AI engine for the citation check. | auto |
| queries | Yes | Queries to cross-reference. 1-20 per call. | |
| end_date | Yes | ISO date for GSC range end, e.g. '2026-05-01'. | |
| site_url | No | Override the GSC siteUrl. Defaults to 'sc-domain:<domain>'. | |
| start_date | Yes | ISO date for GSC range start, e.g. '2026-04-01'. | |
| credentials_path | No | Path to GCP service account JSON. Defaults to env GOOGLE_APPLICATION_CREDENTIALS. |
Output Schema
| Name | Required | Description |
|---|---|---|
| rows | Yes | Per-query GSC + AI citation cross-reference. |
| range | Yes | GSC date range. |
| domain | Yes | Domain analyzed. |
| engine | No | Engine used for the citation check. |
| site_url | Yes | GSC siteUrl used. |
| closest_wins | Yes | Queries where domain ranks in Google top-10 but is not AI-cited (the editorial gap). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly, idempotent, openWorld and non-destructive, so the safety profile is covered. The description adds a genuinely useful behavioral fact beyond that: it requires GCP service account credentials supplied via credentials_path or the GOOGLE_APPLICATION_CREDENTIALS env var, which affects whether the call can succeed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight sentences with zero padding: purpose first, then the value proposition, then the credential prerequisite. Front-loaded and each sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists so return values need not be described, and all seven parameters are covered by the schema. The description supplies purpose, value and auth prerequisite, leaving only sibling routing (vs signals_bing_gap) unaddressed.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so every parameter including the engine enum and date formats is already documented in the schema. The description only restates the credential source and the GSC/citation join, adding little semantics beyond what the schema provides. Baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: it joins GSC performance with citation data and surfaces queries where the domain ranks in Google but is not cited in AI. The phrase 'the closest editorial wins' adds a concrete framing of the output. It does not name the obvious sibling signals_bing_gap, so differentiation is left to inference.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The value framing ('closest editorial wins') implies the scenario where this is useful, but there is no explicit when-to-use or when-not, and no reference to the near-twin signals_bing_gap. The usage context is implied rather than spelled out.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
signals_wikipediaARead-onlyIdempotent
List Wikipedia articles that reference the given domain. Read-only. One HTTPS GET to the Wikipedia API (en.wikipedia.org/w/api.php?action=query&list=exturlusage). No auth required; no API keys; no rate limits beyond Wikipedia's public API fair-use policy (~1 request/second). Returns article titles and URLs. Wikipedia backlinks are the highest-lift signal for LLM training corpora — a domain cited from Wikipedia is far more likely to appear in AI training data and citation pools. Use lang to query non-English Wikipedias.
| Name | Required | Description | Default |
|---|---|---|---|
| lang | No | Wikipedia language subdomain, e.g. 'en', 'de', 'fr'. | en |
| limit | No | Maximum mention rows to return. | |
| domain | Yes | Domain to search for, e.g. 'automatelab.tech' (without protocol). |
Output Schema
| Name | Required | Description |
|---|---|---|
| lang | Yes | Wikipedia language subdomain used. |
| total | Yes | Number of Wikipedia articles referencing this domain. |
| domain | Yes | Domain that was searched. |
| mentions | Yes | List of Wikipedia articles that cite the domain. |
| fetched_at | Yes | UTC ISO-8601 timestamp. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Goes well beyond the annotations by disclosing the exact endpoint (en.wikipedia.org/w/api.php?action=query&list=exturlusage), the absence of auth/API keys, the external rate limit (~1 request/second fair-use), and the return shape (article titles and URLs). With annotations already covering the safety profile, this adds genuinely useful operational context.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Tight and front-loaded: the action comes first, then the implementation detail, then the rationale. The LLM-training-corpus sentence is longer than strictly necessary but carries decision-relevant motivation, so it largely earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return-value explanation is not required, yet the description still summarizes it. Combined with the endpoint, rate limit, auth status, and lang guidance, nothing an agent needs to invoke this correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so baseline is 3, but the description adds a usage hint for lang ('Use lang to query non-English Wikipedias') that is not present in the schema's own field description. It does not characterize limit, leaving that to the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a precise verb and resource ('List Wikipedia articles that reference the given domain'), which cleanly separates it from the broader signals_* and citations_* siblings. An agent can identify the tool's job without reading the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides clear context for why and when to reach for it ('highest-lift signal for LLM training corpora') and how to extend it with lang. However, it never names a sibling it competes with or states an exclusion, so it stops short of explicit alternative routing.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
49 tool updates
v0.1.2- Removed
ai_overview - Removed
am_i_cited - Removed
answer_box_position - Added
audit_crawler_access - Added
audit_llms_txt - Added
audit_schema - Changed
audit_sitemap1 field changed- changed
Output schema / (root)Previous value: -nullNew value: +{ + "$schema": "http://json-schema.org/draft-07/schema#", + "additionalProperties": false, + "properties": { + "audited": { + "description": "Number of URLs that were scored.", + "type": "number" + }, + "average_score": { + "description": "Mean predict_citation score across scored URLs.", + "type": "number" + }, + "errors": { + "description": "URLs whose audit threw an error.", + "items": { + "additionalProperties": false, + "properties": { + "error": { + "type": "string" + }, + "url": { + "type": "string" + } + }, + "required": [ + "url", + "error" + ], + "type": "object" + }, + "type": "array" + }, + "fetched_at": { + "description": "UTC ISO-8601 timestamp.", + "type": "string" + }, + "sitemap_url": { + "description": "The sitemap URL that was audited.", + "type": "string" + }, + "total_urls": { + "description": "Total URLs found in the sitemap.", + "type": "number" + }, + "worst_first": { + "description": "Up to 20 lowest-scoring URLs, worst first.", + "items": { + "additionalProperties": false, + "properties": { + "grade": { + "type": "string" + }, + "score": { + "type": "number" + }, + "signals": { + "additionalProperties": {}, + "type": "object" + }, + "top_fix": { + "description": "Top suggested fix for this URL.", + "type": "string" + }, + "url": { + "type": "string" + } + }, + "required": [ + "url", + "score", + "grade", + "signals" + ], + "type": "object" + }, + "type": "array" + } + }, + "required": [ + "sitemap_url", + "fetched_at", + "total_urls", + "audited", + "average_score", + "worst_first", + "errors" + ], + "type": "object" +}
- Added
audit_sitemap_map - Added
audit_structured_data - Removed
canonical_competitor_set - Removed
check_citations - Removed
citation_evidence - Removed
citation_freshness_score - Removed
citation_provenance - Removed
citation_trend - Added
citations_check - Added
citations_evidence - Added
citations_freshness - Added
citations_predict - Added
citations_provenance - Added
citations_trend - Removed
cited_for - Removed
cited_for_diff - Removed
compare_domains - Removed
compete_for_query - Added
competitors_canonical_set - Added
competitors_compare - Added
competitors_compete - Removed
crawler_access_audit - Added
domain_am_i_cited - Added
domain_cited_for - Added
domain_cited_for_diff - Removed
gsc_citation_gap - Removed
llms_txt_generator - Added
panel_run - Added
panel_track - Removed
predict_citation - Added
report_visibility - Removed
run_panel - Removed
schema_audit - Added
signals_ai_overview - Added
signals_answer_box - Added
signals_bing_gap - Added
signals_gsc_gap - Added
signals_wikipedia - Removed
sitemap_citation_map - Removed
structured_data_repair - Removed
track_queries - Removed
wikipedia_mentions
24 tool updates
v0.1.0- First observed
ai_overview - First observed
am_i_cited - First observed
answer_box_position - First observed
audit_sitemap - First observed
canonical_competitor_set - First observed
check_citations - First observed
citation_evidence - First observed
citation_freshness_score - First observed
citation_provenance - First observed
citation_trend - First observed
cited_for - First observed
cited_for_diff - First observed
compare_domains - First observed
compete_for_query - First observed
crawler_access_audit - First observed
gsc_citation_gap - First observed
llms_txt_generator - First observed
predict_citation - First observed
run_panel - First observed
schema_audit - First observed
sitemap_citation_map - First observed
structured_data_repair - First observed
track_queries - First observed
wikipedia_mentions
TDQS
Scored across 26 tools
Multiple tools overlap: citations_check, domain_am_i_cited, and report_visibility all measure citation presence; citations_provenance and competitors_canonical_set both aggregate cross-engine citations; audit_sitemap vs audit_sitemap_map and competitors_compare vs competitors_compete are easily confused. Descriptions clarify but require careful reading, making misselection likely.
Consistent snake_case with prefix categories (citations_, domain_, signals_, etc.), but within categories patterns mix verbs and nouns (citations_check vs citations_provenance), and domain_am_i_cited breaks the verb_noun convention. Readable but not fully predictable.
26 tools is heavy for a single server; many are variations on citation checking or auditing, indicating over-proliferation rather than a tightly scoped set. Could be consolidated.
Covers the domain broadly: checking, predicting, auditing crawlers/schema/sitemaps, competitor analysis, panels, trends, and signals from GSC/Bing/Wikipedia. Minor gaps like cache management or panel deletion, but core workflows are well supported.
Maintenance
Related MCP Connectors
Zotero MCP server for Claude and ChatGPT: search, citations, safe writes, PDF passages and pages.
Free web search for AI agents. No API key required. Hosted MCP in active development.
Free OpenAI-compatible inference with signed provenance receipts and 3 focused MCP tools.
Hosted MCP server connecting claude.ai, ChatGPT and other AI apps to your own computer
Related MCP Servers
- AlicenseAqualityCmaintenanceAutomateLab AI-SEO audits, scores, and rewrites web pages for AI citation eligibility, AEO and GEO. No API keys or registration. Works with Claude, Cursor, Codex, and other MCP clients. Product and documentation: https://automatelab.tech/products/mcp/ai-seo/2058 npm3MIT
- AlicenseNot gradedqualityCmaintenanceZero-auth multi-source research MCP server that enables web search, reading URLs, PDFs, GitHub repos, and querying Hacker News, Stack Overflow, Semantic Scholar, and YouTube transcripts without API keys.11Apache 2.0
- AlicenseAqualityAmaintenanceAI visibility tracker MCP server. Track brand citations across ChatGPT, Claude, Perplexity, Gemini & Google AI Overviews. Self-host on Cloudflare Workers. GEO/AEO.12807 npm46MIT
- AlicenseNot gradedqualityFmaintenanceMCP server enabling local-first web search, fetch, extract, and caching with citeable excerpts, no API key required. Supports research workflows for agents and apps.301 npm1MIT