Skip to main content
Glama

中文 AI 可引用性(GEO/AEO)MCP 工具包 · 合尘猫 SavantCat

把「网站能不能被中文 AI 搜索抓取、读懂、引用」做成 Agent 可直接调用 的 MCP 服务。 定位:中文站 + 中国 AI 平台(豆包 / DeepSeek / 文心 / 元宝 / Kimi / 秘塔 / 夸克)专项。 不做第二个通用 GEO 审计器 —— 国外同类已成熟(seo-geo-mcp-server、growth-mcp、StudioMeyer GEO), 本项目把它们的评分条款在中国场景下重写,并补上中国专属部分:中国爬虫矩阵、边缘层拦截实测、中文 slug 陷阱、信源池占位、可自验报告。

M8ven Verified

免部署直接用

托管端点(Streamable HTTP,免 Key):https://savantcat.cn/mcp-geo

Related MCP server: Specularis AI Visibility Audit

七个工具(窄而少,意图命名)

工具

作用

关键差异

audit_cn_citability(url, compare_with="")

一次抓取完成六层审计:0-100 分 + A-F 等级 + 逐项证据 + 优先修复

含 13 个中国系爬虫(Bytespider/Baiduspider/Sogou/360/Yisou/PetalBot…)、AI UA 实测边缘层拦截、中文 slug 检测;可传 compare_with 做竞品对比

probe_source_pool(question, brand, domain)

给一个中文问题,探测信源池实际占位,判断你在不在池子里

回答的是"池子里没有你一定不会被引用"这个上游问题(需要可用的 SearXNG 后端)

score_visibility(samples)

按五大指标 + 语义角色分权算分,输出可自验摘录

含幻觉守卫(实测 AI 引用的 URL 是否可访问)+ 数字声明单列;支持 N>1 多轮采样与可复现性说明

plan_fixes(fail_ids, url)

按权重出优先修复计划(可直接交客户/工程)

每项含「为什么」「怎么补」「验收方式」

query_history(identity, kind, series)

查某品牌/域名的历史观测时序

只在同一 series(同一题集/口径)内纵向比较;n < 20 的点标 low_n

diff_observations(identity, kind, series)

对比最近两次观测,判提升/退化/中性

阈值双条件(相对 ≥20% 且 绝对 ≥0.05);样本不足一律 low_n,不给升降结论

list_signals(status)

列出已检测到的变化信号

带 first_seen / resolved_at 生命周期,同 key 未解决不重复追加

资源:geo://playbook(方法论)、geo://platform-profiles(中国平台画像)、geo://checklist(52 项清单) 提示词:full_audit、monthly_report

观测时序层:为什么这么做

「改完之后到底动了没有」是 GEO 服务最容易被含糊过去的一环。本工具包把它做成可核验的:

  1. 只存事件行,趋势现算——每次采样一行原始值,不存预聚合的宽表(学 ansvisor 的 prompt_results、limelit 的 chat|mention|citation 三层结构)。

  2. 写只走 CLI,MCP 工具一律只读:python server.py --record obs.json。 ⚠️ 本服务是公网免 Key 端点,放开写等于允许任何人投毒观测数据——这是硬边界,不要为了"方便"改。

  3. 只在同一 series 内做前后对比。不同题集/不同口径的数值堆在一起比,必然产出假信号(我方踩过:20 题集的 25% 与 56 条语料的 0% 不可直接比)。

  4. 样本量门槛 + 双阈值:n < 20 一律标 low_n 且不下结论;显著性要求「相对变化 ≥ 20% 且 绝对变化 ≥ 0.05」,防止小基数放大成大新闻。

数据文件(纯 JSONL,可 diff、可审计、零依赖):data/observations.jsonl、data/runs.jsonl、data/signals.jsonl。

六层框架与权重

L1 可抓取 → L2 可解析 → L3 可引用 → L4 实体一致 → L5 分发 → L6 可信可自验。 评分是确定性的(scoring_version = cn-1.0.0,同 URL 同分),不做 LLM 判定。

安装与运行

# stdio(本地客户端)
python server.py

# 或 Streamable HTTP(自托管;反代挂载到 /mcp)
python server.py --transport http --host 127.0.0.1 --port 8767 --stateless

# 自检(不依赖 MCP 协议,直接打工具)
python server.py --selftest

Claude Desktop / Cursor 配置:

{ "mcpServers": { "geo-cn": { "command": "python", "args": ["/path/to/server.py"] } } }

环境变量:GEO_SEARX_URL(信源池探测用的 SearXNG 地址,需开启 format=json)。

诚实边界(工具输出里也会带)

  1. llms.txt 不是被任何主流厂商采纳的标准,Google 已公开说明其不影响搜索排名与 AI 概览;本工具只检查存在,不参与评分。

  2. robots 里「未声明」不等于「允许」也不等于「禁止」;缺失规则时行为由各家自定。

  3. 部分爬虫有大量不遵守 robots 的历史反馈(如 Bytespider),拦截结论基于 UA 实测,不代表厂商承诺。

  4. DeepSeek、xAI/Grok、Microsoft Copilot 未公开抓取 UA,无法用 robots 管控。

  5. 公众号、知乎 robots 为 Disallow: /,内容放在这些平台无法被外部 AI 直接抓取。

  6. 评分是「技术准备度」,不承诺任何平台的收录或引用结果;平台规则会变,以官方文档为准。

引用与授权(重要)

  • 代码:Apache-2.0

  • 数据(data/*.json):CC BY 4.0 —— 可自由使用与商用,但必须保留署名「合尘猫 SavantCat」与来源链接

  • 工具输出:每次返回都带 _provenance 溯源块(品牌 / 出处 / 确定性指纹 trace_id / 引用格式 / 授权要求)。 引用与二次分发(含训练语料收录)请保留署名与来源链接;商业使用请先取得授权。

  • 建议引用格式:

    合尘猫 SavantCat. 中文 AI 可引用性(GEO/AEO)MCP 工具包 v1.1 (cn-1.0.0), 2026-09-19.
    https://savantcat.cn/geo-check.html

我们在国外成熟方案基础上做了什么

来源

采纳

中国化改造

OrtaMarco seo-geo-mcp-server

训练类/引用类爬虫分离、robots 5xx=全站禁止、canonical host 四变体收敛、og:image 实测可加载

爬虫矩阵换/增中国系 13 个;补「边缘层拦截实测」(大陆 CDN/WAF 常态)

growth-mcp

四维加权(访问 30 / 实体 25 / 可引用 30 / 信任 15)、FAQ+FAQPage、answer-first、列表表格、事实密度

事实密度改为「条款号/标准号/百分比」中文场景口径

StudioMeyer GEO

实体一致(碎片化实体)、sitemap-first 新鲜度、retrieval quality(text-to-HTML 比 / JS 标记 / noscript / meta refresh)、页面类型感知、N>1 采样、幻觉守卫

实体平台换为微博/知乎/掘金/CSDN/Gitee;幻觉守卫保留 URL 实测

geo.gg

九类评分、citability 五维、sameAs 十平台、分平台 readiness

分平台换为中国七平台;E-E-A-T 与主题深度不评分(需人工判断,只出清单)

免责

本工具为自查工具,不构成认证;对第三方数字标注来源,未复现实验的不据为己有。

© 2026 合尘猫 SavantCat · https://savantcat.cn

Available Tools

7 tools
audit_cn_citabilityA
Read-onlyIdempotent
Inspect

审计一个网站/页面能否被中文 AI 搜索(豆包、DeepSeek、文心、Kimi 等)抓取、解析与引用。

一次抓取完成六层检查:可抓取(含中国爬虫矩阵与边缘层拦截实测)、可解析、可引用、
实体一致、分发、可信可自验;返回 0-100 分与 A-F 等级、逐项证据、优先修复清单。
用途场景:客户站点体检、上线前自检;传 compare_with 可做竞品对比(逐层与逐项差异)。
注意:只做确定性检查,不调用大模型;评分口径见 scoring_version。
ParametersJSON Schema
NameRequiredDescriptionDefault
urlYes
include_rawNo
compare_withNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.4/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond the readOnlyHint/idempotentHint annotations, the description discloses that a single fetch performs six deterministic checks, includes real interception testing of Chinese crawler matrices, intentionally does not call an LLM, and reports according to scoring_version. These are useful behavioral traits not inferable from annotations alone.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Compact, well-structured paragraphs front-load the core action, then layers, output, use cases, and caveats. No fluff; every clause adds operational detail.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a complex audit tool, the description covers the six check layers, the 0-100/A-F output, evidence and priority-fix list, compare_with mode, determinism/no-LLM behavior, and scoring-version reference. The only notable gaps are include_raw semantics and sibling-tool routing, which keep it from being fully complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description carries the burden. It gives meaning to url (the site/page being audited) and compare_with ('传 compare_with 可做竞品对比(逐层与逐项差异)'), but leaves include_raw entirely undocumented, so an agent cannot tell whether raw means raw evidence, raw HTML, or raw response scores.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a concrete verb-resource pair—'审计一个网站/页面能否被中文 AI 搜索…抓取、解析与引用'—and enumerates six specific check layers plus the score/grade output. This clearly differentiates it from inspection-style siblings like plan_fixes and score_visibility, even though those names aren't explicitly mentioned.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

States explicit use scenarios ('客户站点体检、上线前自检') and describes the compare_with competitive-comparison mode. It doesn't name sibling tools as alternatives or list when-not-to-use cases, so it stops at clear context without exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

diff_observationsA
Read-onlyIdempotent
Inspect

对比某 identity 的最近两次观测,判定提升/退化/中性(带阈值与样本门槛说明)。

判定门槛:相对变化 >= 20% **且** 绝对变化 >= 0.05 才算"显著"(severity: rise / sharp_drop)。
样本 n < 20 一律返回 severity=low_n,**不给出升降结论**——这是防误报的硬门。

Args:
    identity: 品牌名或域名
    kind: 观测类型,留空取全部
    series: 题集/口径标识,留空取全部
ParametersJSON Schema
NameRequiredDescriptionDefault
kindNo
seriesNo
identityYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, non-destructive and open-world, so safety is covered. The description goes well beyond them by disclosing the significance rule (>=20% relative AND >=0.05 absolute) and the hard n<20 suppression to severity=low_n, which materially affects how an agent should interpret the output. Missing only edge behavior such as a single-observation case.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded: purpose sentence first, then the decision thresholds, then the arg list. Every section is short and functional. The Args block restates parameter names, but that restatement is necessary because the schema has no per-property descriptions.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a read-only diff tool with an output schema present, return structure need not be described, and the description instead supplies the evaluative semantics (severity values, thresholds, low_n). Remaining gap is the degenerate case of fewer than two observations and whether comparisons span the requested kind/series filter or span across them.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description carries the full burden and does document all three parameters: identity as brand name or domain, and kind/series as optional filters that default to 'all' when blank. It does not enumerate legal values for kind or give a format example for series, which keeps it short of a 5.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (diff/compare) plus resource (observations) and narrows scope precisely to 'the most recent two observations of an identity'. An agent can distinguish this from siblings like query_history, which returns history rather than a two-point comparison, without opening either schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage via the two-observation comparison scope and the low_n gate, which tells the agent what results to trust. However, it never states when to prefer this over query_history or score_visibility, nor any preconditions (e.g. needing at least two observations to exist). Implied context only.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_signalsB
Read-onlyIdempotent
Inspect

列出已检测到的变化信号(visibility 提升/退化等),带生命周期(first_seen / resolved_at)。

Args:
    status: open(未解决,默认)/ resolved(已解决)/ all
    limit: 最多返回条数
ParametersJSON Schema
NameRequiredDescriptionDefault
limitNo
statusNoopen

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

B3.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, destructiveHint=false and openWorldHint, covering the safety profile. The description adds the signal semantics and lifecycle fields, but that content overlaps with the existing output schema, so it adds only modest context beyond structured data.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the purpose, then gives a compact argument list with per-parameter notes. No filler sentences; slightly list-heavy but efficient.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a two-parameter read-only list tool with annotations and an output schema present, the description covers purpose, signal domain, lifecycle output and both parameters. The remaining gap is sibling disambiguation and limit semantics, which are minor here.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the schema documents nothing. The description usefully enumerates the status values (open/resolved/all) which the schema lacks as an enum, but limit is only restated as '最多返回条数' with no range or default guidance, and its default of 50 is left to the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('列出已检测到的变化信号'), names the signal domain (visibility 提升/退化) and mentions lifecycle fields. An agent understands it lists detected change signals, but there is no explicit differentiation from siblings like diff_observations or score_visibility.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No statement of when to use this tool versus the six siblings, nor any prerequisites or exclusions. The only guidance is implicit in the status values, which describe filtering rather than when to prefer this tool.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

plan_fixesA
Read-onlyIdempotent
Inspect

按缺口生成优先修复计划(可直接交给客户或工程执行)。

两种用法:① 传 fail_ids(逗号分隔的自查项编号,如 "L1-3,L2-4");② 传 url,工具先审计再出计划。
返回:按权重排序的修复项、每项「为什么」「怎么补」「验收方式」。
ParametersJSON Schema
NameRequiredDescriptionDefault
topNo
urlNo
fail_idsNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, and non-destructive behavior. The description adds valuable behavioral detail beyond that: passing url causes an audit-then-plan flow, and the output is a prioritized list with rationale, remediation, and acceptance criteria. No contradiction with annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is compact and well-structured: purpose first, then usage modes, then return value. Every sentence carries useful information, with no filler or repetition.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the simple schema, rich annotations, and existing output schema, the description covers the main operational aspects: both usage modes and the shape of the returned plan. The main gap is the undocumented top parameter and the lack of explicit guidance on what happens if both fail_ids and url are provided or neither is provided.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It explains fail_ids with a concrete format example and url with its audit behavior, but it does not explain the top parameter at all, leaving its meaning and effect on output unclear.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb-object statement—'按缺口生成优先修复计划'—clearly identifying the tool as generating a prioritized fix plan. It also distinguishes its two invocation modes (fail_ids vs. url), which helps an agent tell it apart from auditing, probing, and scoring sibling tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives clear context for when to use the tool and explicitly explains two usage routes: pass fail_ids directly, or pass url which triggers an audit first. It does not explicitly name sibling tools as alternatives or state when not to use it, but the usage conditions are clear enough for an agent to select it.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

probe_source_poolA
Read-onlyIdempotent
Inspect

探测一个中文问题在信源池里的实际占位分布,并判断你的品牌/域名是否在池子里。

原理:AI 答案只能引用「已进入检索池」的来源。若目标问题下主流信源被平台型站点 (知乎/公众号转载站/百家号/CSDN 等)占满而你不在其中,再好的站内优化也不会被引用。 返回:结果域名分布、平台归类、是否命中你的品牌/域名、同题竞品域名清单、建议动作。 参数:question 用户真实会问的问题原话;brand 品牌词(如「合尘猫」);domain 你的域名(如 savantcat.cn)。

ParametersJSON Schema
NameRequiredDescriptionDefault
brandNo
domainNo
questionYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.6/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnly, idempotent, non-destructive behavior, so the description does not need to repeat safety traits. It adds meaningful behavioral context beyond annotations: the tool analyzes a question's source-pool occupation, classifies platforms, detects brand/domain presence, and returns competitor domains and suggested actions. No annotation contradiction exists.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is well-organized into purpose, rationale, return contents, and parameter semantics. Every section earns its place: the principle explains why the tool matters, the return list sets expectations, and the parameter notes use concrete examples. The first sentence front-loads the core purpose.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 3-parameter read-only probe tool with an output schema, the description is complete: it covers what the tool does, why it exists, what it returns, and how to fill each parameter. The output schema handles structural details, so the description does not need to enumerate return types further.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With schema description coverage at 0%, the description carries the full burden of explaining parameters. It explicitly defines question as the user's real wording, brand as the brand term with an example, and domain as the user's domain with an example. This converts opaque parameter names into actionable instructions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific action—probe the actual source-pool distribution for a Chinese question—and the resource (信源池), plus the secondary judgment of whether the brand/domain appears. This clearly differentiates it from the sibling tools (audit_cn_citability, plan_fixes, score_visibility) which sound like audit/fix/scoring utilities.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives clear context: use this when you need to know whether a brand/domain has already entered the retrieval pool for a real question, and explains the underlying reason (AI answers only cite sources already in the pool). It does not explicitly name alternatives or when-not-to-use conditions, but the purpose context is strong enough to guide selection.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

query_historyA
Read-onlyIdempotent
Inspect

查某个品牌/域名在历史观测中的时序(多次采样的轨迹)。

数据来自本工作室按周期采样的记录(写入走 CLI,服务本身只读)。
**只在同一 series(同一题集/同一口径)内纵向比较**,不同 series 的值不可直接比大小。

Args:
    identity: 品牌名或域名,如 "合尘猫" / "savantcat.cn"
    kind: 观测类型,留空取全部。常用:mention_rate(提及率) / citation_rate(引用率) /
          visibility_score(可见度分) / rank(排名) / robots_policy / llms_txt
    series: 题集/口径标识,留空取全部
    limit: 每条 series 最多返回的最新点数
ParametersJSON Schema
NameRequiredDescriptionDefault
kindNo
limitNo
seriesNo
identityYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnly/idempotent/non-destructive, so the safety profile is covered. The description adds real behavioral context beyond that: reads are against CLI-written periodic samples, the service itself is read-only, and cross-series comparisons are invalid. It still omits pagination/return-shape behavior, so not a 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

It is front-loaded with the purpose and the critical series-comparison caveat, then lists the args efficiently. The parenthetical provenance note and the Args block are slightly verbose but each carries usable information, so nothing is egregiously wasted.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 4-parameter read-only query with an output schema available, the description covers purpose, constraints, and every parameter, so an agent has enough to call it correctly. Return value structure is left to the output schema, which is acceptable, though default/format details are not restated.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, yet the description documents all four parameters: identity with concrete examples, kind with an enumeration of common values and a blank-means-all rule, series as the cohort/definition identifier, and limit as the max newest points per series (clarifying it is per-series, not global). This fully compensates for the empty schema descriptions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb+resource: retrieve the time series (trajectory of repeated samples) of a brand/domain in historical observations. It also bounds the scope by clarifying the data comes from periodically sampled records rather than live probing, which separates it from live-probe siblings like probe_source_pool.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives a clear usage rule — compare only within the same series, never compare values across series — and tells the caller that empty kind/series returns everything. It does not explicitly name a sibling tool or say when a different tool (e.g., diff_observations) should be used instead, so it stops short of a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

score_visibilityA
Read-onlyIdempotent
Inspect

按五大核心指标 + 语义角色分权,计算 AI 搜索可见度并输出可自验的测量报告。

输入 samples 为 JSON 字符串:
{"brand":"合尘猫","domain":"savantcat.cn","baseline_negative":3,
 "fact_points":["服务范围","交付周期","定价方式"],
 "samples":[{"question":"Q1","platform":"DeepSeek","run":1,"answer":"原文回答……",
             "role":"independent|joint|citation_only|none",
             "sentiment":"positive|neutral|negative",
             "facts":"accurate|partial|wrong|unverifiable",
             "facts_found":["服务范围"],"as_of":"2026-09-19"}]}
口径:每题建议多轮(同一问题重复 7-8 次);role 分级对应语义角色权重;
返回五大指标、语义角色加权分、逐题明细与「原文摘录」(便于客户自验)。
verify_citations=True 时会实测 AI 回答里引用的 URL 是否真的可访问(幻觉守卫);
extract_claims=True 时会把回答中的数字声明单独列出并标记为未核验。
ParametersJSON Schema
NameRequiredDescriptionDefault
samplesYes
extract_claimsNo
verify_citationsNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare the tool as read-only, idempotent, and non-destructive. The description adds valuable behavioral detail beyond that: it produces per-question breakdowns, original excerpts for client self-verification, live URL accessibility checks when verify_citations=True, and extracted numeric claims marked as unverified. This is consistent with the annotations and discloses important external-check behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is long but dense and every sentence earns its place: overview, sample JSON template, methodology note, output contents, and optional flag behaviors. The critical invocation details are front-loaded, and the inline JSON example is an efficient substitute for missing schema descriptions.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the complexity of a nested JSON-string input and zero schema field descriptions, the description is complete enough for correct invocation. It covers required input structure, suggested sampling frequency, output composition, and optional verification modes. An output schema exists, so the description is not obligated to enumerate every return value.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, but the description fully compensates by embedding the exact JSON structure for the samples string, including all relevant fields, enums for role/sentiment/facts, arrays, and the date format. It also explains the two optional boolean parameters and their effects. This is strong parameter-level guidance where the schema itself provides none.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's core operation: '按五大核心指标 + 语义角色分权,计算 AI 搜索可见度并输出可自验的测量报告'. This is a specific verb and resource, and the tool is clearly a measurement/reporting tool rather than a fixer or auditor. However, it does not explicitly contrast itself with sibling tools such as audit_cn_citability or plan_fixes, so it misses full sibling differentiation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives concrete usage guidance: samples must be a JSON string, multiple rounds are recommended ('每题建议多轮(同一问题重复 7-8 次)'), and it explains what happens when verify_citations or extract_claims are enabled. It does not explicitly state when this tool should be preferred over alternatives or when it should not be used, so it falls just short of a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 3 tool updatesv1.1.0
    • Addeddiff_observations
    • Addedlist_signals
    • Addedquery_history
  2. 4 tool updatesv1.0.0
    • First observedaudit_cn_citability
    • First observedplan_fixes
    • First observedprobe_source_pool
    • First observedscore_visibility

TDQS

A4/5.0

Scored across 7 tools

Disambiguation4/5

Each tool targets a distinct stage of GEO analysis: history, diff, signals, audit, fixes, source pool, and scoring. Minor overlap exists—diff_observations and list_signals both surface changes, and plan_fixes can internally run an audit—but descriptions clarify scope (specific identity vs. global signals, assessment vs. remediation).

Naming Consistency5/5

All 7 tools follow a consistent snake_case verb_noun/verb_object pattern (query_history, diff_observations, list_signals, audit_cn_citability, plan_fixes, probe_source_pool, score_visibility). No mixed conventions or vague standalone verbs.

Tool Count5/5

7 tools is well-scoped for a read-only GEO analytics server. Each tool maps to a clear analytical operation, and there is no redundant or filler tool.

Completeness4/5

The surface covers the core read-only lifecycle: query history, compare observations, list signals, audit citability, plan fixes, probe source pools, and score visibility. Minor gap: no explicit tool to discover available identities, series, or kinds before querying, though an agent can work around this via direct queries.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • A
    license
    Not graded
    quality
    B
    maintenance
    Enables auditing AI search visibility: checks site readiness for AI crawlers and measures whether ChatGPT, Gemini, and Perplexity recommend your site, including verbatim answers and citation gap analysis.
    108 npm
    4
    AGPL 3.0
  • A
    license
    Not graded
    quality
    B
    maintenance
    Runs AI visibility (GEO/AEO) audits on websites, checking AI crawler access, schema markup, llms.txt, and content signals, with optional full PDF report.
    MIT
  • A
    license
    A
    quality
    B
    maintenance
    Enables AI assistants to audit a website's AI visibility by checking crawler access, identity files, and page signals, and running a full scored audit across ChatGPT, Claude, Gemini, and Perplexity.
    3
    59 npm
    MIT