Skip to main content
Glama

DocuSky MCP

讓 Claude 直接查詢 DocuSky 數位人文學術研究平台 的資料庫。

打包成 .mcpb 之後,使用者雙擊就能安裝,不需要碰終端機。裝好就可以用日常語言問問題:

「DocuSky 上有哪些公開資料庫?」

「在 DaoBudMed6D 裡查『針灸』,看它在佛、道、醫三類文獻的分布」

「把《真誥》那一筆的全文調出來」


給使用者

需要準備什麼

  • Claude Desktop(macOS 或 Windows 桌面應用程式)

  • 不需要 DocuSky 帳號,也不需要會寫程式

網頁版(claude.ai)和手機 App 不支援擴充功能,必須用桌面版。 沒裝的話可以到 claude.ai/download 下載。


步驟一:下載

⬇️ 點此下載 docusky.mcpb

這個連結永遠指向最新版。想看歷史版本或更新說明,到 Releases 頁面

檔案大約 75 KB,副檔名是 .mcpb。瀏覽器可能會提示「不常下載的檔案類型」,選擇保留即可。


步驟二:安裝

三種方式擇一:

  • 雙擊下載好的 docusky.mcpb

  • 把檔案拖進 Claude Desktop 視窗

  • Claude Desktop 選單:設定 → 擴充功能 → 進階設定 → 安裝擴充功能…

會跳出安裝畫面,上面列出這個擴充功能的名稱、說明,以及它提供的七個工具。確認後點安裝。

接著會看到兩個欄位:DocuSky 帳號、DocuSky 密碼。

你的情況

怎麼填

只想查公開資料庫(大多數人)

兩個都留空,直接完成安裝

想查自己在 DocuSky 建的資料庫

填入你的 DocuSky 帳號與密碼

帳密之後隨時可以補填或修改,不必重裝。


步驟三:確認可以用了

設定 → 擴充功能,應該會看到 DocuSky 且開關是開啟狀態。

然後開一個新對話,試著問:

DocuSky 上有哪些公開資料庫?

第一次使用時,Claude 會詢問你是否允許它使用這個擴充功能的工具,選擇允許即可。

如果它列出了正統道藏、大明一統志、淡新檔案那些資料庫,就代表一切正常。


實際用起來是什麼樣子

你不需要記任何指令,直接用日常語言描述你想找什麼。

你問:

在 DaoBudMed6D 裡查「針灸」,看它在佛、道、醫三類文獻的分布

Claude 會去 DocuSky 查完,然後回給你:

文獻集

文件數

Medical 醫書

14

Buddhist 佛典

10

Daoist 道藏

8

共 32 筆。值得注意的是佛道兩家加起來(18)比醫書(14)還多——針灸的討論並不侷限在醫學文本裡。前幾筆命中的是《太上洞淵神咒經》《真誥》《抱朴子》,年代集中在 363–420 年間。

接著你可以繼續問:

把《真誥》那一筆的全文調出來

它就會把整篇文獻讀出來給你。文章很長的話會分段,你說「繼續」就好。

其他可以這樣問的例子:

  • 宋會要輯稿裡關於「市舶司」的記載有哪些?

  • 淡新檔案有沒有提到樟腦的案件?

  • 幫我比較「疫」和「癘」在道藏裡的分布差異


日後管理

你想做什麼

怎麼做

填入或修改 DocuSky 帳密

設定 → 擴充功能 → DocuSky → 設定

暫時停用

設定 → 擴充功能 → 把開關關掉

更新到新版

下載新的 .mcpb 再安裝一次,會直接覆蓋

移除

設定 → 擴充功能 → 解除安裝


它能做什麼

  • 全文檢索 —— 38 個公開資料庫(正統道藏、大明一統志、宋會要輯稿、朝鮮王朝實錄、淡新檔案、馬偕日記……)

  • 分布統計 —— 一個詞在不同文獻集、時代、地點的出現分布

  • 取全文 —— 單篇文獻完整讀出,長文自動分段

  • 標記分析 —— 統計文本裡的人名、地名、時間等標記

  • 私人資料庫 —— 填入帳密後也能查自己在 DocuSky 建的資料庫


檢索語法

平常用日常語言就好,需要精確控制時可以這樣講:

寫法

意思

針灸

全文檢索這個詞

醫 +方

必須同時包含「醫」和「方」

醫 -註

包含「醫」但排除「註」

.all

整個文獻集全部


⚠️ 不要把密碼貼在對話裡

對話內容會被保存。密碼請一律填在擴充功能的設定欄位,那是專門為此設計的,會經過安全儲存且不進入對話紀錄。

如果不小心貼了,建議去 DocuSky 改密碼。


遇到問題

設定裡找不到「擴充功能」

確認你用的是桌面版 Claude,不是瀏覽器開的 claude.ai。

裝好了但 Claude 說查不到 DocuSky

先重新啟動 Claude Desktop。再到設定 → 擴充功能確認 DocuSky 是開啟狀態。

擴充功能顯示啟動失敗

這個擴充功能以 Python 執行,需要系統上有 uv。 多數情況下 Claude Desktop 會自行處理,若確實失敗,可在終端機安裝後重啟 Claude:

curl -LsSf https://astral.sh/uv/install.sh | sh

查不到東西

古籍常有異體字。試試換字(「醫」vs「毉」)、拆成單字,或換一個資料庫。也可以直接請 Claude 幫你想替代詞。

查詢很慢

DocuSky 的全文檢索本來就需要時間,跨大型資料庫時等十幾秒是正常的。


Related MCP server: EKMS MCP Server

給開發者:怎麼打包

npm install -g @anthropic-ai/mcpb
mcpb pack . docusky.mcpb

產出約 75 KB。驗證 manifest:

mcpb validate manifest.json

自動建置

.github/workflows/build-mcpb.yml 會在 push 到 main手動觸發時:

  1. 驗證 manifest.json

  2. 打包 docusky.mcpb

  3. 解開並實際啟動一次,確認七個工具都在(不會連到 DocuSky)

  4. 上傳為 workflow artifact

  5. 發布 GitHub Release,tag 取自 manifest.jsonversion

Release 讓沒有 GitHub 帳號的人也能直接下載。版本號沒變而重跑時,會覆蓋既有的附件而不是失敗。

要發新版本就改 manifest.json 裡的 version,push 之後會自動建立對應的 Release。

專案結構

manifest.json           MCPB manifest(宣告工具、user_config、啟動方式)
server.py               進入點 shim
docusky_mcp/client.py   DocuSky Web API client
docusky_mcp/server.py   MCP 工具層
docusky_mcp/credentials.py  憑證讀取(環境變數優先)
pyproject.toml          相依套件定義
uv.lock                 鎖定版本,啟動時用 --frozen 安裝
.mcpbignore             打包時排除的檔案

server.py 存在的理由:uv 型別的 host 會直接執行進入點檔案,而 docusky_mcp/server.py 用的是相對匯入,直接跑會 ImportError。這個 shim 透過已安裝的套件轉一手。

執行環境

manifest.json 宣告 server.type: "uv",啟動指令是:

uv run --frozen --directory ${__dirname} server.py

依 MCPB 規格,uv 型別由 host 管理 Python 與相依套件。manifest_version 必須是 0.4 —— 0.3 的 schema 只接受 python | node | binary

憑證

user_config 的兩個欄位會被注入成環境變數:

欄位

環境變數

DocuSky 帳號

DOCUSKY_USERNAME

DocuSky 密碼(sensitive: true

DOCUSKY_PASSWORD

留空時兩者皆為空字串,credentials.py 會判定為未登入並退回公開模式。

其他可用的環境變數:

變數

預設值

用途

DOCUSKY_CREDENTIALS

~/.docusky/credentials.json

憑證檔位置(給非 Claude Desktop 的 MCP 客戶端用的後備)

DOCUSKY_BASE_URL

https://docusky.org.tw/DocuSky/webApi

API 根路徑

DOCUSKY_TIMEOUT

120

單次請求逾時(秒)

提供的工具

工具

用途

list_databases

列出可用資料庫

list_corpora

某資料庫下的文獻集與篇數

search_documents

全文檢索,回傳書目與摘錄

get_document

取單篇全文

post_classification

分布統計

tag_analysis

標記統計

check_login

檢查登入狀態

search_documents 刻意不回全文——DocuSky 單篇動輒六千字以上,一次二十筆會塞爆 context。要讀全文請用 get_document,帶入該筆的 n,並沿用同一組 db / query / corpus / page_size


注意事項

  • 本專案使用 DocuSky 的 Web API(docusky.org.tw/DocuSky/webApi/)。該 API 沒有公開文件也沒有版本號,DocuSky 改版時可能需要跟著更新。

  • 部分端點會在 JSON 前面夾帶 PHP 警告訊息,client 有做容錯。

  • 命中數是文件數,不是詞頻。

  • 少數資料庫的「分類」欄位在 DocuSky 端本身就是亂碼,不要據以推論。

  • 部分資料庫仍在建構中,查不到內容是正常的。

  • 請遵守 DocuSky 服務與使用規範,並節制查詢頻率。

Available Tools

7 tools
check_loginA

Report whether DocuSky credentials are configured and whether they work.

Public databases need no login; run this only when private ("USER") databases are unreachable.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the behavioral burden and does well: it frames the tool as a passive diagnostic that reports status rather than mutating anything, and confines its relevance to a specific failure condition. It does not mention possible side effects or latency of the credential check, but the report-only framing adequately signals a safe read.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences, front-loaded with what it does and followed immediately by the condition for running it. No filler or repetition of schema content.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need no explanation, and with zero parameters the only remaining need is invocation guidance, which the description fully supplies. Nothing an agent requires to call this correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so there is nothing for the description to disambiguate beyond the schema. Baseline 4 applies with no parameter gaps.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific operation (report credential configuration and validity) against a specific resource (DocuSky credentials). It is clearly distinguishable from the data-oriented siblings like list_databases and search_documents.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives explicit when-not-to-use ("Public databases need no login") and a precise trigger for when to run it ("only when private ('USER') databases are unreachable"). This is a textbook scoping rule an agent can act on without inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_documentA

Fetch the full text of one search hit, identified by its position in the result set.

result_number is the global 1-based rank of the document for this query — the n field of a search_documents hit. Pass the same db, query, corpus and page_size that produced it, or the ranking will not line up.

Long documents are returned in slices: raise offset by max_chars to page through the body.

Args: db: Database title. result_number: 1-based rank of the document within the query's results. query: The same query string used in search_documents. corpus: The same corpus used in search_documents. target: "OPEN" or "USER". page_size: The same page_size used in search_documents. offset: Character offset into the document body. max_chars: Maximum characters of body text to return. include_raw_xml: Also return the untouched DocuXML content. owner_username: Owner of a friend-shared database, when applicable.

ParametersJSON Schema
NameRequiredDescriptionDefault
dbYes
queryNo.all
corpusNo[ALL]
offsetNo
targetNoOPEN
max_charsNo
page_sizeNo
result_numberYes
owner_usernameNo
include_raw_xmlNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations exist, so the description carries the full burden, and it does disclose the key traits: results are a single hit, long bodies come back in slices, and paging is done by raising offset by max_chars. It also notes include_raw_xml returns the untouched DocuXML. It leaves auth/permission behavior (e.g. what OPEN vs USER target implies) implicit.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the essential constraint (matching search parameters) before the parameter list, and the pagination mechanic is given its own compact sentence. The Args block duplicates some parameter names already visible in the schema, but overall it is tight and earns its length.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return-value documentation is not needed, and the description correctly focuses on the coupling to search_documents and on slice-based paging. The one gap is that permission/access semantics for target and owner_username are left unstated.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% across 10 params, yet the Args block explains every one, with real added meaning for result_number ('global 1-based rank…the n field of a search_documents hit'), offset ('character offset into the document body'), max_chars, include_raw_xml and target ('OPEN' or 'USER'). Some entries are thin restatements (db = 'Database title', page_size = 'the same one'), which keeps this short of a 5.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Fetch the full text of one search hit') and immediately delimits the retrieval mode to a single ranked result, which cleanly separates it from the sibling search_documents. The identifying key (position in the result set) is named in the first sentence.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly states the precondition that db, query, corpus and page_size must match those that produced the hit, 'or the ranking will not line up', and names search_documents as the producing tool. It does not positively state when not to use this tool, but the prerequisite guidance is strong.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_corporaA

List the corpora inside one database, with document counts.

Args: db: Database title, exactly as returned by list_databases. target: "OPEN" or "USER". include_friend_db: Include databases shared by friends (USER only).

ParametersJSON Schema
NameRequiredDescriptionDefault
dbYes
targetNoOPEN
include_friend_dbNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A3.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden. 'List' implies a non-destructive read and it discloses the conditional constraint that include_friend_db applies only to USER target, but it says nothing about authorization needs, scope of returned data, or pagination behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

One summary sentence plus a compact Args block; front-loaded with the operation and scope and nothing redundant. The structure is efficient though not especially polished.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given an output schema exists, return-value detail is unnecessary, and the description adequately covers all three parameters and their interplay. The only material gap is sibling routing and any environment/auth preconditions.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, so the description must compensate and largely does: it gives the accepted values for 'target' (OPEN/USER), explains include_friend_db's effect and its USER-only restriction, and defines 'db' precisely. It does not restate the schema defaults, which is fine.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (List) and resource (corpora) and scopes it to 'inside one database, with document counts', which separates it from list_databases at the parent level. It does not explicitly name a sibling it is not, so it falls short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is implied through the prerequisite that 'db' must be 'exactly as returned by list_databases', which tells the agent the correct call ordering. However, there is no explicit when-to-use/when-not or routing against siblings like search_documents or list_databases.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_databasesA

List DocuSky databases available to this server.

Args: target: "OPEN" for public databases, "USER" for the logged-in account's own. include_friend_db: Also list databases shared by DocuSky "friends" (USER only).

ParametersJSON Schema
NameRequiredDescriptionDefault
targetNoOPEN
include_friend_dbNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A3.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden. It conveys read-only listing semantics and hints at authentication via 'logged-in account', but does not state permission requirements, friend-sharing implications, or result limits. Partial disclosure only.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded one-line purpose followed by tight per-argument notes; no filler. The 'Args:' block is compact and each line earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With only two optional parameters, an output schema covering return values, and explicit parameter meanings, an agent has what it needs to invoke the tool. Minor gaps in auth/permission context remain but are low-impact.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate, and it does: it explains the meaning of both the 'target' enum-like values ('OPEN' vs 'USER') and the 'include_friend_db' flag including its USER-only constraint. This is meaningfully beyond the bare schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource ('List DocuSky databases available to this server') with a clear scope. It is distinguishable from the closest sibling list_corpora by naming a different resource type, though it does not explicitly contrast against it.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It implies when to use each mode by explaining that target='OPEN' yields public databases and 'USER' yields the account's own, which is useful selection guidance. However, it gives no explicit when-to-use versus alternatives or prerequisites beyond that.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

post_classificationA

Break a query's hits down by DocuSky's post-classification facets.

Returns, per facet (corpus, period, place, category…), a distribution of [value, document count, hit count] — the quickest way to see how a term is spread across a database.

Args: db: Database title. query: Search terms, or ".all" for the whole database. corpus: Corpus title, or "[ALL]". target: "OPEN" or "USER". owner_username: Owner of a friend-shared database, when applicable.

ParametersJSON Schema
NameRequiredDescriptionDefault
dbYes
queryNo.all
corpusNo[ALL]
targetNoOPEN
owner_usernameNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A3.9/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden. It does disclose the return shape (per-facet distribution of [value, document count, hit count]), which is meaningful. However, it omits read-only/permission behavior, error conditions, and what the facet set actually contains beyond a vague ellipsis.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Purposes are front-loaded, followed by return format and an arg list. The structure is clean and each section pulls weight, with only mild redundancy in the arg explanations.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Because an output schema exists, the description needn't explain return values, yet it helpfully summarizes them. Parameters are all covered. It stops short of permission requirements or sibling routing, but for a facet-distribution tool this is nearly sufficient.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, so the description must carry parameter meaning — and it does, documenting all five: db, the '.all' query sentinel, corpus with '[ALL]', target accepting 'OPEN'/'USER', and owner_username for friend-shared databases. That is strong compensation, though enum-style constraints are only hinted, not formalized.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The opening sentence gives a specific verb and resource — breaking a query's hits down by post-classification facets — and goes further by naming the facets (corpus, period, place, category). It distinguishes itself from sibling retrieval tools like search_documents or tag_analysis by framing the output as a distribution rather than a hit list.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The phrase 'the quickest way to see how a term is spread across a database' implies when the tool is valuable, but no alternative tool is named and no explicit when-not/exclusion is given. Usage is implied rather than stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

search_documentsA

Full-text search one DocuSky database; returns metadata plus a short excerpt per hit.

Full document text is deliberately omitted — call get_document with a hit's n (and the same db/corpus/query/page_size) to read one in full.

Args: db: Database title. query: Search terms. +term requires, -term excludes, .all matches everything. corpus: Corpus title, or "[ALL]" for every corpus in the database. page: 1-based page number. page_size: Hits per page (1-100). target: "OPEN" or "USER". excerpt_chars: Characters of body text to preview per hit; 0 disables excerpts. owner_username: Owner of a friend-shared database, when applicable.

ParametersJSON Schema
NameRequiredDescriptionDefault
dbYes
pageNo
queryNo.all
corpusNo[ALL]
targetNoOPEN
page_sizeNo
excerpt_charsNo
owner_usernameNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the burden and does disclose key behavior: results are excerpt-only, excerpts are controlled by excerpt_chars (0 disables them), and pagination is 1-based with a 1-100 page_size bound. It does not state permission/auth requirements or read-only nature beyond the friend-shared `owner_username` hint, so it stops short of full transparency.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The core behavior and the get_document hand-off are front-loaded in the first two sentences, followed by a compact structured Args block. Every line carries information an agent needs; nothing is restated from the name or schema.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For an 8-parameter search tool with an output schema available, the description supplies everything else needed to call it correctly – query syntax, corpus/db scoping, pagination, excerpt control, and the pointer to get_document for full text. Nothing material is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate and it does: query syntax is documented ('+term' requires, '-term' excludes, '.all' matches everything), corpus accepts '[ALL]', target is 'OPEN' or 'USER', page is 1-based, page_size is 1-100, and owner_username is scoped to friend-shared databases. All 8 parameters gain meaning absent from the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Full-text search one DocuSky database') plus the return shape ('metadata plus a short excerpt per hit'), which immediately separates it from get_document and the listing siblings. An agent can identify the tool's role without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It explicitly routes the agent: full text is omitted here, so 'call get_document with a hit's `n` (and the same db/corpus/query/page_size)' to read one in full. That is a concrete when-to-use-this vs when-to-use-the-alternative condition rather than implied guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

tag_analysisA

Summarize the DocuXML tags (people, places, dates, custom markup) in a query's hits.

Only meaningful for databases whose documents carry inline tagging; returns an empty result otherwise.

Args: db: Database title. query: Search terms, or ".all" for the whole database. corpus: Corpus title, or "[ALL]". target: "OPEN" or "USER". owner_username: Owner of a friend-shared database, when applicable.

ParametersJSON Schema
NameRequiredDescriptionDefault
dbYes
queryNo.all
corpusNo[ALL]
targetNoOPEN
owner_usernameNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A3.8/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations exist, so the description carries the full behavioral burden. It usefully discloses the empty-result outcome for untagged databases, but says nothing about permissions, whether filtering by owner_username requires a share relationship, or what the summary contains.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the purpose in one sentence, then the key caveat, then a compact Args list. Nothing is padded, though the Args block is essentially a plain list rather than prose that adds interpretation.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return-value detail is not required. With five parameters at 0% schema coverage, the glosses plus the inline-tagging caveat cover most of what an agent needs; only the meaning of target and the sharing/permission model for owner_username are left implicit.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate and largely does: it glosses db, query (including the '.all' sentinel), corpus (including '[ALL]'), target's two allowed values, and owner_username's conditional relevance. The target values remain opaque ('OPEN' vs 'USER' is not explained), leaving a small gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource combination ('Summarize the DocuXML tags ... in a query's hits') and enumerates what counts as a tag (people, places, dates, custom markup). An agent can distinguish it from list_databases or search_documents by resource, though no sibling is named explicitly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The second sentence gives a real exclusion condition: 'Only meaningful for databases whose documents carry inline tagging; returns an empty result otherwise.' That tells the agent when not to expect value, but it never names an alternative tool for untagged databases or non-tag analyses.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 7 tool updatesv0.1.0
    • First observedcheck_login
    • First observedget_document
    • First observedlist_corpora
    • First observedlist_databases
    • First observedpost_classification
    • First observedsearch_documents
    • First observedtag_analysis

TDQS

A4.2/5.0

Scored across 7 tools

Disambiguation5/5

Each tool has a clearly distinct purpose: listing databases, listing corpora, checking login, searching documents, fetching full text, post-classification breakdown, and tag analysis. No two tools overlap significantly; the workflow is well-defined with search_documents and get_document having complementary roles.

Naming Consistency5/5

All tool names follow a consistent snake_case verb_noun pattern (list_databases, list_corpora, check_login, search_documents, get_document, post_classification, tag_analysis). The pattern is predictable and readable throughout.

Tool Count5/5

Seven tools form a well-scoped, cohesive set that covers the core operations for exploring and analyzing DocuSky databases without redundancy. Each tool earns its place in the workflow.

Completeness4/5

The toolset covers the essential lifecycle: discovery (list databases/corpora), authentication (check_login), search (search_documents), retrieval (get_document), and analysis (post_classification, tag_analysis). While read-only operations are comprehensive, there is no support for modifying or creating databases/corpora, which may be a deliberate scope limitation.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Resources

Unclaimed servers have limited discoverability.

Looking for Admin?

If you are the server author, to access and configure the admin panel.

Related MCP Connectors

Related MCP Servers

  • F
    license
    A
    quality
    D
    maintenance
    Enables Claude to search, query, and interact with an Enterprise Knowledge Management System (EKMS). Supports semantic search, knowledge recommendations, relationship graphs, and feedback recording for enterprise knowledge bases.
    7
    -
  • F
    license
    Not graded
    quality
    Not graded
    maintenance
    Enables Claude to interact with Google Docs to list, read, create, search, and update documents in a user's Google Drive. It provides a suite of tools and prompts for document management and content analysis using OAuth 2.0 authentication.
    881
    -