Skip to main content
Glama

DocuSky MCP

讓 Claude 直接查詢 DocuSky 數位人文學術研究平台 的資料庫。

打包成 .mcpb 之後,使用者雙擊就能安裝,不需要碰終端機。裝好就可以用日常語言問問題:

「DocuSky 上有哪些公開資料庫?」

「在 DaoBudMed6D 裡查『針灸』,看它在佛、道、醫三類文獻的分布」

「把《真誥》那一筆的全文調出來」


給使用者

需要準備什麼

  • Claude Desktop(macOS 或 Windows 桌面應用程式)

  • 不需要 DocuSky 帳號,也不需要會寫程式

網頁版(claude.ai)和手機 App 不支援擴充功能,必須用桌面版。 沒裝的話可以到 claude.ai/download 下載。


步驟一:下載

⬇️ 點此下載 docusky.mcpb

這個連結永遠指向最新版。想看歷史版本或更新說明,到 Releases 頁面。

檔案大約 82 KB,副檔名是 .mcpb。瀏覽器可能會提示「不常下載的檔案類型」,選擇保留即可。


步驟二:安裝

三種方式擇一:

  • 雙擊下載好的 docusky.mcpb

  • 把檔案拖進 Claude Desktop 視窗

  • Claude Desktop 選單:設定 → 擴充功能 → 進階設定 → 安裝擴充功能…

會跳出安裝畫面,上面列出這個擴充功能的名稱、說明,以及它提供的九個工具。確認後點安裝。

接著會看到兩個欄位:DocuSky 帳號、DocuSky 密碼。

你的情況

怎麼填

只想查公開資料庫(大多數人)

兩個都留空,直接完成安裝

想查自己在 DocuSky 建的資料庫

填入你的 DocuSky 帳號與密碼

帳密之後隨時可以補填或修改,不必重裝。


步驟三:確認可以用了

到 設定 → 擴充功能,應該會看到 DocuSky 且開關是開啟狀態。

然後開一個新對話,試著問:

DocuSky 上有哪些公開資料庫?

第一次使用時,Claude 會詢問你是否允許它使用這個擴充功能的工具,選擇允許即可。

如果它列出了宋會要輯稿、大明一統志、淡新檔案那些資料庫,就代表一切正常。


實際用起來是什麼樣子

你不需要記任何指令,直接用日常語言描述你想找什麼。

你問:

在 DaoBudMed6D 裡查「針灸」,看它在佛、道、醫三類文獻的分布

Claude 會去 DocuSky 查完,然後回給你:

文獻集

文件數

Medical 醫書

14

Buddhist 佛典

10

Daoist 道藏

8

共 32 筆。值得注意的是佛道兩家加起來(18)比醫書(14)還多——針灸的討論並不侷限在醫學文本裡。前幾筆命中的是《太上洞淵神咒經》《真誥》《抱朴子》,年代集中在 363–420 年間。

接著你可以繼續問:

把《真誥》那一筆的全文調出來

它就會把整篇文獻讀出來給你。文章很長的話會分段,你說「繼續」就好。

其他可以這樣問的例子:

  • 宋會要輯稿裡關於「市舶司」的記載有哪些?

  • 淡新檔案有沒有提到樟腦的案件?

  • 幫我比較「疫」和「癘」在道藏裡的分布差異


日後管理

你想做什麼

怎麼做

填入或修改 DocuSky 帳密

設定 → 擴充功能 → DocuSky → 設定

暫時停用

設定 → 擴充功能 → 把開關關掉

更新到新版

下載新的 .mcpb 再安裝一次,會直接覆蓋

移除

設定 → 擴充功能 → 解除安裝


它能做什麼

  • 全文檢索 —— 38 個公開資料庫(本草經集注、大明一統志、宋會要輯稿、朝鮮王朝實錄、淡新檔案、馬偕日記……)

  • 分布統計 —— 一個詞在不同文獻集、時代、地點的出現分布

  • 取全文 —— 單篇文獻完整讀出,長文自動分段

  • 標記分析 —— 統計文本裡的人名、地名、時間等標記

  • 文字雲 —— 把詞頻畫成 DocuSky 官方 WordCloudLite 的文字雲(同一份數字也能切換成泡泡圖、Top-10 長條圖、表格)

  • 地圖 —— 把地點(地名+WGS84 經緯度,可加時間、說明)整理成 DocuSky 官方 DocuGIS2 的匯入 TSV,貼進去就能看時空分布、時間軸、路線、熱區

  • 雙維度交叉統計(實驗性)—— 同時交叉兩個分類維度,DocuSky 伺服器端功能尚未完整,目前測試過的組合都會被拒絕

  • 私人資料庫 —— 填入帳密後也能查自己在 DocuSky 建的資料庫

內嵌顯示 DocuSky 網頁

在支援 MCP Apps 的 Claude 版本裡,查公開資料庫時(全文檢索、分布統計、標記分析)以及畫文字雲時,Claude 的回覆旁邊會直接內嵌顯示 DocuSky 官方網頁本身的畫面,不只是文字結果。畫面上方有一排這個擴充自己加的工具列:

  • 上一頁/下一頁/跳頁 —— 直接翻 DocuSky 的檢索結果。請用這排按鈕翻頁,不要用 DocuSky 頁面自己那排頁碼:它靠整頁跳轉(window.location.href)換頁,而那在對話內嵌的沙箱 iframe 裡會翻出空白頁。

  • 全螢幕 —— 把內嵌畫面放大成整頁再看,關掉就回到對話裡。

  • 瀏覽器開啟 —— 在真正的瀏覽器打開同一頁,DocuSky 的全部功能(篩選、標記統計、匯出、另開單篇)都在那裡。

查詢類的內嵌只在公開資料庫(不需帳密)上生效——私人資料庫的登入狀態是伺服器端的,內嵌畫面沒辦法一起帶過去,所以私人查詢仍然只會拿到文字結果。文字雲不受這個限制:它畫的是已經拿到手的數字,資料從哪個資料庫來都可以。如果你用的 Claude 版本不支援 MCP Apps,這一切照常運作,只是不會出現內嵌畫面。

地圖(docugis_map)的內嵌長得不一樣:DocuGIS2 沒有辦法用網址帶資料進去,所以內嵌畫面上半是整理好的 TSV 加一顆「複製 TSV」,下半是 DocuGIS2 本體,最後一步要你自己貼上——按複製、在下方地圖左上角的 ⇥ 打開選單、點「1.2 匯入資料 Import Data」、在右邊的框貼上、按[匯入 Import]。貼完之後時間軸、路線、群聚、熱區、匯出都是 DocuGIS2 原生的功能。

少數 client 會用最嚴格的沙箱(沒有 allow-same-origin)來跑內嵌畫面,那種環境下 DocuSky 只有第一頁能正常顯示。擴充會自己偵測到並改成提示你按「瀏覽器開啟」,不會給你一排按了只會變空白的翻頁鈕。DocuGIS2 在那種沙箱裡更直接:整頁卡在轉圈載不起來,所以擴充在偵測到時乾脆不嵌它,只給你 TSV 和「瀏覽器開啟 DocuGIS2」。


檢索語法

平常用日常語言就好,需要精確控制時可以這樣講:

寫法

意思

針灸

全文檢索這個詞

醫 +方

必須同時包含「醫」和「方」

醫 -註

包含「醫」但排除「註」

.all

整個文獻集全部


⚠️ 不要把密碼貼在對話裡

對話內容會被保存。密碼請一律填在擴充功能的設定欄位,那是專門為此設計的,會經過安全儲存且不進入對話紀錄。

如果不小心貼了,建議去 DocuSky 改密碼。


遇到問題

設定裡找不到「擴充功能」

確認你用的是桌面版 Claude,不是瀏覽器開的 claude.ai。

裝好了但 Claude 說查不到 DocuSky

先重新啟動 Claude Desktop。再到設定 → 擴充功能確認 DocuSky 是開啟狀態。

擴充功能顯示啟動失敗

這個擴充功能以 Python 執行,需要系統上有 uv。 多數情況下 Claude Desktop 會自行處理,若確實失敗,可在終端機安裝後重啟 Claude:

curl -LsSf https://astral.sh/uv/install.sh | sh

查不到東西

古籍常有異體字。試試換字(「醫」vs「毉」)、拆成單字,或換一個資料庫。也可以直接請 Claude 幫你想替代詞。

查詢很慢

DocuSky 的全文檢索本來就需要時間,跨大型資料庫時等十幾秒是正常的。


Related MCP server: EKMS MCP Server

給開發者:怎麼打包

npm install -g @anthropic-ai/mcpb
mcpb pack . docusky.mcpb

產出約 82 KB。驗證 manifest:

mcpb validate manifest.json

自動建置

.github/workflows/build-mcpb.yml 會在 push 到 main 或手動觸發時:

  1. 驗證 manifest.json

  2. 打包 docusky.mcpb

  3. 解開並實際啟動一次,確認九個工具都在(不會連到 DocuSky)

  4. 上傳為 workflow artifact

  5. 發布 GitHub Release,tag 取自 manifest.json 的 version

Release 讓沒有 GitHub 帳號的人也能直接下載。版本號沒變而重跑時,會覆蓋既有的附件而不是失敗。

要發新版本就改 manifest.json 裡的 version,push 之後會自動建立對應的 Release。

專案結構

manifest.json           MCPB manifest(宣告工具、user_config、啟動方式)
server.py               進入點 shim
docusky_mcp/client.py   DocuSky Web API client
docusky_mcp/server.py   MCP 工具層
docusky_mcp/ui.py       MCP Apps:ui:// 檢視器資源與 webUrl 產生邏輯
docusky_mcp/credentials.py  憑證讀取(環境變數優先)
pyproject.toml          相依套件定義
uv.lock                 鎖定版本,啟動時用 --frozen 安裝
.mcpbignore             打包時排除的檔案

server.py 存在的理由:uv 型別的 host 會直接執行進入點檔案,而 docusky_mcp/server.py 用的是相對匯入,直接跑會 ImportError。這個 shim 透過已安裝的套件轉一手。

執行環境

manifest.json 宣告 server.type: "uv",啟動指令是:

uv run --frozen --directory ${__dirname} server.py

依 MCPB 規格,uv 型別由 host 管理 Python 與相依套件。manifest_version 必須是 0.4 —— 0.3 的 schema 只接受 python | node | binary。

憑證

user_config 的兩個欄位會被注入成環境變數:

欄位

環境變數

DocuSky 帳號

DOCUSKY_USERNAME

DocuSky 密碼(sensitive: true)

DOCUSKY_PASSWORD

留空時兩者皆為空字串,credentials.py 會判定為未登入並退回公開模式。

其他可用的環境變數:

變數

預設值

用途

DOCUSKY_CREDENTIALS

~/.docusky/credentials.json

憑證檔位置(給非 Claude Desktop 的 MCP 客戶端用的後備)

DOCUSKY_BASE_URL

https://docusky.org.tw/DocuSky/webApi

API 根路徑

DOCUSKY_TIMEOUT

120

單次請求逾時(秒)

提供的工具

工具

用途

list_databases

列出可用資料庫

list_corpora

某資料庫下的文獻集與篇數

search_documents

全文檢索,回傳書目與摘錄

get_document

取單篇全文

post_classification

分布統計

tag_analysis

標記統計

twodim_analysis

雙維度交叉統計(實驗性,見下方說明)

word_cloud

把詞頻畫成 DocuSky WordCloudLite 文字雲

docugis_map

把地點整理成 DocuGIS2 地圖的匯入 TSV

check_login

檢查登入狀態

search_documents 刻意不回全文——DocuSky 單篇動輒六千字以上,一次二十筆會塞爆 context。要讀全文請用 get_document,帶入該筆的 n,並沿用同一組 db / query / corpus / page_size。

twodim_analysis 包的是 DocuSky 2026-01-28 才加上的 getQueryTwodimAnalysisJson.php。實際測試(2026-09-11)發現不管 dim1/dim2 帶什麼組合(包括直接沿用 post_classification 回傳的 facet 代碼,如 COMP/TP1)都會被回覆 {"code": 1, "message": "Currently not support ..."} 拒絕,DocuSky 自己的前端 JS 也還沒接上這個功能的 UI。先留著這個工具,等 DocuSky 補完後不用再改 client 端。

word_cloud 不自己統計也不自己查詢,它只負責把已經有的數字畫出來: tag_analysis 的標記次數、post_classification 的分布(value / docCount)、 或 Claude 從 get_document 全文自行數出來的詞頻,傳成 {"針灸": 120, "湯液": 48} 這種 term → 次數的對照表即可。

這個工具包的是 DocuSky 的 WordCloudLite。 讀過它的原始碼(2026-09-12)後確定,它只吃兩個 URL 參數:url=<某個.csv> 和 data=<詞,值;詞,值;...>。其餘設定(標題、背景色、隱藏控制列……)只能透過 postMessage 傳,而它的 handler 會擋掉所有非 docusky.org.tw 的 origin,內嵌用的 沙箱 iframe 過不了這關;url= 又需要一個公開可讀的 CSV,stdio MCP server 沒地方放。 所以走 data=。實作上有兩個坑,都已實測確認:

  • data= 傳進去的值在頁面裡仍然是字串(它的 parser 只做 v.split(',')), 但畫圖那段用的是 d3.max(data, d => d.value),而 d3 v5 對字串是字典序比較。 像 100 / 90 / 9 這組,最大值會變成 "9",所有字級跟著爆掉,畫面全白。 解法是把每個值補零到同樣位數("090" < "100"),字典序就跟數值序一致, 後面的 d.value / maxValue 對補零字串也算得出正確比例。

  • DocuSky 的 Apache 對 ~15 KB 的網址回 200、~24 KB 回 414,所以產生的網址會從 最小的詞開始砍,砍到長度安全為止。

另外 WordCloudLite 在詞數多的時候會刻意隨機取樣一部分來畫(見它的 plotWordCloud),所以想讓每個詞都出現,大約傳 20–40 個詞最穩。

docugis_map 同樣不查詢、不做地理編碼:座標要由呼叫端給(使用者提供、 search_documents 的 placeInfo、地名資料庫,或 Claude 自己確定知道的地點), 不確定的地點就不要編——寧可留空或問使用者。它把 rows 整理成 DocuGIS2 匯入用的 TSV(id name x y date text 加上任何額外欄位,x 是 WGS84 經度、y 是緯度), 回傳 tsv 與 webUrl。

DocuGIS2 的匯入器比想像中挑(2026-09-12 逐項實測,細節寫在 docusky_mcp/ui.py 的註解裡):date 欄只要有一格是空的或只寫年份(1887),那一列就會被整列丟掉, 但整個 date 欄不存在時每一列都進得去。所以 build_docugis_tsv 會把純年份補成 YYYY-01,而且在只有部分資料有日期時預設不輸出 date 欄(地圖完整、沒有時間 軸),要時間軸就傳 include_dates=true,代價是沒日期的那幾列不會出現。1887-01、 1887/1/1、18870101、-0200-01-01(西元前)都吃得下。

DocuGIS2 也是這個專案裡唯一不能用網址餵資料的 DocuSky 工具:它 56 支 script 沒有一支讀 location.search,也沒有註冊 message 監聽器,所以 WordCloudLite 那套 ?data= 在這裡完全不成立(它認得的 index.html?f=<id> 要先把 資料寫進一個共用的公開帳號、資料就此公開,所以刻意不用)。剩下能走的就是它的貼上 框,這也是為什麼地圖的 ui:// 資源是另一份 HTML(DOCUGIS_HTML)。

search_documents、post_classification、tag_analysis 這三個工具,在 target="OPEN" 時回傳的 JSON 裡多了一個 webUrl 欄位,指向 docusky.org.tw 上對應的查詢頁(webApi/webpage-open-3in1.php,用 spType 切換成一般搜尋 / 分布統計 / 標記分析檢視)。支援 MCP Apps (docusky_mcp/ui.py)的 client 會把這個 URL 用一個沒有外部依賴、手刻 postMessage 協定的小型 ui:// 資源嵌成 iframe 顯示;不支援的 client 就只是多一個可以忽略或當連結用的欄位,行為與加這個功能前完全一樣。word_cloud 回傳的 webUrl 走的是同一個 ui:// 資源與同一條 CSP(frameDomains 已經涵蓋 docusky.org.tw)。

手刻那份 postMessage 有三個地方是照著 ext-apps 的 schema 與實測結果來的 (都寫在 docusky_mcp/ui.py 的模組 docstring 裡):

  • ui/initialize 的 params 三個欄位都是必填:appInfo、appCapabilities、 protocolVersion。0.3.0 只送了 appCapabilities,會驗證 params 的 host 會直接擋掉交握,結果就是什麼都不顯示——0.3.1 修好了這點。

  • DocuSky 內建的頁碼是整頁跳轉,在沙箱 iframe 裡會翻成空白,所以翻頁改成由 工具列把 &page=N 直接設進 iframe 的 src(這條路實測可行)。

  • 內嵌畫面預設高度太矮,所以會用 ui/notifications/size-changed 要 720px, 並用 ui/request-display-mode 提供全螢幕切換。

驗證方式是照 ext-apps 的 examples/basic-host(外層 proxy iframe + 內層 sandbox iframe 的雙層架構)架一個假 host,對真正的 docusky.org.tw 跑過 allow-scripts allow-same-origin allow-forms 與只有 allow-scripts 兩種沙箱。 target="USER" 時 webUrl 會是 null——DocuSky 用伺服器端的 session cookie 認證私人資料庫,內嵌用的瀏覽器分頁沒有那個 cookie,硬塞連結只會顯示 「未登入」,所以私人查詢乾脆不給這個欄位。


注意事項

  • 本專案使用 DocuSky 的 Web API(docusky.org.tw/DocuSky/webApi/)。該 API 沒有公開文件也沒有版本號,DocuSky 改版時可能需要跟著更新。

  • 部分端點會在 JSON 前面夾帶 PHP 警告訊息,client 有做容錯。

  • 命中數是文件數,不是詞頻。

  • 少數資料庫的「分類」欄位在 DocuSky 端本身就是亂碼,不要據以推論。

  • 部分資料庫仍在建構中,查不到內容是正常的。

  • 請遵守 DocuSky 服務與使用規範,並節制查詢頻率。

Available Tools

10 tools
check_loginA

Report whether DocuSky credentials are configured and whether they work.

Public databases need no login; run this only when private ("USER") databases are unreachable.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the behavioral burden and does well: it frames the tool as a passive diagnostic that reports status rather than mutating anything, and confines its relevance to a specific failure condition. It does not mention possible side effects or latency of the credential check, but the report-only framing adequately signals a safe read.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences, front-loaded with what it does and followed immediately by the condition for running it. No filler or repetition of schema content.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need no explanation, and with zero parameters the only remaining need is invocation guidance, which the description fully supplies. Nothing an agent requires to call this correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so there is nothing for the description to disambiguate beyond the schema. Baseline 4 applies with no parameter gaps.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific operation (report credential configuration and validity) against a specific resource (DocuSky credentials). It is clearly distinguishable from the data-oriented siblings like list_databases and search_documents.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives explicit when-not-to-use ("Public databases need no login") and a precise trigger for when to run it ("only when private ('USER') databases are unreachable"). This is a textbook scoping rule an agent can act on without inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

docugis_mapA

Put places on a map in DocuSky's own DocuGIS2 tool.

Takes coordinates you already have — from the user, from a placeInfo in search_documents, from a gazetteer, or from your own knowledge of where a place is — and returns the TSV that DocuGIS2's import box accepts. This tool does no geocoding and no searching: never invent coordinates for a place you are unsure of; leave that row out, or ask the user.

DocuGIS2 cannot be fed data through a URL, so the last step belongs to the user: MCP Apps-capable hosts show the TSV with a copy button next to the embedded tool, and the user pastes it into DocuGIS2's box and presses 匯入 (the rendered map then does time filtering, routes, clustering, heatmaps and export). Other hosts get the same TSV as text plus webUrl, to paste into DocuGIS2 in a browser. Tell the user that paste step is theirs.

Args: rows: One entry per place. name plus x (WGS84 longitude) and y (latitude) are required; date, text and any other keys are optional and become extra columns in DocuGIS2's popups. Common aliases are understood (lng/lon/經度 -> x, lat/緯度 -> y, 地名 -> name). Example: [{"name": "鹿港", "x": 120.434, "y": 24.057, "date": "1784-01-01", "text": "鹿港開港"}]. title: A name for this map, shown above the tool. include_dates: Leave unset to decide automatically. DocuGIS2 drops any row whose date cell is empty or a bare year, so a half-dated set ships without the date column (every place on the map, no time axis). Pass true to keep the timeline and lose the undated rows, false to drop dates entirely. max_rows: Keep at most this many places (default 2000).

ParametersJSON Schema
NameRequiredDescriptionDefault
rowsYes
titleNo
max_rowsNo
include_datesNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations exist, so the description carries the full burden and does so richly: it discloses it performs no geocoding/searching, that DocuGIS2 cannot be fed via URL, the host-dependent return path (embedded TSV with copy button vs. text + webUrl), and the non-obvious rule that DocuGIS2 silently drops rows whose date is empty or a bare year.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with purpose, then sources, then the delivery/paste workflow, then an Args section. It is long, but with 0% schema coverage and no annotations nearly every sentence is load-bearing, and the Args block is cleanly scoped per parameter.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, yet the description still usefully frames what comes back (TSV plus webUrl) and who completes the final step. For a tool whose real action happens in an external UI, nothing an agent needs to call it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, and the description compensates thoroughly: it documents required fields (name, x, y), optional fields (date, text, arbitrary keys becoming popup columns), alias handling (lng/lon/經度, lat/緯度, 地名), a concrete example, and the exact semantics of include_dates and max_rows.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Put places on a map in DocuSky's own DocuGIS2 tool') and immediately frames what it produces (the TSV DocuGIS2's import box accepts). The explicit 'does no geocoding and no searching' boundary cleanly separates it from siblings like search_documents.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explains exactly when to reach for it and where coordinates come from (user, placeInfo from search_documents, gazetteer, own knowledge), and gives a clear when-not: never invent coordinates, leave the row out or ask the user. It also names the downstream alternative path (MCP Apps hosts vs. other hosts) and the human paste step.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_documentA

Fetch the full text of one search hit, identified by its position in the result set.

result_number is the global 1-based rank of the document for this query — the n field of a search_documents hit. Pass the same db, query, corpus and page_size that produced it, or the ranking will not line up.

Long documents are returned in slices: raise offset by max_chars to page through the body.

Args: db: Database title. result_number: 1-based rank of the document within the query's results. query: The same query string used in search_documents. corpus: The same corpus used in search_documents. target: "OPEN" or "USER". page_size: The same page_size used in search_documents. offset: Character offset into the document body. max_chars: Maximum characters of body text to return. include_raw_xml: Also return the untouched DocuXML content. owner_username: Owner of a friend-shared database, when applicable.

ParametersJSON Schema
NameRequiredDescriptionDefault
dbYes
queryNo.all
corpusNo[ALL]
offsetNo
targetNoOPEN
max_charsNo
page_sizeNo
result_numberYes
owner_usernameNo
include_raw_xmlNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations exist, so the description carries the full burden, and it does disclose the key traits: results are a single hit, long bodies come back in slices, and paging is done by raising offset by max_chars. It also notes include_raw_xml returns the untouched DocuXML. It leaves auth/permission behavior (e.g. what OPEN vs USER target implies) implicit.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the essential constraint (matching search parameters) before the parameter list, and the pagination mechanic is given its own compact sentence. The Args block duplicates some parameter names already visible in the schema, but overall it is tight and earns its length.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return-value documentation is not needed, and the description correctly focuses on the coupling to search_documents and on slice-based paging. The one gap is that permission/access semantics for target and owner_username are left unstated.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% across 10 params, yet the Args block explains every one, with real added meaning for result_number ('global 1-based rank…the n field of a search_documents hit'), offset ('character offset into the document body'), max_chars, include_raw_xml and target ('OPEN' or 'USER'). Some entries are thin restatements (db = 'Database title', page_size = 'the same one'), which keeps this short of a 5.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Fetch the full text of one search hit') and immediately delimits the retrieval mode to a single ranked result, which cleanly separates it from the sibling search_documents. The identifying key (position in the result set) is named in the first sentence.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly states the precondition that db, query, corpus and page_size must match those that produced the hit, 'or the ranking will not line up', and names search_documents as the producing tool. It does not positively state when not to use this tool, but the prerequisite guidance is strong.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_corporaA

List the corpora inside one database, with document counts.

Args: db: Database title, exactly as returned by list_databases. target: "OPEN" or "USER". include_friend_db: Include databases shared by friends (USER only).

ParametersJSON Schema
NameRequiredDescriptionDefault
dbYes
targetNoOPEN
include_friend_dbNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A3.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden. 'List' implies a non-destructive read and it discloses the conditional constraint that include_friend_db applies only to USER target, but it says nothing about authorization needs, scope of returned data, or pagination behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

One summary sentence plus a compact Args block; front-loaded with the operation and scope and nothing redundant. The structure is efficient though not especially polished.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given an output schema exists, return-value detail is unnecessary, and the description adequately covers all three parameters and their interplay. The only material gap is sibling routing and any environment/auth preconditions.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, so the description must compensate and largely does: it gives the accepted values for 'target' (OPEN/USER), explains include_friend_db's effect and its USER-only restriction, and defines 'db' precisely. It does not restate the schema defaults, which is fine.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (List) and resource (corpora) and scopes it to 'inside one database, with document counts', which separates it from list_databases at the parent level. It does not explicitly name a sibling it is not, so it falls short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is implied through the prerequisite that 'db' must be 'exactly as returned by list_databases', which tells the agent the correct call ordering. However, there is no explicit when-to-use/when-not or routing against siblings like search_documents or list_databases.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_databasesA

List DocuSky databases available to this server.

Args: target: "OPEN" for public databases, "USER" for the logged-in account's own. include_friend_db: Also list databases shared by DocuSky "friends" (USER only).

ParametersJSON Schema
NameRequiredDescriptionDefault
targetNoOPEN
include_friend_dbNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A3.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden. It conveys read-only listing semantics and hints at authentication via 'logged-in account', but does not state permission requirements, friend-sharing implications, or result limits. Partial disclosure only.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded one-line purpose followed by tight per-argument notes; no filler. The 'Args:' block is compact and each line earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With only two optional parameters, an output schema covering return values, and explicit parameter meanings, an agent has what it needs to invoke the tool. Minor gaps in auth/permission context remain but are low-impact.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate, and it does: it explains the meaning of both the 'target' enum-like values ('OPEN' vs 'USER') and the 'include_friend_db' flag including its USER-only constraint. This is meaningfully beyond the bare schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource ('List DocuSky databases available to this server') with a clear scope. It is distinguishable from the closest sibling list_corpora by naming a different resource type, though it does not explicitly contrast against it.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It implies when to use each mode by explaining that target='OPEN' yields public databases and 'USER' yields the account's own, which is useful selection guidance. However, it gives no explicit when-to-use versus alternatives or prerequisites beyond that.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

post_classificationA

Break a query's hits down by DocuSky's post-classification facets.

Returns, per facet (corpus, period, place, category…), a distribution of [value, document count, hit count] — the quickest way to see how a term is spread across a database. Facet codes returned here (e.g. "COMP", "TP1") are also what twodim_analysis's dim1/dim2 arguments expect.

For a public ("OPEN") database, the result also carries a webUrl to the same breakdown on docusky.org.tw — MCP Apps-capable hosts render it inline.

Args: db: Database title. query: Search terms, or ".all" for the whole database. corpus: Corpus title, or "[ALL]". target: "OPEN" or "USER". owner_username: Owner of a friend-shared database, when applicable.

ParametersJSON Schema
NameRequiredDescriptionDefault
dbYes
queryNo.all
corpusNo[ALL]
targetNoOPEN
owner_usernameNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden, and it does substantial work: it discloses the exact return structure per facet, the meaning of facet codes, and the conditional `webUrl` behavior for OPEN databases with MCP Apps rendering. It omits auth/permission requirements and any statement that the operation is read-only, which is the main remaining gap.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the one-line purpose, then return shape, then cross-tool relevance, then args — a clean structure with no filler sentences. It runs somewhat long for a five-parameter tool, but each block carries information the agent needs.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Params are fully documented, the return layout is described even though an output schema exists, and the webUrl/host-rendering nuance is captured. Missing only the operational context an unannotated tool would normally need, such as auth state and the fact that this is a non-mutating read.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, so the description must compensate, and the Args section documents all five parameters — including non-obvious sentinel values (".all" for the whole database, "[ALL]" for corpus) and the OPEN/USER target distinction. The owner_username entry ("when applicable") is vague about which target or sharing mode triggers it.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource: breaking a query's hits down by DocuSky post-classification facets, with a concrete description of the output shape ([value, count, count] per facet). It also positions itself relative to siblings by noting that its facet codes feed twodim_analysis's dim1/dim2 arguments, so an agent can place it in the tool family without opening another schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

"The quickest way to see how a term is spread across a database" gives a clear usage condition, and the twodim_analysis cross-reference tells the agent how this tool's output feeds a downstream call. However, it never contrasts itself with sibling analysis tools like word_cloud or tag_analysis, so the agent must infer which breakdown tool to pick.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

search_documentsA

Full-text search one DocuSky database; returns metadata plus a short excerpt per hit.

Full document text is deliberately omitted — call get_document with a hit's n (and the same db/corpus/query/page_size) to read one in full.

For a public ("OPEN") database, the result also carries a webUrl to the same search on docusky.org.tw — MCP Apps-capable hosts render it inline.

Args: db: Database title. query: Search terms. +term requires, -term excludes, .all matches everything. corpus: Corpus title, or "[ALL]" for every corpus in the database. page: 1-based page number. page_size: Hits per page (1-100). target: "OPEN" or "USER". excerpt_chars: Characters of body text to preview per hit; 0 disables excerpts. owner_username: Owner of a friend-shared database, when applicable.

ParametersJSON Schema
NameRequiredDescriptionDefault
dbYes
pageNo
queryNo.all
corpusNo[ALL]
targetNoOPEN
page_sizeNo
excerpt_charsNo
owner_usernameNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. It discloses return shape, deliberate omission of full text, excerpt behavior, pagination parameters, and that OPEN databases return a webUrl rendered by MCP Apps hosts. It does not mention authentication or permission requirements, but the read-only search nature is implicit.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is front-loaded with the core purpose and then structured with a clear Args section. Despite being longer due to eight undocumented parameters, every sentence earns its place by adding search semantics, output behavior, or parameter meaning. There is no wasted or repetitive text.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given eight parameters, 0% schema coverage, and an output schema, the description is complete enough for correct invocation. It explains all arguments, the excerpt/full-text distinction, and the public webUrl behavior. The output schema can carry return-value details, and the remaining auth context is not essential to call the search correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate for all eight parameters. It does so thoroughly: it defines db, query with `+term`/`-term`/`.all` syntax, corpus `[ALL]`, 1-based page, page_size 1-100, target OPEN/USER, excerpt_chars with 0 disabling excerpts, and owner_username for friend-shared databases. This adds substantial meaning beyond the bare schema titles and defaults.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The first sentence states a specific verb, resource, and scope: full-text search of one DocuSky database returning metadata plus a short excerpt. It immediately distinguishes itself from get_document by explaining that full document text is deliberately omitted. An agent can tell exactly what this tool does without opening other schemas.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly routes full-document reading to get_document with a hit's `n` and matching db/corpus/query/page_size. It also explains query syntax and target modes, giving clear context for when to use this search tool versus the full-document alternative. Nothing important is left to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

tag_analysisA

Summarize the DocuXML tags (people, places, dates, custom markup) in a query's hits.

Only meaningful for databases whose documents carry inline tagging; returns an empty result otherwise.

For a public ("OPEN") database, the result also carries a webUrl to the same summary on docusky.org.tw — MCP Apps-capable hosts render it inline.

Args: db: Database title. query: Search terms, or ".all" for the whole database. corpus: Corpus title, or "[ALL]". target: "OPEN" or "USER". owner_username: Owner of a friend-shared database, when applicable.

ParametersJSON Schema
NameRequiredDescriptionDefault
dbYes
queryNo.all
corpusNo[ALL]
targetNoOPEN
owner_usernameNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations, so the description must carry the full burden. It discloses a key behavioral trait: returns empty for untagged databases. It also explains the OPEN-database side effect (webUrl to docusky.org.tw, rendered in MCP Apps-capable hosts). It does not discuss read-only nature, permissions, or rate limits, but 'Summarize' implies a read operation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded core purpose, followed by meaningful preconditions and the OPEN/webUrl caveat, then an Args block. Efficient overall, though the webUrl/MCP Apps aside could be trimmed for some hosts.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description covers purpose, preconditions, side effects, and parameter semantics, which is substantial for a 5-param tool with 0% schema coverage. An output schema exists, so return-value details are not required. It could go slightly further on parameter usage (owner_username conditions) and target semantics, but is largely complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, so the description must document all 5 params — and it mostly does: db, query (with '.all' sentinel), corpus ('[ALL]' sentinel), target ('OPEN'/'USER'), and owner_username ('when applicable'). The sentinel values and enum-like values for target are essential semantics the bare schema lacks. Minor gap: no guidance on when owner_username applies beyond 'friend-shared database'.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (Summarize) and resource (DocuXML tags: people, places, dates, custom markup), scoped to a query's hits. The parenthetical enumeration of tag types makes the resource concrete and distinguishable from sibling analysis tools like word_cloud or twodim_analysis.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly states when it is meaningful ('only meaningful for databases whose documents carry inline tagging') and what happens otherwise ('returns an empty result otherwise') — a clear precondition. However, it doesn't compare itself to sibling analysis tools (word_cloud, twodim_analysis), leaving the agent to infer which analysis tool to pick.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

twodim_analysisA

Cross-tabulate a query's hits by two DocuSky classification facets at once.

EXPERIMENTAL: DocuSky added this endpoint on 2026-01-28, but live testing on 2026-09-11 found it rejects every dim1/dim2 pair tried so far — including facet codes taken straight from post_classification's own output, such as "COMP"/"TP1" — with {"code": 1, "message": "Currently not support ..."}. DocuSky's own front-end has not wired up a UI for it yet either. Call post_classification first to see which facet codes exist for a database, try them here, and expect DocuSky may still refuse the combination while this feature is unfinished on their end.

Args: db: Database title. dim1: First classification facet code (a key from post_classification's "facets", e.g. "COMP" or "TP1"). dim2: Second classification facet code to cross with the first. query: Search terms, or ".all" for the whole database. corpus: Corpus title, or "[ALL]". target: "OPEN" or "USER". owner_username: Owner of a friend-shared database, when applicable.

ParametersJSON Schema
NameRequiredDescriptionDefault
dbYes
dim1Yes
dim2Yes
queryNo.all
corpusNo[ALL]
targetNoOPEN
owner_usernameNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden and does so well: it discloses that the endpoint is EXPERIMENTAL (added 2026-01-28), that live testing on 2026-09-11 found it rejects every pair tried, the exact error shape returned, and that DocuSky's own UI is not wired up. This is unusually rich behavioral context beyond any structured field.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Purpose is front-loaded and the Args block is cleanly structured. The experimental warning is longer than typical but each sentence carries decision-relevant information (date, error payload, alternative workflow), so the length is largely justified rather than padding.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 7-parameter, 3-required tool with an output schema already present, the description covers purpose, prerequisites, parameter meanings, and failure behavior comprehensively. Nothing an agent needs in order to either call it correctly or decide to avoid it is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate, and it does: the Args block documents all seven parameters, adding meaning the schema lacks (dim1/dim2 as keys from post_classification's 'facets', query '.'all' for the whole database, corpus '[ALL]', target 'OPEN'/'USER' values absent from the schema's enum-less properties). A few entries (db, corpus, owner_username) remain thin, so it falls short of a 5.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The first sentence gives a specific verb and resource ('Cross-tabulate a query's hits by two DocuSky classification facets at once'), which distinguishes it cleanly from sibling analysis tools like tag_analysis and word_cloud. An agent can identify the operation without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It explicitly prescribes the prerequisite workflow ('Call post_classification first to see which facet codes exist') and warns about the expected failure mode and the condition under which the tool is unusable. This is exactly the when-to-use / when-not-to-use guidance an agent needs.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

word_cloudA

Draw a word cloud of term frequencies in DocuSky's own WordCloudLite tool.

Takes numbers you already have and turns them into a DocuSky visualization: tag counts from tag_analysis, a facet distribution from post_classification (its value / docCount pairs), or terms you counted yourself in text from get_document. This tool does no counting and no searching of its own.

Returns a webUrl to the rendered cloud — MCP Apps-capable hosts show it inline, other hosts can offer it as a link. The rendered page also has buttons for bubble, top-10 bar and table views of the same data, plus Save.

WordCloudLite draws a random subset of a large term list, so pass roughly 20-40 terms when every word should appear.

Args: terms: Term -> weight, e.g. {"針灸": 120, "湯液": 48}. Weights must be positive; non-integers are rescaled proportionally. Commas and semicolons in a term are replaced with spaces (the tool's URL format uses them as separators). max_terms: Keep at most this many of the heaviest terms (default 150). A very long list is also trimmed further to fit DocuSky's URL limit.

ParametersJSON Schema
NameRequiredDescriptionDefault
termsYes
max_termsNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden and delivers well beyond the schema: it discloses the webUrl return, inline-vs-link rendering by host, extra bubble/bar/table views and Save, the random-subset caveat for large lists, term rewriting of commas/semicolons, and URL-length trimming.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the one-line purpose, then progressively deeper detail; the Args block repeats the schema shape but earns its place through added semantics. Slightly long (the 'takes numbers you already have' sentence is near-redundant with the no-counting line), but nothing is filler-critical.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Despite only 2 params and an existing output schema, the definition covers purpose, upstream data contracts, return behavior, rendering differences, and the practical term-count/URL-limit caveats an agent needs to build a working call. Nothing material is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate and does: it gives the terms shape (term -> weight) with an example, requires positive weights, explains proportional rescaling of non-integers, documents the separator-stripping behavior, and states max_terms default 150 plus secondary trimming for DocuSky's URL limit.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('Draw a word cloud of term frequencies') and pins down the engine (DocuSky's WordCloudLite). It explicitly distinguishes itself from siblings by asserting 'This tool does no counting and no searching of its own', separating it from tag_analysis, search_documents, etc.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Names the concrete upstream sources (tag_analysis counts, post_classification value/docCount pairs, terms counted from get_document) so the agent knows the correct pipeline position, and rules out using it for counting/searching. It does not name any alternative visualization tool, but the when-to-use condition is clear.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 1 tool updatev0.4.0
    • Addeddocugis_map
  2. 2 tool updatesv0.3.0
    • Addedtwodim_analysis
    • Addedword_cloud
  3. 7 tool updatesv0.1.0
    • First observedcheck_login
    • First observedget_document
    • First observedlist_corpora
    • First observedlist_databases
    • First observedpost_classification
    • First observedsearch_documents
    • First observedtag_analysis

TDQS

A4.2/5.0

Scored across 10 tools

Disambiguation4/5

Tools target distinct actions (search, get, list, analyze, map, login) with clear boundaries. Minor potential confusion between post_classification and twodim_analysis, but descriptions clarify their different purposes (facet breakdown vs. cross-tabulation).

Naming Consistency5/5

All names follow a consistent verb_noun pattern (e.g., search_documents, get_document, list_databases, post_classification, twodim_analysis). No deviations in casing or style.

Tool Count5/5

10 tools is a well-scoped set for a document search and analysis server, covering search, retrieval, metadata listing, analysis, visualization, mapping, and login without redundancy.

Completeness4/5

Core lifecycle is covered: list databases/corpora, search, retrieve full text, analyze facets/tags, visualize, map, and check login. Missing direct document export or bulk operations, but agents can work around these gaps.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • F
    license
    A
    quality
    D
    maintenance
    Enables Claude to search, query, and interact with an Enterprise Knowledge Management System (EKMS). Supports semantic search, knowledge recommendations, relationship graphs, and feedback recording for enterprise knowledge bases.
    7
    -
  • F
    license
    Not graded
    quality
    Not graded
    maintenance
    Enables Claude to interact with Google Docs to list, read, create, search, and update documents in a user's Google Drive. It provides a suite of tools and prompts for document management and content analysis using OAuth 2.0 authentication.
    1,236 npm
    -