ExcelMCP
ExcelMCP
AIエージェントのための生きたExcelインテリジェンスレイヤー。 OneDriveフォルダを指定するだけで、エージェントが現在の数値に基づいて、それらのスプレッドシートについて平易な英語で質問できるようになります。
これが解決する問題
ほとんどのスプレッドシート連携は、データを別の場所にコピーすることで機能します。ワークブックを取り込み、チャンク化し、セルの値を埋め込み、全体をベクターデータベースに保存します。その瞬間から、エージェントはスナップショットに関する質問に答えることになります。誰かが午前9時に在庫シートを更新しても、エージェントは火曜日の数値を引用し続けます。
ExcelMCPは問題を2つに分割します。
構造はキャッシュされます。 ファイル名、シート名、列ヘッダー、ヘッダー行の開始位置、日付を保持する列、シート同士の関係、そして低カーディナリティ列ごとの個別ラベルの小さなサンプル — これが、何百ものほぼ同一のシート間でのルーティングを意味のあるものにします。これらはめったに変更されず、保存コストも安く、エージェントが何を要求すべきかを知るために必要なものです。(サンプリングされたラベルは、構造が値に触れる唯一の場所です。正確な境界はディスクに保存されるものに詳述されています。)
データは決してキャッシュされません。 数値を返すすべてのツールコールは、Microsoft Graph APIにアクセスし、ライブでデータを取得します。古くなるデータキャッシュも、遅れる同期ジョブも、ディスクから提供される回答もありません。
すべてのレスポンスには、metadata.fetched_atタイムスタンプとis_cached: falseフラグが含まれ、モデルが帯域内で最新データを見ていることを認識できるようにします。
Related MCP server: Microsoft 365 MCP Server
仕組み
自然言語の質問が埋め込まれ、コサイン類似度によってシートの説明と照合され、次に列名とサンプリングされた値との字句的な重複によって再ランク付けされます。これにより、20のワークブックが1つのスキーマを共有する場合でも、ルーティングが意味を保ちます。これらのシート、そしてそれらのシートのみがライブでフェッチされます。フィルタリングと集計は、新しくフェッチされたフレームに対してpandasで行われます。単一値の質問は行パイプラインを完全にスキップします。lookupは1つのキー列と1つの行を読み取り、その出所とともにセルを返します。
要件
Python 3.10 以降
OneDriveを使用するMicrosoft 365アカウント
uv、または必要に応じて通常のpip
インストール
リポジトリルートから:
git clone https://github.com/Karunya-Muddana/ExcelMCP.git
cd ExcelMCP
uv sync # install dependencies
uv build # build the wheel
pip install dist/excelmcp-0.3.0-py3-none-any.whlまたは、ビルドせずにソースから直接インストール:
pip install .コンパイラのステップやビルドする必要のあるネイティブ拡張はありません。ベクトル検索はhnswlibではなくNumPyのコサインスキャンで実行されるため、C++ツールチェーンのないマシンでもpip installが機能します。
セットアップ
ウィザードを一度実行します:
excelmcp-setup以下の4つの手順を実行します:
Microsoftデバイスフローサインイン。コードが発行され、ブラウザに貼り付けると、トークンキャッシュが
0600パーミッションで~/.excelmcp/token.jsonに保存されます。インデックスを作成するOneDriveフォルダ(例:
/ERP)。そのフォルダ内のすべての
.xlsxをスキャンして、構造グラフと埋め込みを構築します。マシンに既にインストールされているAIエージェントを検出し、選択したエージェントの設定エントリを書き込みます。
自動設定可能なエージェント
エージェント | 設定ファイル |
Claude Code |
|
Claude Desktop |
|
Cursor |
|
Windsurf |
|
Gemini CLI |
|
Codex CLI |
|
VS Code (Copilot) | VS Codeユーザー |
Cline | 拡張機能 |
Continue |
|
Goose |
|
Zed |
|
Hermes |
|
既存の設定ファイルは変更前にバックアップされます。エージェントがリストにない場合、ウィザードは自分で貼り付けるための正確なJSONまたはTOMLブロックを出力します。
その他のウィザードコマンド
excelmcp-setup list-agents # show what was detected
excelmcp-setup install --only cursor # register with one agent, skip the rescan
excelmcp-setup doctor # diagnose a broken install
excelmcp-setup uninstall # remove ExcelMCP from every agent config
excelmcp-setup --folder /ERP --yes # fully non-interactive
excelmcp-setup --dry-run # print the changes, write nothingエージェントに公開されるツール
ツール | ネットワーク | 機能 |
| なし | ワークスペースの完全な構造: ファイル、シート、列、テーブル領域、関係、命名バリアント、スキャン経過時間。即時。 |
| なし | 同上、1つのファイルに限定、最終スキャン時点のおおよその行数付き。即時。 |
| 高負荷 | OneDriveを再クロールし、構造、サンプリング値、関係、埋め込みを再構築。 |
| ライブ | 自然言語の質問。ベクトル類似度と字句的再ランク付けによってルーティング。 |
| ライブ | 1回の呼び出し → ファイル/シート/セルの出所と信頼度シグナルを含む1つのセル値。 |
| ライブ | 1つのGraphリクエストで1つのアドレス指定されたセル。 |
| ライブ | 1つのシートをフェッチし、条件に一致する行を返す。 |
| ライブ | 1つのシートをフェッチし、グループ化して集約、 |
| ライブ | すべてのファイルから一致するシートをフェッチし、合計にまとめる。 |
| ライブ | 既知の関係から提案されたキー列で2つのシートをマージ。 |
| ライブ | トランザクションタイプに基づく符号付き合計 — 1回の呼び出しで正味在庫。 |
2つの構造ツールはローカルグラフを読み取るため、無料で即時応答です。ライブとマークされたものは、呼び出しのたびにAPIにアクセスします。
使用方法
サーバーが登録されると、通常通りエージェントと会話するだけです。内部的には、以下のような呼び出しが行われます。
まずは方向性を確認。エージェントは列名を推測する前に常にこれを行うべきです。同じ名前の会社は2つとないからです:
get_workspace_graph(folder_path="/ERP")答えがどこにあるかわからずに質問する:
query("what are the top 10 products by sales value", folder_path="/ERP")既知のシートをフィルタリングする:
filter_sheet(
file_name="Inventory.xlsx",
sheet="Stock",
conditions={"Status": "Low", "Quantity": "<50"},
folder_path="/ERP",
sort_by="Quantity",
limit=100,
)サポートされている条件演算子(すべてAND結合):
形式 | 意味 |
| 完全一致 — 大文字小文字と空白を区別しません。厳密な場合は |
| 部分一致、リテラル部分文字列、正規表現ではありません |
| より大きい( |
| 日付境界、ISO-8601形式、検出された日付列で機能します |
| リストされた値のいずれか |
| 包括範囲、数値または日付 |
| 複合境界 |
| nullチェック — 空白と空文字列はnullとしてカウントされます |
存在しない列名や演算子は、静かに0行を返すのではなくエラーを発生させます。これは、エージェントが誤った情報を自信を持って報告する原因となる障害モードです。条件が正当に何も一致しない場合、レスポンスにはzero_match_diagnosticsが含まれます — 各条件が単独で何に一致したか、問題の列に実際に存在する最大20個の値 — これにより、ニアミスが「データなし」と報告される代わりに修正されます。
1回の呼び出しで単一の数値を尋ねる:
lookup(query="contracted rate for Titanium Dioxide under the BESTEX contract",
folder_path="/Contracts")回答は出所(ファイル、シート、セルアドレス、一致した行)と信頼度フィールドとともに返されます。複数の一致行がある場合はambiguousとすべての行が返され、シート間で矛盾がある場合はconflictとすべてのバージョンが値なしで返され、キーのスペルミスはあいまいな提案を返します。このツールは裸の数値を返すことはありません。
1つのファイル内でグループ化して集約する:
aggregate(
file_name="Sales.xlsx",
sheet="Q1",
group_by="Region",
value_col="Revenue",
operation="sum",
folder_path="/ERP",
)ワークスペース内のすべてのファイルで同じシートを合計する:
cross_file_aggregate(
sheet="Q1",
value_col="Revenue",
operation="sum",
folder_path="/ERP",
conditions={"Status": "Closed"},
)cross_file_aggregateは、合計とともにファイルごとの内訳を返し、ファイルが読み取れなかった場合はskipped_filesを、正確なシート名を含まないすべてのファイルに対してdid_you_mean候補を含むunmatched_filesを返します。これにより、部分的な合計が静かに間違っているのではなく、明らかに部分的であることがわかります。これは、一部のファイルでシート名がSalesであり、他のファイルでSales 2024である場合も含みます。集計する前にget_workspace_graphでsheet_name_variantsを確認して、その断片化を事前に把握してください。
エージェントプレイブック
サーバーのインストールは簡単な半分です。agents/フォルダは残りの半分をカバーしています: これらのツールを持つエージェントにプロンプトを与える方法、各ホストに配線する方法、そして動作した後に何を自動化するか。
カスタムエージェント、サブエージェント、 | |
ジョブ別に分類されたコピペ用プロンプト:オリエンテーション、ストレートな回答、分析、検証、レポート、データ品質。最後にアンチプロンプト(もっともらしく見えて確実に間違った答えを生成する言い回し)のセットが付いています。 | |
チェーンがエンドツーエンドで機能することを証明する最初のセッション。データが実際にライブであることを自分で確認する方法も含みます。 | |
12のサポートされているホスト構成のそれぞれに書き込まれる内容、確認方法、ホストごとのクセ、およびホストなしでサーバーをプログラム的に駆動する方法。 | |
どのツールを手に取るか、セマンティックルーティングが実際にどのようにシートを選択するか、条件構文で表現できないこと、そして自信満々の間違った答えを生み出すデータ形状。 | |
症状の解説:PATHの問題や403エラーから、文字化けした列名や2倍になる合計値まで。 | |
4つのすぐにスケジュールできるルーチン:毎日の在庫チェック、週次売上ダイジェスト、月末調整、データ品質監査。それぞれにプロンプト、スケジュール、よくある問題が含まれています。 |
サーバーに組み込まれたガードレール
サーバーには、MCPインストラクションとして一連の運用ルールが同梱されており、ホストモデルが最初の呼び出しを行う前に読み取ります。これらは、LLMがスプレッドシートの質問を間違える具体的な方法に対処するために存在します。
ファイル名、シート名、列名を決して推測しない。グラフから発見すること。
ファイルをまたいだ数値を頭の中で合計しない。
cross_file_aggregateを呼び出してツールに任せること。openpyxl、pandas.read_excel、またはローカルファイルシステムに決してアクセスしない。ファイルはこのマシン上にありません。トランザクション形式のデータで数量列を生のまま合計しない。
deriveを使用し、トランザクションタイプを明示的に指定すること。日付列はISO-8601文字列として到着し、サーバーによってシリアル値から変換済みです。シリアル演算を手動で行わないこと。
単一の数値の場合は
lookupを呼び出し、それが返す出典情報を引用すること。ambiguousおよびconflictの結果を表示し、値を自分で選ばないこと。結果が完全であると主張する前に、
truncatedおよびtotal_matchedフィールドを確認すること。
サーバーインストラクションを無視するホスト、および自分で構築するカスタムエージェントについては、この内容を独自のプロンプトに明記する必要があります。agents/system-prompt.mdを参照してください。
構成
変数 | デフォルト | 目的 |
| 組み込み | Azure ADアプリケーションクライアントID |
|
| テナント。個人アカウントの場合は |
| 未設定 | ツール呼び出しで |
|
| すべてのコードパスにわたる、Microsoft Graphへの最大同時リクエスト数。 |
組み込みのクライアントIDは、デバイスコードフローに使用されるパブリッククライアントです。シークレットは含まれず、設計上すべての認証リクエストで可視であり、このリポジトリに含めても安全です。同意画面に組織名を表示したい場合は、独自のアプリ登録に置き換えてください。
ディスクに保存されるもの
~/.excelmcp/
token.json MSAL token cache. Auth material only, written 0600.
graph.json Structure graph: item IDs, sheet names, column headers,
used-range dimensions, date column types, per-sheet
table regions, inferred and formula-declared
relationships — and sampled values (see below).
vectors.npy Embedded sheet descriptions for semantic routing.
metadata.json Labels and lexical terms tying each embedding to a sheet.
relationships.yaml Optional, written by you: declared join relationships.バージョン0.3.0時点での、キャッシュなし主張の正直なバージョン。 データの行、セルグリッド、クエリ可能な値はディスクに保存されません。すべての回答は常にライブフェッチから提供されます。意図的な例外が1つあります:graph.jsonはサンプリングされた値を保存します。これは、低カーディナリティ列(クライアント名、ステータス、材料名、単位)あたり最大50個の異なるテキストラベルで、スキャン時に取得されます。これらは、質問をルーティングする際に構造的に同一のシートを区別できるようにするため、lookupがすべてをダウンロードせずにどのシートに「BESTEX」が含まれているかを特定できるようにするため、そして関係を列名から推測するのではなく値の重複から推測できるようにするために存在します。これらはルーティングの証拠であり、データキャッシュではありません。これらから質問に答えることは決してなく、ワークスペーススキャンによって完全に更新されます。グラフはまた、シートごとの構造フィンガープリント(ヘッダー列と使用範囲アドレス)を保存し、これはドリフト検出専用です。さらに、バージョン0.3.0ではリージョンマップも保存します:シート上の各テーブル本体の行スパン。これは、シート自体のSUM/COUNT/AVERAGE数式が参照する範囲と、シート間の数式が読み取るアドレスから導出されます。これらは行番号とセルアドレスであり、内容ではありません。値を読み取って生成されることはありません。存在する場合のリージョンのlabelは、サンプリングされた値に次ぐ2番目の意図的な例外です。リージョンのすぐ上のセクションバナーセルから読み取られた数語(「NAPHTHALENE」、「OLEUM 65%」)で、モデルが行番号から推測するのではなく、どのテーブルを指しているかを名前で指定できるように保持されます。これはシートのレイアウトを記述する構造メタデータであり、行データではありません。サンプリングされた値がすでに描いているのと同じ区別です。これらがディスクに保存される以上のことを望まない場合は、そのフォルダをスキャンしないでください。境界を確認したい場合は、graph.jsonは小さくて読み取り可能なので、見てみてください。
Windowsでは、os.chmodは読み取り専用ビットのみを切り替えるため、0600モードはそこでのベストエフォートであり、実際の保護は%USERPROFILE%のデフォルトのユーザーごとのACLです。macOSとLinuxでは、モードはコンテンツが書き込まれる前に一時ファイルに適用されるため、トークンが一時的にワールドリーダブルとして存在することはありません。
テスト
# offline unit tests, no network and no credentials required
pytest tests/test_unit.py
# live integration tests against a workspace you have already scanned, opt in
EXCELMCP_TEST_FOLDER=/ERP pytest tests/test_live_integration.py -v統合テストスイートは、EXCELMCP_TEST_FOLDERが設定されていない場合に自分自身をスキップするため、プレーンなpytestの実行はオフラインのままです。
プロジェクトレイアウト
agents/ prompts, host guides, and schedulable routines
auth.py MSAL device flow, token cache, proactive refresh
graph_client.py Graph API wrapper, 429 backoff, shared concurrency gate
structure.py Structure discovery, value sampling, relationship inference
embeddings.py FastEmbed vectors, NumPy cosine search, lexical rerank
query_engine.py Conditions, live fetch, aggregation, joins, derive
lookup.py Single-cell lookup pipeline and get_cell
ranges.py A1-notation range arithmetic
main.py FastMCP tool definitions and server entry point
cli.py Setup wizard, agent detection, config writing
agents.py Per agent config formats and file locations
storage.py Atomic writes, stderr logging, config directory handlingコントリビューション
Issueとプルリクエストは歓迎します。別のエージェントのサポートを追加する場合、agents.pyだけを触る必要があります。AgentSpecを設定パス、エントリ形状、検出ヒントとともに追加してください。
ライセンス
MIT。 LICENSEを参照してください。
Available Tools
11 toolsaggregateA
Fetches a sheet LIVE and runs a grouped aggregation. Operations: sum, count, mean, min, max. group_by is one column name or a list of them. conditions filters rows before aggregating (same grammar as filter_sheet, including the object form). having filters the AGGREGATED rows afterwards, e.g. having={"Revenue": ">1000"} keeps only groups whose aggregate exceeds 1000. SINGLE FILE ONLY. For totals across multiple files you MUST use cross_file_aggregate instead — never use this tool and then manually add results across files. Get column names from get_workspace_graph first. Returns rows plus a truncated flag. If conditions matched zero rows, zero_match_diagnostics shows what each condition matched alone and the values actually present — correct the condition and retry.
| Name | Required | Description | Default |
|---|---|---|---|
| sheet | Yes | ||
| having | No | ||
| group_by | Yes | ||
| file_name | Yes | ||
| operation | Yes | ||
| value_col | Yes | ||
| conditions | No | ||
| folder_path | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Despite no annotations, the description discloses key behavioral traits: live data fetch, output includes a 'truncated flag', and zero_match_diagnostics behavior when no rows match. It also explains condition grammar and having filter semantics with an example, providing substantial context beyond the schema.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense but well-structured with line breaks, and each sentence adds necessary information (operations, parameters, single-file constraint, diagnostics). It is slightly long but avoids redundancy and earns its length.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the complexity (8 params, grouped aggregation), the description covers essential context: multi-file exclusion, column names source, condition/having grammar, return flags, and error diagnostics. An output schema exists, so return details need not be repeated. Folder_path is the only minor omission, but optional and less critical.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description carries the burden. It explains group_by (column name or list), conditions (filter_sheet grammar), having (post-aggregation filter with example), and lists operations. However, folder_path and file_name/sheet are not explicitly described, leaving minor gaps for those parameters.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's core action: 'Fetches a sheet LIVE and runs a grouped aggregation' and lists supported operations. It also distinguishes itself from siblings by explicitly limiting to a single file and pointing to cross_file_aggregate for multi-file operations.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides explicit when-to-use and when-not-to-use guidance: 'SINGLE FILE ONLY' and 'For totals across multiple files you MUST use cross_file_aggregate instead'. It also references filter_sheet grammar for condition syntax and advises fetching column names from get_workspace_graph first.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
cross_file_aggregateA
MANDATORY for any total spanning more than one file. Fetches relevant sheets from ALL files in PARALLEL, applies filter conditions, returns the aggregate total.
WHEN YOU MUST CALL THIS:
Any total, sum, count, or average across multiple files
Any cross-file comparison or consolidation
Verifying a total you calculated from individual files
NEVER calculate cross-file totals by:
Adding individual filter_sheet results in your head
Using Python to sum numbers from separate tool calls
Guessing based on partial data
Always call this AND show per-file breakdown so the user can verify both agree. If they differ, flag it.
ONLY files whose sheet is named EXACTLY sheet are
included in the total. Files without that exact sheet
are listed in unmatched_files, with their actual sheet
names and did_you_mean candidates — they are NEVER
silently included. If the response has a warning,
skipped_files, or unmatched_files, surface that to the
user: the total may be incomplete. Check
sheet_name_variants in get_workspace_graph first to see
naming fragmentation before aggregating.
| Name | Required | Description | Default |
|---|---|---|---|
| sheet | Yes | ||
| operation | Yes | ||
| value_col | Yes | ||
| conditions | No | ||
| folder_path | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Given no annotations, the description discloses key behaviors: parallel fetching, exact sheet-name matching, listing unmatched files with did_you_mean candidates, and never silently including them. It also warns that warning/skipped_files/unmatched_files indicate incomplete totals and mandates surfacing them to the user. It does not explicitly state read-only nature, but there are no mutations implied.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is longer than average but well-structured with bolded headings and lists, making it scannable. Each sentence carries actionable guidance, though some redundancy exists (e.g., repeated emphasis on showing per-file breakdown). Overall, it earns its length without being bloated.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool has an output schema, so return values need not be described, yet the description references response fields (unmatched_files, skipped_files, warning) for error handling and gives a cross-tool prerequisite. It does not explain all parameters or link to filter_sheet's conditions structure, but it is highly comprehensive for a complex tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must compensate. It clarifies that `sheet` must match exactly and mentions 'filter conditions' conceptually, but it does not explain `value_col`, `operation` options, `conditions` structure, or `folder_path`. It adds some semantic context beyond the bare schema but leaves significant parameter gaps.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with 'MANDATORY for any total spanning more than one file' and explicitly states it fetches sheets from all files, applies filter conditions, and returns the aggregate total. It distinguishes from siblings like filter_sheet and aggregate by contrasting its cross-file scope with single-file alternatives.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides an explicit 'WHEN YOU MUST CALL THIS' list (any cross-file total/sum/count/average, cross-file comparison, verifying totals) and a 'NEVER calculate' list (adding filter_sheet results, Python summing, guessing). It also instructs to check sheet_name_variants in get_workspace_graph first, naming a prerequisite tool.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
deriveA
Computes a NET value over transaction types in one call: sum of sign * groupwise_sum(quantity_col) across the given components. This is how a stock figure like receipts + purchases − consumption − returns becomes ONE call with the arithmetic done in pandas, instead of five filter_sheet calls added up in your head (which RULE 3 forbids).
components is a list of {"conditions": {...same grammar as filter_sheet...}, "sign": 1 or -1, "label": "receipts"} conditions (optional) pre-filters the sheet before any component applies. The response includes a per-component breakdown with rows_matched. A component that matched ZERO rows is flagged and warned about — check the spelling of the transaction type before trusting the net.
| Name | Required | Description | Default |
|---|---|---|---|
| sheet | Yes | ||
| group_by | Yes | ||
| file_name | Yes | ||
| components | Yes | ||
| conditions | No | ||
| folder_path | No | ||
| quantity_col | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It discloses that computation is done in pandas, includes a per-component breakdown with rows_matched, and flags zero-match components with a warning. It does not explicitly state that the operation is read-only (no file modification), but given its nature and the context, that is a minor omission. Overall, it provides strong behavioral safeguards beyond a simple summary.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact but rich. It front-loads the core purpose, then provides an example, breaks down the components structure, and adds a critical warning. A few words are slightly repetitive ('one call' appears twice), but each sentence adds value and the structure is logical, so it earns a high score without being overly verbose.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool has 7 parameters and is genuinely complex, but the description covers the main behavioral contract: the net computation, component structure, optional conditions, and output warnings. It doesn't elaborate on folder_path or file_name, but those are self-explanatory. Given the output schema exists (per context signals), the description does not need to explain return values in detail. This is fairly complete for a tool of this complexity.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description is the only source of parameter meaning. It thoroughly explains the complex 'components' list (conditions, sign, label), clarifies 'conditions' as optional pre-filter, and ties 'quantity_col' to the groupwise_sum. This goes well beyond the bare schema and compensates entirely for the lack of schema descriptions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('Computes') and clearly defines the resource ('NET value over transaction types'). It explicitly contrasts with filter_sheet by showing how it replaces five calls, which strongly distinguishes it from siblings. This is a textbook example of purpose clarity.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly states when to use this tool: to compute a net figure from signed components in one call, and tells the agent to avoid multiple filter_sheet calls (citing RULE 3). It names the alternative filter_sheet and implies that the tool is the right choice for this pattern. No exclusion criteria are missing; it is very clear.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
filter_sheetA
Fetches a specific sheet LIVE from OneDrive and returns rows matching the given conditions. Always live — no cache. Use when you already know which file and sheet to query. Get column names from get_workspace_graph first.
Condition formats (string form): Exact match: {"ColumnName": "value"} Contains: {"ColumnName": "~value"} (literal, not regex) Comparisons: {"ColumnName": ">100"} (also >=, <, <=) Date bounds: {"Batch Date": ">=2026-01-01"} (ISO-8601)
Condition formats (object form, combinable): IN list: {"Status": {"in": ["Closed", "Shipped"]}} Range: {"Qty": {"between": [10, 500]}} Date range: {"Batch Date": {">=": "2026-01-01", "<": "2026-04-01"}} Null check: {"Notes": {"is_null": false}} Contains: {"Name": {"contains": "oxide"}}
Multiple conditions are ANDed together; multiple operators inside one object are ANDed too. An unknown column name or operator is an error, not an empty result.
MATCHING IS NORMALISED, NOT STRICT: exact string matches ignore case and surrounding whitespace ("closed" matches "Closed "), because Excel cells carry stray whitespace constantly. Pass exact_case=true for byte-for-byte matching. Contains (~) is case-insensitive. If zero rows match, the response includes zero_match_diagnostics showing what each condition matched on its own and the values actually present in the column — use it to correct a near-miss and retry instead of concluding the data does not exist. At most 1000 rows are returned; check the truncated and total_matched fields in the response.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | ||
| sheet | Yes | ||
| sort_by | No | ||
| file_name | Yes | ||
| conditions | Yes | ||
| exact_case | No | ||
| folder_path | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description takes full responsibility for behavioral disclosure. It covers the live/cache behavior, case-insensitive normalized matching, exact_case flag, literal-contains semantics, error behavior for unknown columns/operators, zero_match_diagnostics, and the 1000-row limit with truncated/total_matched fields.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is long but earns its length: a clear purpose sentence, structured condition formats with examples, then matching semantics and edge-case behavior. Each section serves a distinct need, and examples are concrete.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given seven parameters, nested conditions, and an output schema, this description covers the high-risk behaviors: error semantics, matching rules, zero-match diagnostics, and response truncation. The presence of an output schema covers return-value structure, and the description supplements it with total_matched/truncated details. The only gaps are sort_by/folder_path semantics, which are minor.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It does, thoroughly, for the conditions parameter: exact/contains/comparison/date/IN/between/is_null/contains object forms, AND semantics, and normalization. It also explains exact_case and limit behavior. However, sort_by and folder_path are not explicitly described beyond the schema, a minor gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb-resource pair: 'Fetches a specific sheet LIVE from OneDrive and returns rows matching the given conditions.' It also preempts sibling confusion by noting to get column names from get_workspace_graph first and stating 'Use when you already know which file and sheet to query.'
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explicitly says 'Use when you already know which file and sheet to query,' and directs users to get_workspace_graph for column names, setting clear context. It doesn't spell out when not to use it relative to query/aggregate/join_sheets, but the specificity of the condition syntax and the mention of the 1000-row limit imply boundaries.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_cellA
Reads EXACTLY ONE cell, LIVE, by address. One Graph request, tiny payload, no ambiguity. Use when the location is already known — follow-up questions, scheduled routines, anything where lookup or filter_sheet already established the address earlier. address is A1 notation ("B7") or the name of a workbook-scoped named range that resolves to one cell. Serial dates arrive converted to ISO-8601; check resolved_type. A multi-cell address is an error — use filter_sheet for ranges.
| Name | Required | Description | Default |
|---|---|---|---|
| sheet | Yes | ||
| address | Yes | ||
| file_name | Yes | ||
| folder_path | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries full burden. It discloses that it reads exactly one cell, performs a live Graph request, converts serial dates to ISO-8601 with a resolved_type check, and treats multi-cell addresses as errors. These are concrete behavioral details beyond a simple read hint.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences earn their place: purpose, usage, then parameter and behavior details. Front-loaded and free of fluff.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's simplicity and the presence of an output schema, the description covers all key aspects: exact behavior, usage context, parameter semantics, and an edge case. No critical gaps for an AI agent to select and invoke it.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must compensate. It thoroughly explains the address parameter (A1 notation or named range) and its constraints, though file_name, sheet, and folder_path rely on their naming for meaning. Adds clear value for the most complex parameter.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Reads'), a precise resource ('EXACTLY ONE cell'), and key qualifiers ('LIVE, by address'), distinguishing it from sibling range tools like filter_sheet. It emphasizes 'no ambiguity' to set expectations.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explicitly states when to use the tool: 'when the location is already known' for follow-up questions or scheduled routines, and names filter_sheet as the alternative for ranges. This provides a clear when-to-use vs when-not-to.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_workspace_graphA
Returns the cached file structure — all filenames, sheet
names, column headers, and cross-sheet relationships
(inferred at scan time from matching column names plus
overlapping sampled values, merged with any the user
declared in ~/.excelmcp/relationships.yaml; each carries
a confidence score and its evidence).
INSTANT — makes no API call. Reads from local graph.json.
ALWAYS call this first at session start to orient yourself.
Shows you exactly which files exist, what sheets they have,
and what columns are in each sheet. The structure varies
for every company — never assume, always discover.
Also returns sheet_name_variants: groups of sheet names
that differ only in case or whitespace across files —
check it before any cross-file operation, because those
match by exact sheet name.
Each sheet carries a regions list: the table bodies found
in it, derived from the sheet's own SUM/COUNT formulas, in
absolute sheet rows. A sheet with more than one region holds
several separate tables (also listed in multi_region_sheets),
so a plain aggregate over it adds up blocks that were never
meant to be summed — read its unclaimed_rows and check which
region you mean before totalling anything.
layout_confidence is "unconfirmed" wherever the region map
came from formulas alone and nothing has verified it.
Use this before any filter_sheet call when unsure which
file or column to query.
| Name | Required | Description | Default |
|---|---|---|---|
| folder_path | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden and does so thoroughly. It discloses caching ('makes no API call', 'Reads from local graph.json'), warns about unconfirmed layout_confidence, explains multi-region sheets and the risk of summing separate tables, and notes sheet_name_variants as a matching caveat.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with the core function and all sentences add behavioral value. However, it is somewhat verbose and contains overlapping usage advice ('ALWAYS call this first' and 'Use this before any filter_sheet call'), so it is not maximally concise.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description is exceptionally complete for a tool with an output schema and no annotations. It covers the return content, performance characteristics, caching path, confidence scoring, multi-region quirks, and recommended invocation order, leaving little ambiguity about the tool's role.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema defines one parameter, folder_path, but the description never mentions it. Schema description coverage is 0%, and the description provides no compensation—an agent would not know when or why to provide folder_path, or what happens if omitted.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('Returns') naming a clear resource ('cached file structure') and enumerates exact content (filenames, sheet names, column headers, cross-sheet relationships). It also distinguishes itself from siblings by highlighting that it is cached and requires no API call.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicit guidance is given: 'ALWAYS call this first at session start' and 'Use this before any filter_sheet call when unsure which file or column to query.' This makes the intended usage context and sequencing very clear.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
inspect_fileA
Returns structural metadata for one specific file — sheet names, column headers, and each sheet's approx_row_count AS OF THE LAST SCAN (this tool makes no API call, so the count is not live; treat it as an order-of-magnitude hint, not a current figure). INSTANT — reads from cached graph.json. Use before filter_sheet when you need to confirm the exact column names available in a specific file.
| Name | Required | Description | Default |
|---|---|---|---|
| file_name | Yes | ||
| folder_path | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description fully discloses behavior: it 'makes no API call', 'reads from cached graph.json', and the count is 'not live' and should be treated as an 'order-of-magnitude hint'. This reveals staleness and performance characteristics beyond what annotations would provide, ensuring the agent understands the tool's limitations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is concise, front-loaded with the primary purpose, and every sentence earns its place: it covers the output, the caveat about non-live counts, the performance characteristic (INSTANT), and a usage example. No fluff or redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description provides rich behavioral context (caching, staleness, speed) and usage guidance, and an output schema exists to detail return values. However, the folder_path parameter is not explained, and the description does not mention potential error cases or prerequisites. These gaps are minor given the tool's simplicity.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must explain parameters. It implicitly covers file_name via 'one specific file', but it does not mention folder_path at all. This leaves one of two parameters unexplained, which is a significant gap in parameter semantics.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose: 'Returns structural metadata for one specific file — sheet names, column headers, and each sheet's approx_row_count.' This is specific with a verb ('returns') and resource ('one specific file'), and it distinguishes the tool from siblings by emphasizing structural metadata and its use before filter_sheet.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly advises 'Use before filter_sheet when you need to confirm the exact column names available in a specific file.' This provides a clear when-to-use scenario and a named alternative. It also notes that the tool makes no API call, implying it is for quick checks rather than live operations.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
join_sheetsA
Joins two sheets LIVE on key columns and returns the merged rows, with filter_sheet's truncation contract (total_matched, truncated, limit). Omit left_on/right_on to let the server pick keys from the workspace's known relationships — it uses a declared or high-confidence inferred relationship and REFUSES with the candidate list when confidence is low, rather than guessing. The keys actually used and their source are in data.keys. Key matching is normalised (case, whitespace, 45 vs 45.0); null keys never join. join_type: inner, left, right, outer. Colliding column names get _left/_right suffixes. Use this instead of stitching filter_sheet results together yourself.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | ||
| left_on | No | ||
| right_on | No | ||
| join_type | No | inner | |
| left_file | Yes | ||
| left_sheet | Yes | ||
| right_file | Yes | ||
| folder_path | No | ||
| right_sheet | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden of behavioral disclosure. It details the live join behavior, truncation contract, refusal with candidate list when confidence is low, key normalization, null key handling, join types, and column suffixing. This is exceptionally transparent for a data tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single dense paragraph that front-loads the core purpose, then logically covers optional behavior, key handling, and an explicit usage recommendation. Every sentence adds valuable information without redundancy or fluff.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description is highly complete given the tool's complexity, covering core behavior, edge cases, and output details. However, it does not explain the 'folder_path' parameter, which is part of the schema. While not critical to the main join functionality, this leaves a minor gap in the overall contextual picture.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema has zero description coverage, so the description must compensate. It explains the semantics of left_on/right_on, join_type, limit, and the data.keys output field. It even covers edge cases like colliding column names and normalized matching, adding rich meaning beyond the raw parameter names.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool 'Joins two sheets LIVE on key columns and returns the merged rows,' naming a specific verb and resource. It also distinguishes itself from sibling tools by explicitly recommending this tool over stitching filter_sheet results together.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description includes explicit guidance on when to use this tool versus alternatives: 'Use this instead of stitching filter_sheet results together yourself.' It also explains the optional behavior of omitting key parameters and the server's decision-making process, giving the agent clear context for invocation.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
lookupA
ONE-CALL semantic lookup: finds a single cell value anywhere in the workspace and returns it WITH PROVENANCE (file, sheet, cell address, the matched row) and a confidence signal. Reads only the key column and the matched row — never whole sheets.
Two ways to call it:
Natural language: lookup(query="contracted rate for Titanium Dioxide under the BESTEX contract"). The server resolves the key value against values sampled at scan time and picks the return column lexically. Works best when the query contains a literal value that appears in the data (a client, a material).
Explicit: lookup(key_column="Material", key_value= "Titanium Dioxide", return_column="Contracted Rate"). Use this when the query form reports it could not parse, or for values too rare to be sampled. scope={"file": ..., "sheet": ...} narrows the search.
READ confidence BEFORE using the value: "high" — single row matched; corroborating sheets (if any) agree. provenance.corroborated_by lists them. "ambiguous" — the key matched SEVERAL ROWS. value is null; every row is in alternatives. Never pick one silently. "conflict" — several sheets DISAGREE. value is null; every version is in alternatives. Surface the conflict to the user. found=false — key not found; suggestions holds fuzzy near-misses (retry with exact spelling), or ambiguity explains why routing failed. NEVER present a value from this tool without citing provenance.file, provenance.sheet and provenance.cell.
| Name | Required | Description | Default |
|---|---|---|---|
| query | No | ||
| scope | No | ||
| key_value | No | ||
| key_column | No | ||
| folder_path | No | ||
| return_column | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full behavioral disclosure burden, and it does so thoroughly: it states that only the key column and matched row are read, never whole sheets; it explains the confidence signals (high/ambiguous/conflict) and their consequences; and it mandates citing provenance. This goes far beyond what the schema alone could convey.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is long but well-structured and front-loaded with the core behavior, then branches into invocation modes, confidence semantics, and a hard safety rule. The line breaks and indented sections make it scannable, and every paragraph adds necessary information rather than padding.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity and the lack of annotations, the description is exceptionally complete: it covers both call styles, scope, confidence interpretation, fallback suggestions, and provenance requirements. An output schema exists for return-value structure, so not restating the full return object is acceptable. The omitted folder_path parameter is minor and does not undermine the overall completeness.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema has 0% description coverage and bare properties, so the description must compensate. It adds real meaning to query, key_column, key_value, return_column, and scope with examples and semantic roles. The only gap is folder_path, which is never mentioned, and the description does not explicitly state the mutual exclusivity of query versus explicit key parameters.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with 'ONE-CALL semantic lookup: finds a single cell value anywhere in the workspace and returns it WITH PROVENANCE...' — a specific verb and resource that clearly distinguishes this from generic query or get_cell tools. The two invocation modes (natural language vs explicit keyed) leave no ambiguity about what the tool does.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives explicit guidance on when to choose the natural-language mode versus the explicit keyed mode, including the trigger 'Use this when the query form reports it could not parse, or for values too rare to be sampled.' It also explains how scope narrows the search. However, it does not explicitly compare against sibling tools or state when NOT to use lookup, so it falls just short of full alternative-based guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
queryA
Natural language question with automatic RAG routing. Embeds your question, retrieves candidate sheets by vector similarity, reranks them by lexical overlap with column names and sampled values, fetches the winners LIVE from OneDrive, and returns results. Use for exploratory questions when you do not know which specific file or sheet contains the answer. n_results controls how many sheets are fetched (default 5); min_score drops weak matches. CHECK data.routing: when routing_ambiguous is true the top candidates scored within a tie margin and the choice between them is effectively arbitrary — confirm with inspect_file or ask the user instead of trusting one. Including a distinctive literal value in the question (a client name, a material) strongly improves routing. Response metadata.fetched_at confirms this is live data. For known file/sheet combinations use filter_sheet instead.
| Name | Required | Description | Default |
|---|---|---|---|
| question | Yes | ||
| min_score | No | ||
| n_results | No | ||
| folder_path | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and goes beyond a basic statement of function. It discloses the RAG mechanism, that data is fetched live from OneDrive, and importantly warns about non-deterministic behavior when routing_ambiguous is true, where the choice is 'effectively arbitrary.' This level of behavioral disclosure is exceptional.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is longer than the ideal but every sentence earns its place: purpose, mechanism, usage, parameters, ambiguity warning, routing tip, and alternative tool. It is front-loaded with the primary purpose and structured clearly, though it could be tightened slightly without losing value.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a complex tool with no annotations and a rich output schema, the description covers the essential decision factors: when to use it, how it works, how to interpret routing ambiguity, and how to confirm live data. It also points to the output schema via metadata.fetched_at, making it sufficiently complete for an agent to invoke correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema has zero descriptions, so the description must compensate. It explains n_results (count, default 5) and min_score (threshold for weak matches), and the question parameter is self-evident. However, folder_path is not described at all, leaving its role to inference from its name, which is a small gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states this is a natural-language query tool with automatic RAG routing, describes the full pipeline (embed, retrieve, rerank, fetch live), and explicitly distinguishes it from filter_sheet for known file/sheet combinations. This leaves no ambiguity about what the tool does and when it is the right choice.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly tells the agent to use this tool for exploratory questions when the specific file or sheet is unknown, and names filter_sheet as the alternative for known cases. It also provides actionable guidance for ambiguous routing ('confirm with inspect_file or ask the user') and tips for improving routing accuracy.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
scan_workspaceA
Rescans the OneDrive folder and rebuilds the structure index and embeddings. SLOW — makes many API calls. ONLY call when: new .xlsx files have been added to OneDrive, or existing sheet names or column headers have changed. DO NOT call this at session start. DO NOT call this before every query. The workspace is already indexed from setup. Use get_workspace_graph for instant structure access.
| Name | Required | Description | Default |
|---|---|---|---|
| folder_path | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full transparency burden. It discloses that the operation is SLOW and makes many API calls, and explains that it rebuilds the index and embeddings. It doesn't detail side effects (e.g., does it overwrite the existing index?), but it provides strong behavioral context and performance warnings.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Every sentence earns its place: first states the action, then the performance warning, then precise call conditions, then explicit what-not-to-dos, and finally the alternative tool. It's front-loaded and highly scannable.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Despite its short length, the description covers all decision-relevant context: when to use, when not to use, performance implications, and a link to a faster alternative. Since an output schema exists, the description doesn't need to detail return values. This fully equips an agent to decide correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema shows one optional folder_path parameter with no description, and the description never mentions this parameter. Since schema_description_coverage is 0%, the description should clarify whether folder_path is the OneDrive root or a subfolder, but it does not. The only implicit hint is 'the OneDrive folder', leaving the parameter's behavior ambiguous.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and object ('Rescans the OneDrive folder') and explains the purpose (rebuilds structure index and embeddings). It clearly distinguishes itself from get_workspace_graph by positioning itself as an occasional maintenance operation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives explicit, highly actionable usage criteria: only call when new .xlsx files are added or sheet/column names change, and do not call at session start or before every query. It also points to get_workspace_graph as the instant-access alternative.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
11 tool updates
v0.3.0- First observed
aggregate - First observed
cross_file_aggregate - First observed
derive - First observed
filter_sheet - First observed
get_cell - First observed
get_workspace_graph - First observed
inspect_file - First observed
join_sheets - First observed
lookup - First observed
query - First observed
scan_workspace
TDQS
Scored across 11 tools
Each tool targets a distinct operation: structure discovery, metadata inspection, rescanning, natural-language query, filtered row fetch, grouped aggregation, cross-file totals, joins, signed net calculations, single-cell address reads, and semantic cell lookup. The descriptions include explicit guidance on when to use each tool and warn against alternatives (e.g., aggregate vs. cross_file_aggregate). There is no meaningful overlap or ambiguity between tool purposes.
All tool names follow a consistent verb_noun pattern in snake_case: get_workspace_graph, inspect_file, scan_workspace, filter_sheet, join_sheets, get_cell, etc. Short verb-only names like query, aggregate, derive, and lookup are also consistent in style and fit the pattern of using a single verb when the object is implied. No camelCase or mixing of conventions.
With 11 tools, the server is well-scoped for an Excel workspace analysis tool. Each tool serves a clear and necessary purpose, covering discovery, querying, aggregation, joining, and cell-level access. The count is comfortably within the 3-15 range and does not feel bloated or sparse.
The tool surface covers the full read/analysis lifecycle: workspace structure discovery (get_workspace_graph, inspect_file, scan_workspace), flexible data retrieval (query, filter_sheet, get_cell, lookup), grouped aggregation (aggregate), multi-file totals (cross_file_aggregate), complex net computations (derive), and joins (join_sheets). There are no obvious missing operations for the apparent purpose of reading and analyzing Excel data.
Maintenance
Related MCP Connectors
Official Microsoft MCP Server to query Microsoft Entra data using natural language
A comprehensive Model Context Protocol (MCP) server that enables AI assistants to interact with yo…
Remote streamable-HTTP MCP server running on a single Cloudflare Worker. Your assistant gets live Airbnb, Amazon, Booking.com, Google Flights, Maps and Reddit data, social search on X, Instagram and TikTok, the Meta Ad Library, and image/video generation without any keys. Connect your own accounts to let it send WhatsApp or Telegram messages, work an IMAP inbox, manage Meta Ads campaigns and publish to X and LinkedIn. OAuth 2.1 with PKCE; stored credentials are AES-256-GCM encrypted.
A Model Context Protocol server for Wix AI tools
Related MCP Servers
- AlicenseAqualityDmaintenanceA Model Context Protocol server that enables AI assistants to read from and write to Microsoft Excel files, supporting formats like xlsx, xlsm, xltx, and xltm.69,662 npm1,041MIT
- AlicenseCqualityAmaintenanceA Model Context Protocol server that enables interaction with Microsoft 365 services (Excel, Calendar, Mail, OneDrive, Teams, etc.) through the Graph API, allowing AI assistants to manage Microsoft 365 resources via natural language.18854,304 npm1,028MIT
- AlicenseNot gradedqualityDmaintenanceA Model Context Protocol server that enables AI agents to create, read, and modify Excel workbooks without requiring Microsoft Excel installation.MIT
- AlicenseBqualityDmaintenanceA Model Context Protocol server that enables AI agents to freely operate Excel spreadsheets, providing tools for workbook creation, cell manipulation, formatting, formula handling, and data export.1167 npmISC