Skip to main content
Glama

excel_analyze_dataset

Read-onlyIdempotent

Analyze Excel XLSX datasets to profile column types, missing values, numeric summaries, duplicates, outliers, top categories, and correlations without altering the file.

Instructions

Profile an XLSX dataset without changing it: column types, missing values, numeric summaries, duplicate rows, IQR outliers, top categories, and Pearson correlations. Descriptive analysis only; formulas need cached results or native recalculation.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
pathYesAbsolute path or path relative to the server working directory
rangeNo
sheetYes
maxRowsNo
hasHeaderNo
topValuesNo

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observedv0.10.2

TDQS

B3.3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, destructiveHint=false and closed-world, so the safe/non-mutating profile is fully covered by structured data. The description's genuinely additive disclosure is that it performs no recalculation and that formulas only yield values if cached — a real behavioral limitation. Beyond that it says nothing about cost, large-sheet behavior, or truncation at maxRows.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, zero padding; the verb and scope come first and the enumeration is dense but each item conveys an actual output. The caveat sentence is short and earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description does useful work by enumerating the produced statistics, so return content is reasonably covered. But for a 6-parameter tool with 17% schema coverage, the near-total silence on range, maxRows, hasHeader, and topValues leaves real gaps in how to invoke it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 17% (just 'path'), and the description mentions none of range, maxRows, hasHeader, sheet, or topValues. 'Top categories' loosely maps to topValues, but maxRows, hasHeader, and the A1-style range pattern are left entirely undocumented in both schema and prose.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('Profile') plus resource ('XLSX dataset') and enumerates the exact statistics produced (column types, missing values, numeric summaries, duplicates, IQR outliers, top categories, Pearson correlations), which clearly separates it in substance from excel_query_dataset or excel_inspect_workbook. It never names a sibling or makes an explicit comparative claim, so differentiation is inferential rather than stated.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The clause 'Descriptive analysis only; formulas need cached results or native recalculation' gives one real applicability constraint that implies excel_native_batch is the fallback for computed values. However, there is no explicit when-to-use statement and no named alternative, so the agent must infer routing.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.