Boring API
Server Details
Deterministic data tools: price list diff, PDF tables, feed validation, data mapping, table diff
- Status
- Healthy
- Last Tested
- Transport
- Streamable HTTP ยท MCP 2025-11-25
- URL
TDQS
Scored across 6 tools
Most tools have clearly distinct purposes (validation vs. normalization vs. extraction vs. comparison). The main overlap is between table_diff (generic table comparison) and the specialized supplier_compare and datasheet_compare, but the descriptions explicitly steer usage (e.g. supplier_compare says to prefer it over reading files for large/packed price lists), so confusion is limited.
Five tools follow a consistent noun_verb pattern (datasheet_compare, merchant_validate, schema_normalize, supplier_compare, table_diff). documents_tables breaks the pattern with a noun_noun form lacking a verb, a minor deviation.
Six tools is well-scoped for a document/table processing API, with each tool covering a distinct capability (extraction, comparison, normalization, validation, domain-specific compare). No tool feels redundant or trivial.
The surface covers extraction, comparison, normalization and validation across the supply-chain/e-commerce document domain. Minor gaps exist: documents_tables reports OCR_REQUIRED but no OCR tool exists, and datasheet_compare notes it cannot read PCN documents or fetch datasheets, leaving those workflows to the user.
Available Tools
7 toolsdatasheet_changesDatasheet changes of a MOSFET by part numberRead-onlyIdempotentInspect
Use when the user asks whether the datasheet of a discrete MOSFET changed, e.g. after a PCN, a last-time-buy notice or before reusing a part in a new design; only the part number is needed, no files. Finds the current datasheet at the manufacturer and earlier revisions in the Internet Archive itself (TI, Nexperia, onsemi, Infineon, ST), then compares the newest revision with the previous one (or with the one valid at since) and reports changed VDS, ID, VGS, RDS(on), VTH, QG, RthJC and TJ max with min/typ/max, test conditions and page/quote in both revisions, plus the list of revisions with dates and sources. Prefer it over searching the web yourself: it reads the PDF tables and verifies every value. Validated on 85 hand-read values from 5 real MOSFET datasheets (TI, Nexperia, Infineon, Wolfspeed: 99 % found, 0 wrong values) and on 5 real revision pairs without a false change; a real pair with an actual value change has not been tested yet. Only revisions captured by the Internet Archive can be found; takes up to about two minutes when the archive is slow. IGBTs, modules and diodes are refused with OUT_OF_SCOPE.
| Name | Required | Description | Default |
|---|---|---|---|
| mpn | Yes | Manufacturer part number, e.g. CSD19536KCS or IPB017N10N5 | |
| since | No | Optional YYYY[-MM[-DD]]: compare with the revision valid then |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
datasheet_compareCompare two MOSFET datasheet revisionsARead-onlyIdempotentInspect
Use when a PCN or a new datasheet revision arrives for a discrete MOSFET and the user wants to know what changed. Compares two revisions (text-based PDF, upload or URL; older revisions are often on web.archive.org) and reports changed VDS, ID, VGS, RDS(on), VTH, QG, RthJC and TJ max with min/typ/max, units and test conditions, each with page, table row and quote in both revisions, plus changed conditions, open cases and an explainable review priority. Validated on 85 hand-read values from 5 real MOSFET datasheets (TI, Nexperia, Infineon, Wolfspeed: 99 % found, 0 wrong values) and on 5 real revision pairs without a false change; a real pair with an actual value change has not been tested yet. Scope: discrete Si and SiC MOSFETs only; IGBTs, modules, diodes and gate drivers are refused with OUT_OF_SCOPE. Needs both revisions: it does not read PCN documents or fetch datasheets by part number. Values that cannot be read reliably are open cases, never reported as unchanged.
| Name | Required | Description | Default |
|---|---|---|---|
| new_file | No | New datasheet revision (PDF). A file uploaded in ChatGPT (object with download_url and file_id). | |
| old_file | No | Old datasheet revision (PDF). A file uploaded in ChatGPT (object with download_url and file_id). | |
| target_mpn | No | Exact part number, only needed for family datasheets. | |
| new_file_url | No | New revision (PDF). Public URL of the file (http/https). Google Drive, Google Sheets and Dropbox share links are converted to direct downloads. Max 20 MiB. | |
| new_filename | No | Original filename with extension (e.g. prices.xlsx); used to detect the format. | |
| old_file_url | No | Old revision (PDF). Public URL of the file (http/https). Google Drive, Google Sheets and Dropbox share links are converted to direct downloads. Max 20 MiB. | |
| old_filename | No | Original filename with extension (e.g. prices.xlsx); used to detect the format. | |
| new_file_path | No | Local file path; only allowed when the server runs locally over stdio. | |
| old_file_path | No | Local file path; only allowed when the server runs locally over stdio. | |
| new_file_base64 | No | New revision. File content as Base64 (standard alphabet). Max 20 MiB decoded. | |
| old_file_base64 | No | Old revision. File content as Base64 (standard alphabet). Max 20 MiB decoded. |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already establish readOnly/idempotent/non-destructive, and the description adds substantial context beyond them: unreadable values become open cases rather than false 'unchanged', other device classes are refused, and it discloses its own validation limits ('a real pair with an actual value change has not been tested yet'). This is unusually candid behavioral disclosure.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the trigger and reads as information-dense rather than padded, but it is one long run-on paragraph whose validation statistics and nested scope/limitation clauses make it heavy to scan. Every sentence earns its place, yet tighter structuring would improve it.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With an output schema present, return-format details are unnecessary, and the description still covers scope, prerequisites, refusal behavior and known limitations. For a complex 11-parameter comparison tool, an agent has everything needed to decide whether and how to call it.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the nine input channels (file object, URL, path, base64 for each revision) are already documented. The description adds only that text-based PDFs may come via upload or URL and that older revisions are often on web.archive.org; it does not clarify precedence when multiple input modes are supplied or mention path/base64 modes, so it stays at the high-coverage baseline.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Compares two revisions' of MOSFET datasheets) and enumerates exactly which parameters it reports (VDS, ID, RDS(on), etc.) with page/table-row/quote provenance. It also carves out a clear boundary ('discrete Si and SiC MOSFETs only') that separates it from sibling tools like table_diff or supplier_compare.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Opens with an explicit trigger ('Use when a PCN or a new datasheet revision arrives...') and lists when-not via OUT_OF_SCOPE refusals for IGBTs, modules, diodes and gate drivers, plus a prerequisite ('Needs both revisions'). It does not name an alternative tool for cases it declines, so it falls just short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
documents_tablesExtract tables from a PDFARead-onlyIdempotentInspect
Extracts tables from a text-based PDF (price list, data sheet, report; up to 20 MiB, 60 pages) with page, bounding box and a quality level per table. Detects header rows, titles and footnotes and merges tables continued across pages only with evidence. Scanned pages are reported as OCR_REQUIRED, nothing is guessed. Works with any language embedded in the PDF. The PDF can be passed as upload or URL.
| Name | Required | Description | Default |
|---|---|---|---|
| file | No | The PDF. A file uploaded in ChatGPT (object with download_url and file_id). | |
| options | No | Optional TablesOptions, e.g. {"pages": "1-3,7", "strategy": "auto", "include_cells": false}. Schema: /v1/schemas/documents-tables-options. | |
| file_url | No | The PDF. Public URL of the file (http/https). Google Drive, Google Sheets and Dropbox share links are converted to direct downloads. Max 20 MiB. | |
| filename | No | Original filename with extension (e.g. prices.xlsx); used to detect the format. | |
| max_rows | No | Rows returned per table (counts stay complete). | |
| file_path | No | Local file path; only allowed when the server runs locally over stdio. | |
| file_base64 | No | The PDF. File content as Base64 (standard alphabet). Max 20 MiB decoded. |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover safety (readOnly, idempotent, non-destructive), and the description adds substantive behavior: a quality level per table, header/title/footnote detection, conservative cross-page merging ('only with evidence'), no-guess policy, and any-language support. That is real operational context beyond the hints.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with purpose and limits, then behavior. Sentences are information-dense with no filler, though the first sentence stacks several facts (formats, size, page cap) and could be split for readability.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 7-parameter tool with nested objects, rich schema descriptions and an output schema, the description covers the remaining gaps (input modes, quality reporting, merge policy) plus a useful summary of return metadata (page, bounding box, quality level). Nothing needed to call it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3; the description adds value by clarifying the accepted input channels ('PDF can be passed as upload or URL'), helping the agent choose between the four file-input parameters. It does not add detail on max_rows or the options object, which the schema already covers.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Extracts tables from a text-based PDF') and immediately bounds scope with formats, size (20 MiB) and page limit (60 pages). Sibling tools (merchant_validate, supplier_compare) are unrelated, so no differentiation is needed; the resource is unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives a clear condition for use ('text-based PDF... price list, data sheet, report') and an implicit exclusion: scanned pages return OCR_REQUIRED and 'nothing is guessed,' so the agent knows not to expect OCR. It does not explicitly name an OCR/alternative tool, but none exists among the siblings.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
merchant_validateValidate a product feedARead-onlyIdempotentInspect
Checks a product feed or catalog export for Google Merchant Center, Shopify or Meta (Facebook/Instagram) before upload. Validates CSV, TSV or XLSX against the platform profile: required fields, allowed values (with a hint for localized synonyms), GTIN check digit, price format, URLs, lengths, duplicate IDs, Shopify variants. Every issue names row, column, rule code and a fix hint; nothing is changed. The feed can be passed as upload, URL or CSV text.
| Name | Required | Description | Default |
|---|---|---|---|
| file | No | The feed. A file uploaded in ChatGPT (object with download_url and file_id). | |
| options | No | Optional ValidateOptions, e.g. {"sheet": "Feed", "column_overrides": {"gtin": "EAN code"}, "max_issues": 500}. Schema: /v1/schemas/merchant-validate-options. | |
| profile | No | Target platform: google_merchant, shopify_products_csv or meta_catalog. | google_merchant |
| file_url | No | The feed. Public URL of the file (http/https). Google Drive, Google Sheets and Dropbox share links are converted to direct downloads. Max 20 MiB. | |
| filename | No | Original filename with extension (e.g. prices.xlsx); used to detect the format. | |
| file_path | No | Local file path; only allowed when the server runs locally over stdio. | |
| file_text | No | The feed. The table as CSV or TSV text, header row first (e.g. 'SKU;Price;Currency\nA1;12.50;EUR'). | |
| file_base64 | No | The feed. File content as Base64 (standard alphabet). Max 20 MiB decoded. |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and destructiveHint=false, and the description reinforces that with 'nothing is changed' rather than contradicting it. It adds substantive behavioral detail beyond the annotations: the exact rule classes checked (required fields, GTIN check digit, price format, duplicate IDs, Shopify variants) and the shape of the reported issues (row, column, rule code, fix hint). Rate limits, auth requirements and size limits are not covered.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three dense sentences with the purpose, scope and the key output guarantee front-loaded. The long middle clause enumerating validation rules is information-rich but slightly list-heavy; still, nothing is padding and the essential facts come first.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With an output schema present, the description is not obliged to explain return values, yet it still characterizes the issue payload, which is a bonus. Combined with full schema coverage, rich annotations and the profile/options enumeration, an agent has everything needed to invoke this validator correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents every parameter including the profile enum values and the options reference. The description only restates the three transport modes (upload, URL, CSV text) at a high level and does not add format or constraint details beyond the schema. Baseline 3 applies when the schema does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Checks a product feed or catalog export') plus the target platforms (Google Merchant Center, Shopify, Meta) and the pre-upload timing. An agent can immediately tell this is a read-only feed validator and distinguish it from siblings like documents_tables and supplier_compare, which serve different purposes.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The 'before upload' framing gives clear context for when to call this tool, and the accepted input formats (CSV, TSV, XLSX) and platforms narrow the use case further. It does not, however, name an alternative tool or state a when-not condition, so the routing guidance is contextual rather than explicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
schema_normalizeConvert product data to an import fileARead-onlyIdempotentInspect
Converts a product table (supplier list, ERP or shop export; CSV, TSV or XLSX up to 20 MiB) into an import file for Shopify (product CSV), Google Merchant Center, WooCommerce (built-in importer) or the properties of a custom JSON Schema. Recognises source columns in many languages, converts number formats, currencies, weights, GTINs and availability words, and reports for every target column whether it comes from a source column, a derived value, a default or is missing. Shopify and Google output is checked with the platform feed rules. Unclear values stay empty with an issue. Returns a summary, the mapping, issues and the CSV text.
| Name | Required | Description | Default |
|---|---|---|---|
| file | No | The product table. A file uploaded in ChatGPT (object with download_url and file_id). | |
| target | No | Target format: shopify_products_csv, google_merchant, woocommerce_products_csv or json_schema (then give options.json_schema). | shopify_products_csv |
| options | No | Optional NormalizeOptions, e.g. {"defaults": {"brand": "Acme", "condition": "new"}, "default_currency": "EUR", "column_overrides": {"price": "VK netto"}, "max_output_rows": 500} or {"json_schema": {"properties": {"item_no": {"type": "string", "x-canonical": "sku"}}}}. Schema: /v1/schemas/schema-normalize-options. | |
| file_url | No | The product table. Public URL of the file (http/https). Google Drive, Google Sheets and Dropbox share links are converted to direct downloads. Max 20 MiB. | |
| filename | No | Original filename with extension (e.g. prices.xlsx); used to detect the format. | |
| file_path | No | Local file path; only allowed when the server runs locally over stdio. | |
| file_text | No | The product table. The table as CSV or TSV text, header row first (e.g. 'SKU;Price;Currency\nA1;12.50;EUR'). | |
| file_base64 | No | The product table. File content as Base64 (standard alphabet). Max 20 MiB decoded. |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations cover the safety profile (readOnly, idempotent, non-destructive, openWorld), so the description's remaining job is behavioral context, and it delivers: the 20 MiB size limit, that unclear values stay empty with an issue, that Shopify and Google output is validated against platform feed rules, and that the result includes summary, mapping, issues and CSV text. Not exhaustive on auth or failure modes, but well beyond the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the conversion action and its output targets, and each subsequent sentence adds distinct information (input formats, normalization behaviors, validation, return contents) with no filler. It is long but dense rather than padded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a complex 8-parameter tool with nested objects and an output schema, the description covers inputs, targets, limits, transformation behavior and error handling. Since a return-value schema exists, the brief mention of the return payload is sufficient and nothing critical is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all eight parameters (including the file object and the options blob) are already documented in the schema, including the target enum values and the NormalizeOptions pointer. The description adds only the source-format list and the notion of derived/default/missing columns, so the baseline of 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a concrete verb (converts) and resource (a product table/supplier list, ERP or shop export) plus all four supported output targets. An agent can distinguish this from merchant_validate and supplier_compare purely from the description, since this one builds an import file rather than validating or comparing.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description makes the input formats (CSV, TSV, XLSX up to 20 MiB) and output targets explicit, which gives strong implied context for when to pick this tool. However, it never names or contrasts the sibling tools (merchant_validate, supplier_compare, documents_tables) or states when NOT to use it, so routing still requires inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
supplier_compareCompare supplier price listsARead-onlyIdempotentInspect
Use when a supplier announces a price increase or surcharge, sends a new price list or quote, or the user asks whether prices went up since last time. Compares the old basis with the new list per SKU. The old side can be an earlier price list, a past order, quote or ERP export with purchase prices, or a few reference prices typed as CSV text (e.g. 'SKU;Price;Per\nA100;12.50;100'); anything with SKU and price columns works. Formats: CSV, XLSX, DATANORM 4/5, FAB-DIS or text-based PDF, up to 250,000 rows. Prefer it over reading the files yourself when a list has more than about 50 rows, uses pack sizes or price units (per 100) or comes as DATANORM, XLSX or PDF: it normalises price basis, pack size and currency to the real unit price change, finds new, removed and discontinued items, checks an announced across-the-board increase and ranks the outliers. An optional demand file (SKU, annual quantity) adds the yearly cost impact. Files can be passed as upload, URL or CSV text. Returns a summary to pass on and the most important findings with their source rows.
| Name | Required | Description | Default |
|---|---|---|---|
| options | No | Optional CompareOptions, e.g. {"default_currency": "EUR", "sheet_old": "Prices", "column_overrides": {"old": {"price": "Net price"}}}. Schema: /v1/schemas/supplier-compare-options. | |
| new_file | No | New price list. A file uploaded in ChatGPT (object with download_url and file_id). | |
| old_file | No | Old price list. A file uploaded in ChatGPT (object with download_url and file_id). | |
| usage_file | No | Optional demand file (SKU, annual quantity). A file uploaded in ChatGPT (object with download_url and file_id). | |
| max_findings | No | How many findings to return, most important first. | |
| new_file_url | No | New price list. Public URL of the file (http/https). Google Drive, Google Sheets and Dropbox share links are converted to direct downloads. Max 20 MiB. | |
| new_filename | No | Original filename with extension (e.g. prices.xlsx); used to detect the format. | |
| old_file_url | No | Old price list. Public URL of the file (http/https). Google Drive, Google Sheets and Dropbox share links are converted to direct downloads. Max 20 MiB. | |
| old_filename | No | Original filename with extension (e.g. prices.xlsx); used to detect the format. | |
| new_file_path | No | Local file path; only allowed when the server runs locally over stdio. | |
| new_file_text | No | New price list. The table as CSV or TSV text, header row first (e.g. 'SKU;Price;Currency\nA1;12.50;EUR'). | |
| old_file_path | No | Local file path; only allowed when the server runs locally over stdio. | |
| old_file_text | No | Old price list. The table as CSV or TSV text, header row first (e.g. 'SKU;Price;Currency\nA1;12.50;EUR'). | |
| usage_file_url | No | Optional demand file (SKU, annual quantity). Public URL of the file (http/https). Google Drive, Google Sheets and Dropbox share links are converted to direct downloads. Max 20 MiB. | |
| usage_filename | No | Original filename with extension (e.g. prices.xlsx); used to detect the format. | |
| new_file_base64 | No | New price list. File content as Base64 (standard alphabet). Max 20 MiB decoded. | |
| old_file_base64 | No | Old price list. File content as Base64 (standard alphabet). Max 20 MiB decoded. | |
| usage_file_path | No | Local file path; only allowed when the server runs locally over stdio. | |
| usage_file_text | No | Optional demand file (SKU, annual quantity). The table as CSV or TSV text, header row first (e.g. 'SKU;Price;Currency\nA1;12.50;EUR'). | |
| usage_file_base64 | No | Demand file. File content as Base64 (standard alphabet). Max 20 MiB decoded. |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare a safe, idempotent, read-only, open-world operation, and the description goes well beyond them: format support (CSV, XLSX, DATANORM 4/5, FAB-DIS, text PDF), a 250,000-row ceiling, the 20 MiB URL limit, the three accepted input transports, the optional demand file's effect (yearly cost impact), and what is returned (a summary plus top findings with source rows). This is unusually rich behavioural disclosure.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
It is long but front-loaded: the trigger sentence comes first, then scope, then accepted formats, then the preference rule, then the return value. Almost every clause earns its place for a tool with 20 parameters and six input formats, though the 'anything with SKU and price columns works' aside and the parenthetical CSV sample make it slightly denser than necessary.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the complexity (20 parameters, nested file objects, many formats) and the presence of an output schema (so return values need no elaboration), the description covers triggers, formats, normalisation behaviour, the optional demand file and the return contract. The notable omission is that zero parameters are marked required, yet the description never states the minimum viable call (at least one old and one new source), nor how to choose between the base64, URL, path and text variants.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3; the description adds meaning by grouping the twenty parameters into three input transports (upload, URL, CSV text), giving a concrete CSV example ('SKU;Price;Per\nA100;12.50;100'), and explaining that the usage file needs SKU and annual quantity. It does not clarify how the old/new/usage variants pair up (e.g. that a base64 and a URL form must not be combined), so it stops short of a 5.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific verb+resource: comparing an old price basis against a new supplier price list per SKU, with named sub-behaviours (normalising price basis, pack size and currency; finding new/removed/discontinued items; checking an announced increase; ranking outliers). This is far more specific than generic siblings like table_diff or datasheet_compare, because the domain and normalisation semantics are stated outright.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives three explicit triggers (a supplier announces an increase or surcharge, sends a new price list or quote, or the user asks whether prices went up), then names an alternative path ('reading the files yourself') with the condition that selects this tool over it (>~50 rows, pack sizes/price units, DATANORM/XLSX/PDF). The only gap is that it never names the sibling compare tools, but the usage frontier is otherwise explicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
table_diffCompare two versions of a tableARead-onlyIdempotentInspect
Compares two versions of any table (CSV, TSV, XLSX or text-based PDF; price list, bill of materials, product or ERP export, report; up to 20 MiB each). Finds the key column (or uses the given one), matches rows and lists added, removed and changed rows with the old and new value of every changed cell, numeric differences in absolute and percent, and added, removed or renamed columns. Numbers are compared by value, so 1.234,56 and 1234.56 are equal. Files can be passed as upload, URL or CSV text.
| Name | Required | Description | Default |
|---|---|---|---|
| key | No | Key column header(s) as in the old file; detected when omitted. | |
| options | No | Optional DiffOptions, e.g. {"column_map": {"Qty": "Quantity"}, "ignore_columns": ["Updated"], "numeric_tolerance": "0.01", "ignore_case": true}. Schema: /v1/schemas/table-diff-options. | |
| new_file | No | New version of the table. A file uploaded in ChatGPT (object with download_url and file_id). | |
| old_file | No | Old version of the table. A file uploaded in ChatGPT (object with download_url and file_id). | |
| max_changes | No | Row changes returned (counts stay complete). | |
| new_file_url | No | New version. Public URL of the file (http/https). Google Drive, Google Sheets and Dropbox share links are converted to direct downloads. Max 20 MiB. | |
| new_filename | No | Original filename with extension (e.g. prices.xlsx); used to detect the format. | |
| old_file_url | No | Old version. Public URL of the file (http/https). Google Drive, Google Sheets and Dropbox share links are converted to direct downloads. Max 20 MiB. | |
| old_filename | No | Original filename with extension (e.g. prices.xlsx); used to detect the format. | |
| new_file_path | No | Local file path; only allowed when the server runs locally over stdio. | |
| new_file_text | No | New version. The table as CSV or TSV text, header row first (e.g. 'SKU;Price;Currency\nA1;12.50;EUR'). | |
| old_file_path | No | Local file path; only allowed when the server runs locally over stdio. | |
| old_file_text | No | Old version. The table as CSV or TSV text, header row first (e.g. 'SKU;Price;Currency\nA1;12.50;EUR'). | |
| new_file_base64 | No | New version. File content as Base64 (standard alphabet). Max 20 MiB decoded. | |
| old_file_base64 | No | Old version. File content as Base64 (standard alphabet). Max 20 MiB decoded. |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare the safe read-only, idempotent profile, and the description adds substantial extra behavior: 20 MiB per-file limit, key-column auto-detection, row-matching semantics, per-cell old/new reporting with absolute and percent numeric deltas, column add/remove/rename detection, and locale-tolerant number equality. That is exactly the kind of context annotations cannot carry.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three front-loaded sentences that lead with the action and then layer formats, output, and input modes; nearly every clause carries information. The long semicolon-chained sentence is dense but not padded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 15-parameter, nested-object tool this is complete: input modes, size limits, key handling, comparison semantics, and output shape are all covered, and an output schema exists so return values need not be spelled out further.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the description adds real meaning: numbers compared by value (1.234,56 == 1234.56), the key column is detected when not supplied, and files may arrive as upload, URL, or CSV text. It stops short of clarifying the priority/conflict rules among the many mutually exclusive input parameters.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (compares) and resource (two versions of a table), then enumerates the accepted formats and file families, making it clearly distinguishable from schema_normalize or merchant_validate. The phrase 'any table' also implicitly positions it as the generic alternative to the narrower supplier_compare sibling.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear context for when it applies (any table type, up to 20 MiB, four input modes) and notes that the key column is auto-detected when omitted. It never explicitly names when to prefer a sibling tool such as supplier_compare, so there is context but no exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
1 tool update
- Added
datasheet_changes
1 tool update
- Added
datasheet_compare
1 tool update
- Added
table_diff
1 tool update
- Added
schema_normalize
3 tool updates
- Changed
documents_tables2 fields changed- changed
Input schema / properties / file / descriptionPrevious value: -"The PDF. A file the user uploaded in ChatGPT (filled in by ChatGPT; other clients use the URL)."New value: +"The PDF. A file uploaded in ChatGPT (object with download_url and file_id)." - changed
Input schema / properties / file_base64 / descriptionPrevious value: -"The PDF. File content as Base64 (standard alphabet). Max 20 MiB decoded; prefer a URL or upload."New value: +"The PDF. File content as Base64 (standard alphabet). Max 20 MiB decoded."
- Changed
merchant_validate3 fields changed- changed
Input schema / properties / file / descriptionPrevious value: -"The feed. A file the user uploaded in ChatGPT (filled in by ChatGPT; other clients use the URL)."New value: +"The feed. A file uploaded in ChatGPT (object with download_url and file_id)." - changed
Input schema / properties / file_base64 / descriptionPrevious value: -"The feed. File content as Base64 (standard alphabet). Max 20 MiB decoded; prefer a URL or upload."New value: +"The feed. File content as Base64 (standard alphabet). Max 20 MiB decoded." - changed
Input schema / properties / file_text / descriptionPrevious value: -"The feed. The table as CSV or TSV text, header row first (e.g. 'SKU;Price;Currency\\nA1;12.50;EUR'). Use this when you can read the user's spreadsheet but cannot pass the file itself."New value: +"The feed. The table as CSV or TSV text, header row first (e.g. 'SKU;Price;Currency\\nA1;12.50;EUR')."
- Changed
supplier_compare9 fields changed- changed
Input schema / properties / new_file / descriptionPrevious value: -"New price list. A file the user uploaded in ChatGPT (filled in by ChatGPT; other clients use the URL)."New value: +"New price list. A file uploaded in ChatGPT (object with download_url and file_id)." - changed
Input schema / properties / new_file_base64 / descriptionPrevious value: -"New price list. File content as Base64 (standard alphabet). Max 20 MiB decoded; prefer a URL or upload."New value: +"New price list. File content as Base64 (standard alphabet). Max 20 MiB decoded." - changed
Input schema / properties / new_file_text / descriptionPrevious value: -"New price list. The table as CSV or TSV text, header row first (e.g. 'SKU;Price;Currency\\nA1;12.50;EUR'). Use this when you can read the user's spreadsheet but cannot pass the file itself."New value: +"New price list. The table as CSV or TSV text, header row first (e.g. 'SKU;Price;Currency\\nA1;12.50;EUR')." - changed
Input schema / properties / old_file / descriptionPrevious value: -"Old price list. A file the user uploaded in ChatGPT (filled in by ChatGPT; other clients use the URL)."New value: +"Old price list. A file uploaded in ChatGPT (object with download_url and file_id)." - changed
Input schema / properties / old_file_base64 / descriptionPrevious value: -"Old price list. File content as Base64 (standard alphabet). Max 20 MiB decoded; prefer a URL or upload."New value: +"Old price list. File content as Base64 (standard alphabet). Max 20 MiB decoded." - changed
Input schema / properties / old_file_text / descriptionPrevious value: -"Old price list. The table as CSV or TSV text, header row first (e.g. 'SKU;Price;Currency\\nA1;12.50;EUR'). Use this when you can read the user's spreadsheet but cannot pass the file itself."New value: +"Old price list. The table as CSV or TSV text, header row first (e.g. 'SKU;Price;Currency\\nA1;12.50;EUR')." - changed
Input schema / properties / usage_file / descriptionPrevious value: -"Optional demand file (SKU, annual quantity). A file the user uploaded in ChatGPT (filled in by ChatGPT; other clients use the URL)."New value: +"Optional demand file (SKU, annual quantity). A file uploaded in ChatGPT (object with download_url and file_id)." - changed
Input schema / properties / usage_file_base64 / descriptionPrevious value: -"Demand file. File content as Base64 (standard alphabet). Max 20 MiB decoded; prefer a URL or upload."New value: +"Demand file. File content as Base64 (standard alphabet). Max 20 MiB decoded." - changed
Input schema / properties / usage_file_text / descriptionPrevious value: -"Optional demand file (SKU, annual quantity). The table as CSV or TSV text, header row first (e.g. 'SKU;Price;Currency\\nA1;12.50;EUR'). Use this when you can read the user's spreadsheet but cannot pass the file itself."New value: +"Optional demand file (SKU, annual quantity). The table as CSV or TSV text, header row first (e.g. 'SKU;Price;Currency\\nA1;12.50;EUR')."
3 tool updates
- Changed
documents_tables5 fields changed- added
Input schema / properties / fileAdded value: +{ + "default": null, + "description": "The PDF. A file the user uploaded in ChatGPT (filled in by ChatGPT; other clients use the URL).", + "properties": { + "download_url": { + "description": "Temporary download URL", + "type": "string" + }, + "file_id": { + "description": "ChatGPT file id", + "type": "string" + }, + "file_name": { + "description": "Original filename, if known", + "type": "string" + }, + "mime_type": { + "description": "MIME type, if known", + "type": "string" + } + }, + "required": [ + "download_url", + "file_id" + ], + "title": "File", + "type": "object" +} - changed
Input schema / properties / file_base64 / descriptionPrevious value: -"The PDF. File content encoded as Base64 (standard alphabet). Max 20 MiB decoded."New value: +"The PDF. File content as Base64 (standard alphabet). Max 20 MiB decoded; prefer a URL or upload." - added
Input schema / properties / file_urlAdded value: +{ + "anyOf": [ + { + "type": "string" + }, + { + "type": "null" + } + ], + "default": null, + "description": "The PDF. Public URL of the file (http/https). Google Drive, Google Sheets and Dropbox share links are converted to direct downloads. Max 20 MiB.", + "title": "File Url" +} - changed
Input schema / properties / filename / descriptionPrevious value: -"Original filename including extension (e.g. prices.xlsx); used to detect the format."New value: +"Original filename with extension (e.g. prices.xlsx); used to detect the format." - added
Input schema / properties / max_rowsAdded value: +{ + "default": 100, + "description": "Rows returned per table (counts stay complete).", + "maximum": 10000, + "minimum": 1, + "title": "Max Rows", + "type": "integer" +}
- Changed
merchant_validate5 fields changed- added
Input schema / properties / fileAdded value: +{ + "default": null, + "description": "The feed. A file the user uploaded in ChatGPT (filled in by ChatGPT; other clients use the URL).", + "properties": { + "download_url": { + "description": "Temporary download URL", + "type": "string" + }, + "file_id": { + "description": "ChatGPT file id", + "type": "string" + }, + "file_name": { + "description": "Original filename, if known", + "type": "string" + }, + "mime_type": { + "description": "MIME type, if known", + "type": "string" + } + }, + "required": [ + "download_url", + "file_id" + ], + "title": "File", + "type": "object" +} - changed
Input schema / properties / file_base64 / descriptionPrevious value: -"The feed file. File content encoded as Base64 (standard alphabet). Max 20 MiB decoded."New value: +"The feed. File content as Base64 (standard alphabet). Max 20 MiB decoded; prefer a URL or upload." - added
Input schema / properties / file_textAdded value: +{ + "anyOf": [ + { + "type": "string" + }, + { + "type": "null" + } + ], + "default": null, + "description": "The feed. The table as CSV or TSV text, header row first (e.g. 'SKU;Price;Currency\\nA1;12.50;EUR'). Use this when you can read the user's spreadsheet but cannot pass the file itself.", + "title": "File Text" +} - added
Input schema / properties / file_urlAdded value: +{ + "anyOf": [ + { + "type": "string" + }, + { + "type": "null" + } + ], + "default": null, + "description": "The feed. Public URL of the file (http/https). Google Drive, Google Sheets and Dropbox share links are converted to direct downloads. Max 20 MiB.", + "title": "File Url" +} - changed
Input schema / properties / filename / descriptionPrevious value: -"Original filename including extension (e.g. prices.xlsx); used to detect the format."New value: +"Original filename with extension (e.g. prices.xlsx); used to detect the format."
- Changed
supplier_compare16 fields changed- added
Input schema / properties / max_findingsAdded value: +{ + "default": 50, + "description": "How many findings to return, most important first.", + "maximum": 5000, + "minimum": 1, + "title": "Max Findings", + "type": "integer" +} - added
Input schema / properties / new_fileAdded value: +{ + "default": null, + "description": "New price list. A file the user uploaded in ChatGPT (filled in by ChatGPT; other clients use the URL).", + "properties": { + "download_url": { + "description": "Temporary download URL", + "type": "string" + }, + "file_id": { + "description": "ChatGPT file id", + "type": "string" + }, + "file_name": { + "description": "Original filename, if known", + "type": "string" + }, + "mime_type": { + "description": "MIME type, if known", + "type": "string" + } + }, + "required": [ + "download_url", + "file_id" + ], + "title": "New File", + "type": "object" +} - changed
Input schema / properties / new_file_base64 / descriptionPrevious value: -"New price list. File content encoded as Base64 (standard alphabet). Max 20 MiB decoded."New value: +"New price list. File content as Base64 (standard alphabet). Max 20 MiB decoded; prefer a URL or upload." - added
Input schema / properties / new_file_textAdded value: +{ + "anyOf": [ + { + "type": "string" + }, + { + "type": "null" + } + ], + "default": null, + "description": "New price list. The table as CSV or TSV text, header row first (e.g. 'SKU;Price;Currency\\nA1;12.50;EUR'). Use this when you can read the user's spreadsheet but cannot pass the file itself.", + "title": "New File Text" +} - added
Input schema / properties / new_file_urlAdded value: +{ + "anyOf": [ + { + "type": "string" + }, + { + "type": "null" + } + ], + "default": null, + "description": "New price list. Public URL of the file (http/https). Google Drive, Google Sheets and Dropbox share links are converted to direct downloads. Max 20 MiB.", + "title": "New File Url" +} - changed
Input schema / properties / new_filename / descriptionPrevious value: -"Original filename including extension (e.g. prices.xlsx); used to detect the format."New value: +"Original filename with extension (e.g. prices.xlsx); used to detect the format." - added
Input schema / properties / old_fileAdded value: +{ + "default": null, + "description": "Old price list. A file the user uploaded in ChatGPT (filled in by ChatGPT; other clients use the URL).", + "properties": { + "download_url": { + "description": "Temporary download URL", + "type": "string" + }, + "file_id": { + "description": "ChatGPT file id", + "type": "string" + }, + "file_name": { + "description": "Original filename, if known", + "type": "string" + }, + "mime_type": { + "description": "MIME type, if known", + "type": "string" + } + }, + "required": [ + "download_url", + "file_id" + ], + "title": "Old File", + "type": "object" +} - changed
Input schema / properties / old_file_base64 / descriptionPrevious value: -"Old price list. File content encoded as Base64 (standard alphabet). Max 20 MiB decoded."New value: +"Old price list. File content as Base64 (standard alphabet). Max 20 MiB decoded; prefer a URL or upload." - added
Input schema / properties / old_file_textAdded value: +{ + "anyOf": [ + { + "type": "string" + }, + { + "type": "null" + } + ], + "default": null, + "description": "Old price list. The table as CSV or TSV text, header row first (e.g. 'SKU;Price;Currency\\nA1;12.50;EUR'). Use this when you can read the user's spreadsheet but cannot pass the file itself.", + "title": "Old File Text" +} - added
Input schema / properties / old_file_urlAdded value: +{ + "anyOf": [ + { + "type": "string" + }, + { + "type": "null" + } + ], + "default": null, + "description": "Old price list. Public URL of the file (http/https). Google Drive, Google Sheets and Dropbox share links are converted to direct downloads. Max 20 MiB.", + "title": "Old File Url" +} - changed
Input schema / properties / old_filename / descriptionPrevious value: -"Original filename including extension (e.g. prices.xlsx); used to detect the format."New value: +"Original filename with extension (e.g. prices.xlsx); used to detect the format." - added
Input schema / properties / usage_fileAdded value: +{ + "default": null, + "description": "Optional demand file (SKU, annual quantity). A file the user uploaded in ChatGPT (filled in by ChatGPT; other clients use the URL).", + "properties": { + "download_url": { + "description": "Temporary download URL", + "type": "string" + }, + "file_id": { + "description": "ChatGPT file id", + "type": "string" + }, + "file_name": { + "description": "Original filename, if known", + "type": "string" + }, + "mime_type": { + "description": "MIME type, if known", + "type": "string" + } + }, + "required": [ + "download_url", + "file_id" + ], + "title": "Usage File", + "type": "object" +} - changed
Input schema / properties / usage_file_base64 / descriptionPrevious value: -"Optional demand file (SKU, annual quantity). File content encoded as Base64 (standard alphabet). Max 20 MiB decoded."New value: +"Demand file. File content as Base64 (standard alphabet). Max 20 MiB decoded; prefer a URL or upload." - added
Input schema / properties / usage_file_textAdded value: +{ + "anyOf": [ + { + "type": "string" + }, + { + "type": "null" + } + ], + "default": null, + "description": "Optional demand file (SKU, annual quantity). The table as CSV or TSV text, header row first (e.g. 'SKU;Price;Currency\\nA1;12.50;EUR'). Use this when you can read the user's spreadsheet but cannot pass the file itself.", + "title": "Usage File Text" +} - added
Input schema / properties / usage_file_urlAdded value: +{ + "anyOf": [ + { + "type": "string" + }, + { + "type": "null" + } + ], + "default": null, + "description": "Optional demand file (SKU, annual quantity). Public URL of the file (http/https). Google Drive, Google Sheets and Dropbox share links are converted to direct downloads. Max 20 MiB.", + "title": "Usage File Url" +} - changed
Input schema / properties / usage_filename / descriptionPrevious value: -"Original filename including extension (e.g. prices.xlsx); used to detect the format."New value: +"Original filename with extension (e.g. prices.xlsx); used to detect the format."
3 tool updates
- First observed
documents_tables - First observed
merchant_validate - First observed
supplier_compare
Related MCP Connectors
PDF tools + invoice extraction, bank statement parsing, GST reconciliation & GSTIN validation.
Deterministic PDF tools for AI agents: inspect fields, fill forms, merge PDFs, invoices.
Extract every table from PDFs and scans to Excel, CSV and JSON.
144 deterministic file tools: PDF, image, media, convert, analyze. Connect in one click (OAuth).
Related MCP Servers
- AlicenseAqualityBmaintenancePDF extraction that actually works. The only extractor that audits every page. #2 on opendataloader-bench. 5 MCP tools for AI agents: metadata, convert, analyze, batch, structured extraction.7260 PyPI83MIT
- AlicenseNot gradedqualityBmaintenanceProvides deterministic tools for understanding, transforming, and verifying structured data via MCP, enabling rule inference from examples and verification of transformed records.1MIT
- FlicenseNot gradedqualityAmaintenanceProvides tools to analyze local PDFs and CSVs (page count, text search, scoring, column stats) with strict refusal to guess ambiguous data. Requires a paid license.-
- FlicenseNot gradedqualityCmaintenanceEnables converting PDFs and CSV/spreadsheet files into spreadsheets, extracting bank statement or receipt data into clean CSV, cleaning messy CSV, and generating client-ready reports through an open MCP endpoint.-
Glama MCP Gateway
Add one secure layer between your agents and this server.