Skip to main content
Glama
MaurizioLisanti

fatturapa-mcp-server

extract_invoice_data

Extract structured invoice metadata from validated FatturaPA XML documents, parsing headers and line items to return supplier, customer, and invoice fields. Handles missing fields gracefully.

Instructions

Extract key fields from a validated FatturaPA XML document.

Parses the header and first body section to return structured invoice metadata. Never logs or persists XML content — only derived values are returned. Missing fields yield None rather than raising KeyError.

When file_path is given the document is read from disk; the path is checked against the roots configured in FATTURAPA_ALLOWED_ROOTS before any read is attempted. Pass xml_content directly to skip file I/O.

Args: xml_content: Raw XML string of a validated FatturaPA document. ctx: Optional MCP context for structured log emission. file_path: Optional filesystem path to read the document from. Checked against allowed roots before reading.

Returns: An ExtractResult TypedDict with supplier, customer, invoice header fields, and aggregated line_items from all body sections.

Raises: PermissionError: If file_path is outside the configured allowed roots. ValueError: If neither xml_content nor file_path is provided. lxml.etree.XMLSyntaxError: If the XML is not well-formed.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
file_pathNo
xml_contentNo

Output Schema

TableJSON Schema
NameRequiredDescriptionDefault
currencyYes
line_itemsYes
invoice_dateYes
total_amountYes
customer_nameYes
customer_pivaYes
document_typeYes
supplier_nameYes
supplier_pivaYes
invoice_numberYes
customer_tax_codeYes
supplier_tax_codeYes

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observedv0.3.2

TDQS

A4.1/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does so well: it discloses that XML content is never logged or persisted, that missing fields return None instead of raising, that file_path is checked against FATTURAPA_ALLOWED_ROOTS before any read, and it enumerates the exact exception types. This is materially richer than a bare 'extract' claim.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the one-line purpose, then uses conventional Args/Returns/Raises sections, so it is skimmable. The Raises block is slightly heavy for a read-only parse, but every sentence conveys contract detail rather than filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a two-parameter parse tool with an output schema, the description covers input sourcing, the allow-roots security gate, None-on-missing semantics, and all failure modes. An agent has everything needed to call it correctly; only the dual-input precedence edge case is unaddressed, and the output schema handles the return shape.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate, and it defines both parameters (raw validated XML string; optional path validated against allowed roots) plus a ctx argument that appears only in the prose. It stops short of stating precedence when both xml_content and file_path are supplied.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource: 'Extract key fields from a validated FatturaPA XML document,' scoping output to header + first body section. The word 'validated' hints that validate_invoice is a prerequisite, but no sibling is named explicitly, so the agent must infer the pipeline position.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives a real usage fork: pass xml_content to skip file I/O, or file_path to read from disk. However it never states when to prefer this tool over find_invoice_anomalies or generate_invoice_report, and the validation prerequisite is only implied by the adjective 'validated' rather than stated as a condition.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.