extract_invoice_data
Extract structured invoice metadata from validated FatturaPA XML documents, parsing headers and line items to return supplier, customer, and invoice fields. Handles missing fields gracefully.
Instructions
Extract key fields from a validated FatturaPA XML document.
Parses the header and first body section to return structured invoice metadata. Never logs or persists XML content — only derived values are returned. Missing fields yield None rather than raising KeyError.
When file_path is given the document is read from disk; the path is
checked against the roots configured in FATTURAPA_ALLOWED_ROOTS before
any read is attempted. Pass xml_content directly to skip file I/O.
Args: xml_content: Raw XML string of a validated FatturaPA document. ctx: Optional MCP context for structured log emission. file_path: Optional filesystem path to read the document from. Checked against allowed roots before reading.
Returns: An ExtractResult TypedDict with supplier, customer, invoice header fields, and aggregated line_items from all body sections.
Raises: PermissionError: If file_path is outside the configured allowed roots. ValueError: If neither xml_content nor file_path is provided. lxml.etree.XMLSyntaxError: If the XML is not well-formed.
Input Schema
| Name | Required | Description | Default |
|---|---|---|---|
| file_path | No | ||
| xml_content | No |
Output Schema
| Name | Required | Description | Default |
|---|---|---|---|
| currency | Yes | ||
| line_items | Yes | ||
| invoice_date | Yes | ||
| total_amount | Yes | ||
| customer_name | Yes | ||
| customer_piva | Yes | ||
| document_type | Yes | ||
| supplier_name | Yes | ||
| supplier_piva | Yes | ||
| invoice_number | Yes | ||
| customer_tax_code | Yes | ||
| supplier_tax_code | Yes |