get_full_text_visual_tool
Downloads a PDF and converts pages to images for AI analysis of graphs, tables, and layouts. Use when text extraction misses formatting.
Instructions
Multimodal Vision Tool. Downloads a PDF and renders the specified number of pages identically into images for the AI to 'look at'. Use this if the user asks you to analyze a graph, table, format, or layout in the paper. If the text extractor doesn't capture formatting, you can use this tool to 'see' the actual PDF! Returns a sequence of texts and images (multimodal format natively parsed). Warning: High token/vision capacity used per page. Default 3 pages.
Input Schema
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | ||
| max_pages | No |