Skip to main content
Glama
ahamedfo

IBM Content Services MCP Server

by ahamedfo

get_document_pdf_text

Read-only

Extract embedded text from PDF documents directly from their content bytes, bypassing the need for text extract annotations. Returns page-separated text for documents without existing extraction metadata.

Instructions

Extracts the embedded text of a PDF document directly from its content bytes.

Unlike get_document_text_extract (which depends on text extract annotations created by the repository's text extraction service), this tool downloads the PDF content in memory and extracts its embedded text layer directly. It is read-only: no reservation/lock is placed and nothing is written to disk. Use it when a PDF's text is needed but no text extract annotation exists.

Note: scanned PDFs with no embedded text layer will return empty pages — those require OCR or a vision model instead.

:param identifier: The document id or path (required). This can be either the document's ID (GUID) or its path in the repository (e.g., "/Folder1/document.pdf").

:returns: If successful, returns a dictionary containing: - document_id (str): The document's ID. - page_count (int): Number of pages in the PDF. - characters (int): Total characters extracted. - text (str): The extracted text, with pages separated by form-feed markers. If unsuccessful, returns a ToolError with details about the failure.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
identifierYes

Output Schema

TableJSON Schema
NameRequiredDescriptionDefault
resultYes
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true; the description adds that no reservation/lock is placed and nothing is written to disk, and that scanned PDFs return empty pages, providing useful behavioral context beyond annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise, front-loads the core purpose, and efficiently uses sentences to cover usage, parameters, and return. No extraneous information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the output schema exists and the description explains the return dictionary, and the tool is read-only with a single parameter, the description covers all necessary aspects: purpose, when to use, parameter, and return value.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, but the description fully documents the single required parameter 'identifier', explaining it can be a GUID or path, and provides an example, thus adding complete meaning.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states it extracts embedded text from a PDF's content bytes, and distinguishes itself from get_document_text_extract by explaining the difference in method and dependency on text extract annotations.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly advises using when PDF text is needed but no text extract annotation exists, and warns against using for scanned PDFs without embedded text, recommending OCR or vision models instead.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Install Server

Other Tools

Latest Blog Posts

MCP directory API

We provide all the information about MCP servers via our MCP API.

curl -X GET 'https://glama.ai/api/mcp/v1/servers/ahamedfo/ibm-content-services-mcp-server'

If you have feedback or need assistance with the MCP directory API, please join our Discord server