Skip to main content
Glama

PDF Text Extractor for AI Agents

Server Details

x402 PDF text, pages, and metadata for AI/RAG. $0.0015 per successful PDF.

If you are the author of this connector, you can claim ownership by verifying the domain or GitHub account it belongs to. Claimed connector authors can inspect health checks, view analytics, and manage their listing.
Status
Healthy
Last Tested
Transport
Streamable HTTP · MCP 2025-11-25
URL
Repository
industrial-platform-ai/industrial-platform-agent-tools
GitHub Stars
0

TDQS

A3.9/5.0

Scored across 5 tools

Disambiguation4/5

Each tool targets a distinct resource and action: aborting runs, getting run details, retrieving dataset items, fetching key-value records, and running a specific PDF actor. Boundaries are clear, though the generic Apify management tools and the specialized PDF tool serve different scopes, which could mildly confuse an agent expecting only PDF operations.

Naming Consistency3/5

Four tools follow a consistent kebab-case verb-noun pattern (abort-actor-run, get-actor-run, get-dataset-items, get-key-value-store-record), but the fifth tool uses an irregular actor-ID naming style (industrial_platform--pdf-text-intelligence) with a double hyphen and no verb. The mix is readable but breaks predictable convention.

Tool Count5/5

Five tools is well-scoped for running an asynchronous actor and retrieving results. Each tool earns its place: one starts extraction, two manage run status, and two fetch output from different storage types.

Completeness4/5

The set covers the key steps of triggering PDF extraction, monitoring/aborting runs, and retrieving dataset and key-value outputs. Minor gaps exist—e.g., no tool to list available actors or runs—but typical workflows are covered.

Available Tools

5 tools
abort-actor-runAbort Actor runA
DestructiveIdempotent
Inspect

Abort an Actor run that is currently starting or running. For runs with status SUCCEEDED, FAILED, ABORTING, ABORTED, or TIMED-OUT, this call has no effect. The results will include the updated run details after the abort request.

USAGE:

  • Use when you need to stop a run that is taking too long or misconfigured.

USAGE EXAMPLES:

  • user_input: Abort run y2h7sK3Wc

  • user_input: Gracefully abort run y2h7sK3Wc

This tool requires an x402 payment. Include a valid x402 payment signature in the request metadata (_meta["x402/payment"]). Your MCP client must support the x402 payment protocol.

ParametersJSON Schema
NameRequiredDescriptionDefault
runIdYesThe ID of the Actor run to abort.
gracefullyNoIf true, the Actor run will abort gracefully with a 30-second timeout.

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond the annotations (destructive=false reading is already covered by destructiveHint=true and idempotentHint=true), the description discloses the no-effect status list, that the response carries updated run details, and — critically — that a valid x402 payment signature is required in _meta["x402/payment"]. The payment/auth prerequisite is a major behavioral fact not derivable from structured fields.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the core action and the no-op conditions, then groups usage, examples, and payment requirements cleanly. Slight waste in the two near-identical usage examples, which convey essentially the same invocation.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists so return values need not be spelled out, and the annotations already cover the destructive/idempotent profile; the description still adds the x402 payment requirement and the status no-op rule. Minor remaining gap: what happens to partially produced run results after abort is not stated.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so both runId and gracefully are already documented in the schema, and the description adds no syntax, format, or default details for either. The example 'Gracefully abort run' hints at the gracefully flag but does not explain its semantics beyond the schema. Baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (Abort) and resource (Actor run) plus the qualifying state ('currently starting or running') and the inverse state where it is a no-op. An agent immediately knows what the call does and when it will have any effect, without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides an explicit when-to-use rule ('stop a run that is taking too long or misconfigured') and an explicit when-it-does-nothing list (SUCCEEDED, FAILED, ABORTING, ABORTED, TIMED-OUT). Two usage examples show concrete invocations, so routing is unambiguous.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get-actor-runGet Actor runA
Read-onlyIdempotent
Inspect

Get detailed information about a specific Actor run.

Returns run result: status, storages (datasets/keyValueStores alias map), stats, summary, nextStep.

  • summary describes the past (e.g. "SUCCEEDED in 22s. 47 items; 3 fields available.").

  • nextStep prescribes one primary follow-up action with identifiers interpolated (e.g. "Use get-dataset-items with datasetId=...").

  • waitSecs (0–45, default 30) waits up to that many seconds for terminal status before returning.

USAGE:

  • Use to check the status of a run started by any Actor-running tool.

  • Pass waitSecs > 0 to block until terminal (or until the cap elapses).

USAGE EXAMPLES:

  • user_input: Show details of run y2h7sK3Wc

  • user_input: Wait for run y2h7sK3Wc to finish

This tool requires an x402 payment. Include a valid x402 payment signature in the request metadata (_meta["x402/payment"]). Your MCP client must support the x402 payment protocol.

ParametersJSON Schema
NameRequiredDescriptionDefault
runIdYesThe ID of the Actor run.
waitSecsNoMaximum seconds to wait for the run to reach a terminal state (SUCCEEDED, FAILED, ABORTED, TIMED-OUT). 0 returns immediately with the current status. Cap: 45. Default: 30.

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint/idempotent/non-destructive, so the safety profile is covered. The description adds material context beyond that: an x402 payment signature is required in _meta, and waitSecs blocks up to a 45s cap. It doesn't describe rate limits or failure modes of the payment path, so not a 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the purpose, then organizes USAGE, examples, and the payment caveat into labeled blocks with no filler. The two example user_inputs are mildly redundant with the USAGE section, keeping it from a 5.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, yet the description usefully previews the result shape (summary as past-tense, nextStep as prescribed follow-up) and surfaces the payment prerequisite. An agent has everything needed to call it correctly; minor gaps only in error/payment-failure behavior.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% and both parameters are documented in the schema itself (runId and waitSecs with range, default, cap). The description largely restates the waitSecs semantics rather than adding new meaning, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Get detailed information about a specific Actor run') and describes the returned payload. It is clearly distinguishable from siblings like abort-actor-run (different verb) and get-dataset-items (different resource).

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives explicit context ('Use to check the status of a run started by any Actor-running tool') and a clear rule for the blocking behavior ('Pass waitSecs > 0 to block until terminal'). It does not name when NOT to use it or an alternative status-checking path, so it stops short of a full when/when-not treatment.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get-dataset-itemsGet dataset itemsA
Read-onlyIdempotent
Inspect

Get items (rows) from a dataset — the output/results produced by an Actor run. Returns the rows themselves, not dataset metadata, counts, or a schema. When the user provides a datasetId and asks to retrieve results, output, data, or rows, call this tool directly. Default limit is 20. Use clean=true to skip empty items and hidden fields.

USAGE:

  • Use when you need to read data from a dataset (all items or only selected fields).

USAGE EXAMPLES:

  • user_input: Retrieve results from dataset abc123

  • user_input: Get only metadata.url and title from dataset username~my-dataset

This tool requires an x402 payment. Include a valid x402 payment signature in the request metadata (_meta["x402/payment"]). Your MCP client must support the x402 payment protocol.

ParametersJSON Schema
NameRequiredDescriptionDefault
descNoIf true, results are returned in reverse order (newest to oldest).
omitNoComma-separated list of fields to exclude from results.
cleanNoIf true, returns only non-empty items and skips hidden fields (starting with #). Shortcut for skipHidden=true and skipEmpty=true.
limitNoMaximum number of items to return. Default is 20.
fieldsNoComma-separated list of fields to include in results. Fields in output are sorted as specified. Use dot notation for nested objects (e.g. "metadata.url"); the server auto-flattens parent prefixes.
offsetNoNumber of items to skip at the start. Default is 0.
flattenNoComma-separated list of fields to flatten (e.g. flatten="metadata" turns {"metadata":{"url":"x"}} into {"metadata.url":"x"}). Normally derived automatically from dot-notation in `fields`; specify only as a diagnostic override.
datasetIdYesDataset ID or username~dataset-name.

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover the read-only/idempotent/non-destructive profile, so the description's added value is the payment requirement ('requires an x402 payment... include a valid x402 payment signature in _meta') plus the default limit of 20 and the clean=true shortcut semantics. The payment/auth disclosure is meaningful context annotations cannot express, though return/pagination behavior is left unstated.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core purpose and scope, followed by trigger condition, examples, and payment notice. The USAGE/EXAMPLE blocks are somewhat verbose but each line carries usable information; no wasted restatement of annotations.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists so return values need not be explained, and the key operational extras (payment protocol requirement, default pagination, clean flag) are present. Missing only explicit guidance on alternatives/non-use cases for a fully documented multi-param read tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all eight parameters are already documented in the input schema. The description only echoes limit default and clean semantics, adding no syntax or format detail beyond the schema, which matches the baseline 3 when the schema does the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('get items (rows) from a dataset') and explicitly scopes it against adjacent capabilities: 'Returns the rows themselves, not dataset metadata, counts, or a schema.' An agent can distinguish this from metadata/count-style tools without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives a clear trigger condition ('When the user provides a datasetId and asks to retrieve results, output, data, or rows, call this tool directly') plus two concrete usage examples. It does not name an alternative sibling or state when NOT to use it, so it stops short of the top tier.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get-key-value-store-recordGet key-value store recordA
Read-onlyIdempotent
Inspect

Get the value stored under a specific key in a key-value store — a single record, not a listing of all keys. Requires the exact key name. The response preserves the original Content-Encoding; most clients handle decompression automatically.

USAGE:

  • Use when you need to retrieve a specific record (JSON, text, or binary) from a store.

USAGE EXAMPLES:

  • user_input: Get record INPUT from store abc123

  • user_input: Get record data.json from store username~my-store

This tool requires an x402 payment. Include a valid x402 payment signature in the request metadata (_meta["x402/payment"]). Your MCP client must support the x402 payment protocol.

ParametersJSON Schema
NameRequiredDescriptionDefault
recordKeyYesKey of the record to retrieve.
keyValueStoreIdYesKey-value store ID or username~store-name.

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.4/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond the annotations (readOnly/idempotent/non-destructive), it discloses two non-obvious behaviors: the response preserves the original Content-Encoding, and the call requires an x402 payment signature in _meta["x402/payment"]. Payment/auth requirements and response-encoding caveats are exactly the context annotations cannot convey.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core purpose and scoping constraint, then payment requirements last. The USAGE line partially restates the opening paragraph and the examples are somewhat boilerplate, but nothing is bloated or buried.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With a rich annotation set, full schema coverage, and an output schema (so return values need no explanation), the description supplies everything remaining an agent needs: exact-key requirement, encoding behavior, and the x402 payment prerequisite.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% and both parameters (recordKey, keyValueStoreId) are already documented in the schema. The description only adds the 'requires the exact key name' constraint, which is marginal extra meaning — the baseline 3 applies when the schema does the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The first sentence states a specific verb (get), resource (value/record in a key-value store), and explicitly scopes it as a single record rather than a key listing. An agent can classify the tool without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

"Use when you need to retrieve a specific record (JSON, text, or binary) from a store" gives clear triggering context, and the 'not a listing of all keys' clause rules out a competing interpretation. It stops short of naming a concrete alternative tool or explicit when-not conditions, so it is not a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

industrial_platform--pdf-text-intelligenceindustrial_platform/pdf-text-intelligenceA
Destructive
Inspect

This tool calls the Actor "industrial_platform/pdf-text-intelligence" and retrieves its output results. Actor description: Extract embedded PDF text, pages and metadata. Price: $0.0015 per successful PDF.

This tool requires an x402 payment. Include a valid x402 payment signature in the request metadata (_meta["x402/payment"]). Your MCP client must support the x402 payment protocol.

ParametersJSON Schema
NameRequiredDescriptionDefault
urlsYes**REQUIRED** Public HTTP/HTTPS PDF URLs to process. Up to 50 unique PDFs per run. Example values: ["https://www.w3.org/WAI/ER/tests/xhtml/testfiles/resources/pdf/dummy.pdf"]
waitSecsNoMax seconds (0–45, default 30) to cap the wait for the Actor run to reach terminal state. For long-running Actors the response returns at the cap with the current run status; follow `nextStep` to poll via get-actor-run. Set to 0 to fire-and-forget.
max_pagesNoMaximum pages extracted from each PDF. Example values: 200
concurrencyNoMaximum number of PDFs downloaded and parsed concurrently. Example values: 5
max_text_charsNoMaximum retained text characters per PDF. Example values: 500000
timeout_secondsNoMaximum time allowed to download each PDF. Example values: 45

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare openWorldHint=true, readOnlyHint=false and destructiveHint=true, and the description adds genuinely non-derivable behavior: an x402 payment is required, the cost is $0.0015 per successful PDF, and the payment signature must be placed in _meta. That is exactly the kind of auth/billing context annotations cannot express.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The most valuable content (payment requirement and signature placement) is buried after boilerplate 'calls the Actor and retrieves its output results' text that merely restates the name. It is not bloated, but it is not front-loaded around what the agent most needs to act on.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need no explanation, and the description covers the payment gate, per-PDF pricing, and wait/polling semantics. Minor gaps remain, such as failure/refund behavior on unsuccessful PDFs, but nothing essential to correct invocation is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the six documented parameters already carry their own semantics and the baseline would be 3. The description goes beyond the schema by disclosing a hidden, undocumented input channel (the required x402 payment signature in _meta), which is real added value for correct invocation.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states the concrete operation: it calls the 'industrial_platform/pdf-text-intelligence' Actor and retrieves its output, and the embedded Actor description specifies the resource precisely (embedded PDF text, pages, metadata). An agent can distinguish it from generic siblings like get-dataset-items or get-actor-run, though the framing is boilerplate wrapper language rather than a direct verb+resource statement.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to choose this tool over the sibling run/dataset tools, nor on prerequisites such as URL eligibility beyond what the schema already says. The only conditional information given is the polling hint in waitSecs, which is parameter-level, not tool-selection guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 5 tool updates
    • First observedabort-actor-run
    • First observedget-actor-run
    • First observedget-dataset-items
    • First observedget-key-value-store-record
    • First observedindustrial_platform--pdf-text-intelligence

Related MCP Connectors

Related MCP Servers

Try in Browser

Glama MCP Gateway

Add one secure layer between your agents and this server.