Skip to main content
Glama
KitchenSink4AI

KitchenSink4Web

Official

Get Text

get_text
Read-only

Extract readable prose from any page or region with pagination, retrieve hidden content when needed, and get clarity on origins—including iframes and shadow DOM.

Instructions

Extract readable prose from a page or one region of it, paginated by start_index so a long article is read in bounded pieces rather than one unbounded dump. Text arrives as labeled data with its origin stated, and hidden regions are stripped and counted rather than silently dropped or silently included. Hidden content IS retrievable, deliberately: include_hidden=true returns it in a separately labeled section with the hiding technique named per block. There is no silent middle tier, because display:none is a real injection channel; the labeled route is the whole design. Prose inside open shadow roots is read, the same as get_page_view reads it; closed roots are counted and stay unreadable. Prose inside same-origin iframes is read after the main document, each frame under a header naming it and its origin, because one page can now deliver text from several documents and a single origin in the label would be a claim about only one of them.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
pageYes
locationNo
max_charsNo
start_indexNo
include_hiddenNo

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observedv1.0.0

TDQS

A4.3/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only declare readOnlyHint and openWorldHint. The description goes far beyond that: it explains pagination mechanics, how hidden content is handled (stripped, counted, and retrievable via include_hidden with per-block technique naming), the behavior for open vs closed shadow roots, and iframe handling with origin labels. It discloses the absence of a 'silent middle tier' and explains the design rationale. This is exceptionally transparent about side effects and edge cases.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single dense paragraph, but every sentence adds meaningful behavioral information. It leads with the core purpose, then covers pagination, hidden content, shadow roots, and iframes. There is no filler or repetition, though it could be tightened by moving some design rationale to a separate note. It is front-loaded and structured logically.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with 5 parameters and no output schema, the description is remarkably thorough: it explains the return format (labeled data with origin), pagination, hidden content, shadow roots, and iframes. The main gap is the 'location' parameter, which is only vaguely referenced as 'one region of it' without defining its structure. Also, the exact interaction of max_chars with pagination isn't spelled out, but the general behavior is clear. Overall it covers most of what an agent needs.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema has zero description coverage for its 5 parameters. The description compensates for start_index (paginated) and include_hidden (returns hidden content in labeled section), and implicitly covers max_chars via 'bounded pieces'. However, the 'location' parameter is never explained beyond 'one region of it', leaving its structure and semantics undocumented. With 0% schema coverage, the description should have covered all parameters; it covers three partially and one not at all.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource: 'Extract readable prose from a page or one region of it'. It immediately distinguishes itself from siblings by mentioning pagination via start_index and explicitly comparing to get_page_view for shadow roots. This is far more specific than a generic 'get text' and leaves no doubt about what the tool does.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives clear context for when to use the tool (for bounded, paginated prose extraction, including hidden content) and even references get_page_view to clarify shadow-root behavior. However, it never explicitly states 'use get_text when you need X, use get_page_view when you need Y' or lists exclusions. The guidance is implicit through behavioral description rather than an explicit decision rule.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.