Skip to main content
Glama
459,989 tools. Updated 2026-08-17 12:39

"Tools for Extracting Content from Websites and PDFs" matching MCP tools:

  • Fetch raw HTML content from any URL with optional JavaScript rendering for dynamic websites and Single Page Applications.
    MIT
  • Search the web for current information, news, articles, and websites to find up-to-date content, research topics, or answer questions about recent events.
    Apache 2.0
  • Extract web page content and convert it to clean, readable markdown format for analysis, bypassing paywalls and obtaining structured text data from websites.
    Apache 2.0
  • Convert any public URL to Markdown. Extract readable content from web pages, articles, and documentation as structured Markdown.
    MIT
  • Extract structured financial data from investor relations websites and online sources for investment research when APIs are unavailable.
    MIT
  • Render websites to images, PDFs, HTML, or markdown with full control over viewport, content blocking, and metadata extraction.
    MIT

Matching MCP Servers

  • A
    license
    A
    quality
    B
    maintenance
    Extract content from URLs, documents, videos, and audio files using intelligent auto-engine selection. Supports web pages, PDFs, Word docs, YouTube transcripts, and more with structured JSON responses.
    2
    168
    MIT
  • A
    license
    -
    quality
    C
    maintenance
    An MCP server that exposes any REST API to LLMs through runtime discovery, providing tools to browse endpoints, fetch schemas, and apply business rules without hardcoding or schema duplication.
    MIT

Matching MCP Connectors

  • Decision Layer for AI Agents — 58+ tools, Advisor, MCP. Free key: POST /v1/register {}.

  • GOV.UK Content + Search APIs (every gov.uk page + full search)

  • Extract structured content and layout from PDFs, images, and Office files. Supports output formats like HTML, text, and markdown.
    MIT
  • Initiates an asynchronous crawl of a website, extracting content from multiple pages. Use for comprehensive site coverage; monitor progress with returned operation ID.
    MIT
  • Read the full content of a Bear note by ID or title. Includes text extracted from images and PDFs with clear labeling.
    Apache 2.0
  • Retrieve raw, unprocessed text from PDFs, web pages, pasted text, or YouTube transcripts. Use for direct content export and extraction without AI processing.
    MIT
  • Fetch a PDF from a URL and extract its text content. Returns clean plain text, handling compressed and encrypted PDFs.
    MIT
  • Read text content from OneDrive or SharePoint files given drive and item IDs. Handles Office documents and PDFs, truncates to a specified character limit.
    MIT
  • Retrieve a list of websites from your Umami account, including their IDs, names, and domains, to use in subsequent queries.
    MIT
  • Retrieve complete details for a specific Google ad, including text variations, regional stats, impressions, headlines, descriptions, and image URLs. Essential for analyzing ad content and extracting media for visual analysis.
    MIT
  • Retrieve and extract clean text from web pages and PDFs, including JavaScript-rendered content, with automatic fallback to alternative extraction methods.
    MIT