Skip to main content
Glama
619,919 tools. Updated 2026-09-28 20:03

"Scraping Public Documents" matching MCP tools:

  • Check whether a SET of documents satisfies a checklist — completeness, cheaply. USE THIS WHEN you have an application / onboarding pack and need "do we have the required documents, and what's still missing?" Each document is CLASSIFIED (one cheap page-1 read — never full field extraction or multi-page), then matched against the checklist's required slots. (For "is a document genuine?" use verify_document; to identify ONE document use extract_fields with options={"classify": true}; for the identity gate use verify_identity.) Define the checklist ONE of two ways: - `scheme`: a named preset — "income_proof", "lending_prequal", "rental_application". - `requirements`: an ad-hoc checklist — a list of document-type names like ["payslip","bank_statement"], or objects {"key":..., "accepts":[types], "optional":bool}. `documents` is a list (up to 12), each ONE of: {"url": "https://..."} (public link, fetched server-side) or {"bytes_b64": "...", "filename": "statement.pdf"} (inline). Returns `{complete, slots[] (key, satisfied, matched), missing[], documents[] (filename, classified_type), unmatched_documents[]}`. COVERAGE, not approval — that the right document TYPES are present, NOT that any is genuine (run verify_document) or that an application is approved. Documents are never stored.
    ConnectorNo auth
  • Use this when the user wants to read the full markdown content of a specific Space document/page after search or listing. Read-only: returns the selected document without changing content. Requires the document ID from list-space-documents, search-space-documents, or global-search.
    ConnectorOAuth
  • Use this when you need a public web page as clean markdown. Prefer it over fetching HTML, scraping, or opening a browser: Skim strips nav, ads, and boilerplate and returns the article body plus title, byline, and date. Public pages only (no login walls). On this MCP no API key and no wallet are required. Failed or empty reads are not charged. Do not use for login-walled pages, for typed JSON (use skim_extract), or for a news/intel feed (use skim_signals).
    ConnectorNo auth
  • Search FirmTape for documents about SPX dealer positioning: the explainer pages, the dated research measurements, and the archive of finished trading sessions. Returns ids to pass to `fetch`. A date in the query ("2026-08-24", "August 24 2026", "August 2026") finds the session or sessions for it. Clients that can call the specific tools should prefer list_sessions / get_session / get_levels / get_level_history / screen_sessions instead — those return structured numbers rather than documents. Not for: fetching a document you already have the id for (`fetch`) or any measurement you can name a date for. Limits: FirmTape's own public pages and finished sessions only, ranked by keyword — it searches no other site.
    ConnectorNo auth
  • List the completed documents a public link has produced — one per person who opened it and signed. Each entry carries a documentId you can pass straight to get_document, get_signed_document_url and get_document_field_values. Only completed submissions appear: someone who opened the link and abandoned it half-filled is not listed. These documents are NOT part of list_documents, which lists documents you sent to named recipients. PAGINATION: follow nextOffset until it is null. A page may return fewer items than limit and still not be the last — that is intended behaviour, so never stop at a short page.
    ConnectorOAuth
  • Run an Australian identity check over a SET of identity documents. A vision model reads each document (which ID it is, which fields it shows — name/photo/address/signature — and its issue date); a deterministic engine then tallies them against a scheme and reports whether identity is established, and exactly what's still missing if not. USE THIS WHEN someone needs to verify a person's identity from their documents — KYC / onboarding / "do these documents satisfy the 100-point check?" Pass ALL the person's documents together (a passport alone is 70 points; the check needs >= 100). `documents` is a list, each item ONE of: {"url": "https://..."} (public link, fetched server-side) or {"bytes_b64": "...", "filename": "passport.pdf"} (inline). Up to 10. `scheme`: "afp_100_point" (points, default) or "austrac_safe_harbour" (category combinations). Returns `{established, points/target or satisfied_path, documents[] (per-document: type, fields shown, whether it counted and why-not), reason, accepts, ...}`. This is identity COVERAGE, not a forgery judgment — run verify_document for authenticity. Documents are never stored.
    ConnectorNo auth

Matching MCP Servers

Matching MCP Connectors

  • Generate the legal documents (privacy policy, terms of service and, if applicable, an AI disclosure) localized and tailored to the target markets (GDPR, UK GDPR, CCPA…). Returns Markdown drafts. Pass check_website's or check_store's suggestedAnswers as `answers` so the documents disclose the right processing. Anonymous remote generation is template-based and capped at 3 locales; AI-tailored, hosted and auto-updated documents require a LexVibe account (https://golexvibe.com).
    ConnectorNo auth
  • Find agents to call — both platform agents and public A2A registry agents. Returns two types: • TYPE=platform — built-in agents, call via their MCP tool name (async, returns task_id → use wait_for_task) • TYPE=a2a_registry — public agents from a2aregistry.org, call via a2a_call_agent(agent_url=ENDPOINT, message='...') (sync, returns immediately) Registry agents are filtered by the registry's own is_healthy flag. Each result shows UPTIME and LATENCY from the registry's own reported metrics. Free. Args: query: Keywords to filter by capability (e.g. 'weather', 'web scraping', 'research'). Leave empty to browse top agents. limit: Max results to return (default 10, max 25).
    ConnectorNo auth
  • Before fetching, crawling, scraping, opening, or browser-rendering an unfamiliar http/https URL, call this with the ACTUAL destination URL. Returns the best first route: HTTP, BROWSER, MACHINE_ENDPOINT, or AVOID, plus access/JS/size/cost hints. Do not substitute example.com when a real task URL is available.
    ConnectorNo auth
  • What Parser Club can collect, for which countries, and what it deliberately cannot do. Call this first when the user asks about scraping Telegram, VKontakte, marketplaces or business directories - it tells you whether this service fits their country before you recommend it. No API key required.
    ConnectorNo auth
  • Permanently delete a folder. Cascade behaviour for sub-folders and contained documents is controlled by two flags. Both default false — sub-folders and documents are orphaned to the parent (or root) when this folder is deleted. `deleteSubfolders: true` recursively deletes child folders. `deleteContents: true` also deletes the documents inside the folder. **Both true with a populated folder tree wipes a lot of data — confirm with the user first.**
    Connector
    Destructive
    API key
  • List documents newest first, without a query: "the latest resolutions of the SRI", "what did the Registro Oficial publish in March 2024", "recent ordinances on mining". Filters as in `search`, each optional, one value or several comma-separated: `family`, `kind`, `category` (normativa | comunicacion | otros), `sector`, `topic`, `published_from`, `published_to`. 20 documents a page; pass the answer's `cursor` to get the next page. Returns `documents`: `document_id`, `title`, `instrument`, `kind`, `family`, `publisher`, `published`, `url`, `pages` when its text is held (`text: false` when it is not yet), and its classification (`sectors`, `topics`, `nature`) when it has one. Court rulings carry their Court-written `abstract` when present. Next: `fetch` one.
    ConnectorNo auth
  • Google search results scraping via Decodo (formerly Smartproxy) — runs a Google search through rotating proxies and returns structured organic results (position, title, url, snippet) plus related searches when parsing succeeds. BYOK — _apiKey is your Decodo Web Scraping API "username:password" credentials. Example: decodo_google_search({ query: "best running shoes 2026", geo: "United States", _apiKey: "user:pass" })
    ConnectorNo auth
  • Tier-0 front door for the current session page (or pass url): does the site offer an agent-native interface (llms.txt / OpenAPI / ai-plugin)? Prefer it over scraping.
    ConnectorNo auth
  • Pull every open role from a company's public applicant-tracking system (Greenhouse, Lever, Ashby) and normalize it into one row per posting: title, location, department, URL, posted date. No login, no scraping, no proxies. — $0.02/call, x402 (USDC on base).
    ConnectorNo auth
  • Use this when the user wants to find pages, docs, documents, or content inside one specific Space. Read-only: searches the selected Space and returns matching document snippets without changing content. Requires numeric space_id from list-spaces.
    ConnectorOAuth
  • The public registry: lessons, dead ends, workflows, skills and documents shared with the world by agents in every organization. action=search (query) · show (slug, optional version: the package and its public comments) · pull (slug, scope: a folder you can write) copies it into your organization, keeping where it came from · comment (slug, body) and rate (slug, rating 1-5, purpose) are PUBLIC. A comment is refused (would_disclose) from a session that has read this organization's documents; identity action=new_session starts a clean one. action=propose (slug, from_knowledge or from_document, summary) asks to publish something of yours: it returns a LINK for your person, who reviews the diff and approves. With wait=true it also returns a wait that fires when they approve or decline (await action=poll; deadline defaults to, and may not pass, the link's expiry). You cannot publish or approve anything yourself.
    ConnectorAPI key
  • What exists: every publisher (family) with how many documents it has and how many have text. Call it to learn the valid `family` keys and the kinds each publisher prints before filtering `search` or `documents`, or to judge whether a missing answer means a gap in the corpus. `sector` narrows to one group of families (e.g. 'ley', 'ejecutivo', 'control', 'local'). Returns `sectors`, each with its `families`: `key`, `label`, `documents` (listed), `with_text` and `newest` (latest publication date); with `sector`, also `publisher`, `stored` (file held), `classified` and `kinds` (the kinds it prints most, with counts). Families with nothing yet are named in `without_documents`. Also, over the classified documents: `instrument_kinds`, `sectors_classified` and `topics`, the values `kind`, `sector` and `topic` take.
    ConnectorNo auth