Skip to main content
Glama
Dasistaiden

whed-tools

by Dasistaiden

WHED Tools — Higher Education Intelligence Pipeline

An MCP-native pipeline for collecting structured intelligence on higher education institutions, aligned with the IAU World Higher Education Database (WHED) schema.

Scrape → Extract → Validate → Save — the Host LLM performs extraction directly using MCP tools. No external LLM required.

Built on samirsaci/mcp-webscraper.


Overview

Step

How

Scrape

MCP crawl_website or standalone run_scraper.py — schema-driven crawl, PDF extraction

Extract

Host LLM reads scraped content, uses get_extraction_schema + get_db_context

Validate

validate_profile — Pydantic schema + WHED DB picklist checks

Save

save_profile — write to output/structured/


Related MCP server: agent-api-gateway-mcp

Architecture

┌─────────────────────────────────────────────────────────────────┐
│  HOST LLM  (Claude in Cursor / any MCP client)                  │
│                                                                 │
│  crawl_website(url)      →  get_extraction_schema()             │
│  scrape_url(url)             get_db_context(domain)            │
│                                     │                           │
│  Host LLM reads content and fills JSON                          │
│                                     │                           │
│  validate_profile(json)  →  save_profile(domain, json)           │
└─────────────────────────────────────────────────────────────────┘
         │                      │                       │
         ▼                      ▼                       ▼
   output/pages/           schema.py              output/structured/
   output/sites/           db_reference.py

Project Structure

mcp-webscraper/
├── MCP_server/
│   ├── server.py           # MCP entry — 9 tools (scrape + extraction)
│   ├── models/
│   └── utils/
│       └── web_scraper.py  # Scraper (static, Playwright, pdfplumber)
├── schema.py               # SchoolProfile, EXTRACTION_TEMPLATE, FIELD_URL_HINTS
├── db_reference.py         # WHED DB — picklists, reference examples, ground truth
├── run_scraper.py          # Standalone CLI — schema-driven crawl, PDF extraction
├── docs/
│   ├── USAGE_GUIDE.md      # Architecture, flow, outputs, comparison
│   ├── PROJECT_ITERATIONS.md
│   └── MCP_VS_N8N_COMPARISON.md
└── output/
    ├── pages/              # Per-page cache from crawl
    ├── sites/              # Combined site crawl
    ├── structured/         # MCP extraction output
    ├── ground_truth/       # WHED DB exports
    └── stages/             # Human review staging

Prerequisites

  • Python 3.10+

  • uv package manager

  • Cursor (for MCP usage)

  • MySQL with WHED database (optional — for DB grounding and comparison)


Installation

git clone https://github.com/your-username/mcp-webscraper.git
cd mcp-webscraper
uv sync
uv run playwright install chromium

Copy .env.example to .env and add WHED DB credentials (if available).

Connect MCP to Cursor

Add to .cursor/mcp.json:

{
  "mcpServers": {
    "whed-tools": {
      "command": "uv",
      "args": [
        "run",
        "--directory",
        "/path/to/mcp-webscraper",
        "python",
        "MCP_server/server.py"
      ]
    }
  }
}

MCP Tools (whed-tools)

Tool

Description

scrape_url

Fetch HTML from a URL

extract_data

Extract by CSS selector

extract_first

First matching element

batch_scrape

Multiple URLs

crawl_website

Discover and crawl site (schema_filter=True to skip irrelevant pages)

extract_pdf_text

Download a PDF and extract its text content

get_extraction_schema

WHED field template (REQUIRED only)

get_db_context

Picklists + reference example for domain

validate_profile

Pydantic + DB picklist validation

save_profile

Save profile to output/structured/

Example prompt

"Crawl https://www.example.edu and extract a WHED profile. Use get_extraction_schema and get_db_context, then validate and save."


Standalone Scripts

Scrape (schema-driven, with PDFs)

Edit run_scraper.py (TARGET_URL, MODE, etc.), then:

uv run python run_scraper.py
  • Uses schema.FIELD_URL_HINTS to follow only relevant URLs

  • Extracts text from PDFs via pdfplumber


Schema & DB Grounding

  • REQUIRED fields are in EXTRACTION_TEMPLATE; DEFERRED fields are in Pydantic but not prompted.

  • With WHED DB: picklists, few-shot examples, and post-validation reduce hallucination.

  • Edit schema.py to add or reactivate fields.


Documentation

Doc

Content

USAGE_GUIDE.md

Architecture, flow, and outputs

PROJECT_ITERATIONS.md

Evolution from Ollama to MCP-native

MCP_VS_N8N_COMPARISON.md

KPI comparison with N8N + Firecrawl


License

MIT — based on samirsaci/mcp-webscraper.

Available Tools

10 tools
batch_scrapeA
Scrape multiple URLs efficiently.

Args:
    urls: List of URLs to scrape
    javascript: Set to True if the sites need JavaScript rendering
    save_path: Optional file path to save all results as a JSON array (e.g. "C:/Users/me/Desktop/batch.json").
               If a directory is given, a timestamped filename is generated automatically.

Returns:
    List of scraping results for each URL
ParametersJSON Schema
NameRequiredDescriptionDefault
urlsYes
save_pathNo
javascriptNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A3.5/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description must carry the full behavioral burden. It discloses the return shape and the save_path file-writing side effect, but says nothing about rate limits, concurrency, per-URL error handling, permissions, or whether partial failures return partial results.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is front-loaded with a one-line purpose and then structured into Args and Returns sections. Every element is short and useful, with only the word "efficiently" adding minor fluff.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given no annotations and a batch operation, the definition is only partly complete. It covers all parameters and the output schema exists, but it lacks behavioral context on failure modes and usage guidance for choosing this tool over siblings.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, so the description must compensate, and it does: it explains javascript as a JS-rendering switch and save_path as optional with a concrete file-path example and timestamped-directory behavior. It omits URL format constraints and the javascript default, keeping it just short of perfect.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

"Scrape multiple URLs efficiently" states a specific verb (scrape) and resource (URLs), and the word "multiple" distinguishes the batch scope from the single-URL sibling scrape_url. An agent can identify what the tool does without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description offers no when-to-use guidance, no exclusions, and never names an alternative such as scrape_url for single URLs or crawl_website for site traversal. The only implicit cue is the name itself, which is not enough to route between siblings.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

crawl_websiteA
Crawl a website to discover its structure and pages.

Args:
    start_url: Starting URL
    max_pages: Maximum pages to crawl (default 50)
    max_depth: Maximum link depth (default 3)
    same_domain_only: Stay on same domain (default True)
    schema_filter: When True, only follow URLs whose path contains keywords
                   relevant to the WHED schema (about, contact, course, program,
                   faculty, etc.). Skips login pages, media, and unrelated content.
                   Recommended for institution websites.
    save_path: Optional file path to save the site map as JSON (e.g. "C:/Users/me/Desktop/sitemap.json").
               If a directory is given, a timestamped filename is generated automatically.

Returns:
    Site map with discovered pages and statistics
ParametersJSON Schema
NameRequiredDescriptionDefault
max_depthNo
max_pagesNo
save_pathNo
start_urlYes
schema_filterNo
same_domain_onlyNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A3.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does disclose real behavior: schema_filter skips login pages, media and unrelated content, and save_path auto-generates a timestamped filename when given a directory. It omits important operational traits such as rate limiting/politeness, robots.txt handling, authentication needs, and whether the crawl is strictly read-only.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The opening sentence front-loads the purpose, and the Args/Returns docstring layout is scannable and efficient. The Returns line is slightly redundant given an output schema exists, but overall it is well-sized with no padding.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a six-parameter crawl tool with an output schema and no annotations, the description covers purpose, every parameter, and the notable filtering/saving behaviors. The remaining gap is operational safety context (read-only nature, throttling, domain restrictions beyond same_domain_only), which slightly reduces confidence.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate, and it does: all six parameters are documented with defaults and semantics, including the non-obvious schema_filter keyword behavior and the directory-vs-file behavior of save_path. Only start_url and the two numeric limits get minimal treatment, but nothing is left ambiguous.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Crawl a website to discover its structure and pages'), which is clear and distinct from pure content extraction. However, it never differentiates itself from close siblings like scrape_url or batch_scrape, so an agent must infer that this tool maps the site rather than fetching content.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is implied by the crawl-and-discover framing, and schema_filter carries a 'Recommended for institution websites' hint, which is useful context. But there is no explicit statement of when to use this over scrape_url/batch_scrape, nor any prerequisites or exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

extract_dataA
Scrape a webpage and extract specific data using CSS selectors.

Args:
    url: The webpage to scrape
    css_selectors: List of CSS selectors (e.g., ["h1", "a.link", "#content"])
    attributes: List of attributes to extract for each selector (e.g., ["text", "href", "text"])
               If not provided, defaults to "text" for all selectors
    javascript: Set to True for JavaScript-rendered sites
    save_path: Optional file path to save the result as JSON (e.g. "C:/Users/me/Desktop/result.json").
               If a directory is given, a timestamped filename is generated automatically.

Returns:
    Dictionary with extracted data for each selector

Example:
    extract_data(
        url="https://example.com",
        css_selectors=["h1", "a"],
        attributes=["text", "href"]
    )
ParametersJSON Schema
NameRequiredDescriptionDefault
urlYes
save_pathNo
attributesNo
javascriptNo
css_selectorsYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full load, and it does so well: it discloses the default attribute ('text') when attributes is omitted, that javascript=True is needed for JS-rendered sites, and that a directory passed to save_path triggers an auto-generated timestamped filename. It omits any mention of auth requirements, rate limits, or error behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core purpose, then cleanly partitioned into Args/Returns/Example sections. The worked example earns its place by clarifying the two parallel list arguments, though the Args block is somewhat verbose for an agent runtime.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a five-parameter scraping tool with no annotations, the description covers inputs, defaults, output shape, and a usage example. Since an output schema exists, the Returns line is bonus rather than necessity, and the only real gap is operational context (auth, rate limits, failure modes).

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate, and it documents all five parameters with concrete semantics and inline examples (selectors list, parallel attributes list, url target, save path formats). Notably it explains the positional correspondence between css_selectors and attributes, which the schema does not.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The first sentence states a specific verb and resource (scrape a webpage, extract data) plus the mechanism (CSS selectors), which is considerably more precise than the sibling names alone. It does not explicitly contrast itself against close siblings like extract_first or batch_scrape, but the mechanism and single-page scope make its purpose clear.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is implied rather than stated: the 'javascript' arg signals when to use it for JS-rendered sites and 'save_path' signals persistence, but there is no explicit rule for when to prefer this over scrape_url, batch_scrape, or crawl_website. The agent must infer routing from the sibling names.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

extract_firstA
Extract the first matching element from a webpage.
Useful for getting single values like page title, main heading, etc.

Args:
    url: The webpage to scrape
    css_selector: CSS selector for the element (e.g., "h1", "title", "meta[name='description']")
    attribute: What to extract - "text" for content, or attribute name like "href", "content", "src"
    javascript: Set to True for JavaScript-rendered sites
    save_path: Optional file path to save the result as JSON (e.g. "C:/Users/me/Desktop/result.json").
               If a directory is given, a timestamped filename is generated automatically.

Returns:
    Dictionary with the extracted value

Example:
    extract_first(url="https://example.com", css_selector="title", attribute="text")
ParametersJSON Schema
NameRequiredDescriptionDefault
urlYes
attributeNotext
save_pathNo
javascriptNo
css_selectorYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A3.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden, and it does add real context: the javascript flag for rendered sites and the save_path directory-to-timestamped-filename behavior. However, it is silent on critical scraper behavior such as what happens when no element matches (error vs. null), rate limiting, or network/permission requirements.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Purpose is front-loaded in the first sentence, followed by a tightly organized Args/Returns/Example block; each section carries distinct information. The docstring formatting is slightly verbose for a tool description, but nothing is wasted.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so the description need not detail return values, yet it still gives a concise 'Dictionary with the extracted value' summary. For a 5-parameter, unannotated scraping tool the main remaining gap is failure/miss behavior when the selector matches nothing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% and there are 5 parameters, so the description must compensate, and it largely does: every parameter is explained, including valid values for attribute ('text' vs. an attribute name like href/content/src) and CSS selector examples. It omits the default for javascript and does not elaborate on save_path's JSON format beyond one example.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource ('Extract the first matching element from a webpage') and clarifies the single-value use case, which distinguishes it implicitly from bulk siblings like extract_data or batch_scrape. It never names a sibling explicitly, so an agent must infer the boundary from the phrase 'first matching element'.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The line 'Useful for getting single values like page title, main heading, etc.' implies when the tool is appropriate, but there is no explicit guidance about when not to use it or how it differs from extract_data or scrape_url. Usage is suggested rather than specified.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

extract_pdf_textA
Download a PDF and extract its text content using pdfplumber.
Useful for reading course handbooks, academic calendars, prospectuses,
and other PDF documents discovered during crawling.

Args:
    url: Direct URL to a PDF file
    max_size_mb: Skip PDFs larger than this (default 5 MB)
    max_chars: Cap extracted text length (default 30000 chars)

Returns:
    Dictionary with extracted text, page count, and character count.
    Returns success=False if the PDF is too large, image-based, or unreadable.
ParametersJSON Schema
NameRequiredDescriptionDefault
urlYes
max_charsNo
max_size_mbNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does well: it discloses that oversized PDFs are skipped, that text length is capped by default, and that failure (success=False) occurs for too-large, image-based, or unreadable files. It stops short of mentioning network/auth requirements, rate limits, or download behavior, leaving some gaps for a network-and-parse operation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The Args/Returns layout is front-loaded and easy to scan, with the core action stated first and supporting detail after. It is slightly padded by the trailing document-type list and a partially redundant Returns block given an output schema exists, but nothing is egregiously wasteful.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a three-parameter download-and-parse tool with an output schema, the description supplies the missing constraint semantics and the failure modes an agent needs. It is nearly complete, with only peripheral details like timeout or retry behavior left out.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate, and it does: all three parameters are explained with intent and defaults (direct PDF URL, size-based skip threshold at 5 MB, character cap at 30000). The url wording only notes it must be a direct PDF link, so subtle edge cases (redirects, non-PDF content types) remain unaddressed.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific verb sequence (download and extract text) and a precise resource (PDF content via pdfplumber), which cleanly separates it from siblings like scrape_url or extract_data. It also enumerates concrete document types (handbooks, calendars, prospectuses) that sharpen the intended target.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives clear usage context — PDFs discovered during crawling, and the kinds of documents it suits — which helps an agent pick it over generic scraping tools. However, it never names an alternative tool or states when not to use it, so the routing guidance is implied rather than explicit.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_db_contextA
Return WHED database reference data for a given institution domain.

Provides two types of grounding to reduce hallucination:
  1. Picklists — valid enum values (institution types, funding, divisions, etc.)
  2. Reference example — a complete record from the same country

Args:
    domain: Institution website domain (e.g. 'www.ampa.edu.au')

Returns:
    Dictionary with picklists, country code, and a reference example
ParametersJSON Schema
NameRequiredDescriptionDefault
domainYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A3.7/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It discloses that it returns reference data intended to reduce hallucination and breaks the return into picklists and a reference example. However, it does not mention error behavior, rate limits, or permission requirements. Since an output schema exists, return format details are less critical, but the description could do more to explain the grounding use case.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is front-loaded with the core purpose, then concisely lists the two types of grounding and provides Args and Returns. Every sentence earns its place, with no waste. The structure is clear and scannable.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given there is an output schema, the description needn't detail return values, and it does mention the return structure. The input parameter is well explained. The main gap is usage guidance relative to siblings, but overall it is complete enough for an agent to invoke it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

There is only one parameter, and the schema description coverage is 0%. The description compensates by explaining that 'domain' is an 'Institution website domain (e.g. \'www.ampa.edu.au\')', which adds format and example meaning beyond the bare schema. This is close to the baseline 4 for zero params, but here it's one param with added semantics, so 4 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a clear verb and resource: 'Return WHED database reference data for a given institution domain.' It further distinguishes the output into picklists and a reference example, which clarifies what kind of data this returns compared to siblings that scrape or extract. It doesn't explicitly differentiate from siblings like get_extraction_schema, but the purpose is specific enough.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies when it's useful ('to reduce hallucination') but does not explicitly state when to use this tool versus alternatives such as get_extraction_schema or validate_profile. There are no exclusions or prerequisites mentioned. The context is implied but not spelled out.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_extraction_schemaA
Return the WHED extraction schema (REQUIRED fields only).

The host LLM should use this template to know which fields to extract
from scraped website content. Each field includes its type and priority.

Typical workflow:
  1. crawl_website / scrape_url  → get site content
  2. get_extraction_schema       → know what to extract
  3. get_db_context(domain)      → get allowed values & reference example
  4. (Host LLM extracts data)
  5. validate_profile(json)      → check the extraction
  6. save_profile(domain, json)  → persist the result
ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.6/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the burden, and it discloses that only required fields are returned and that each field carries type and priority metadata. It is a deterministic, no-argument read tool and an output schema exists, so deeper return-format detail is unnecessary; only operational traits like caching or freshness are left unstated.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The core statement is front-loaded in the first sentence, and the workflow list is compact and high-value for cross-tool routing. Step 2's gloss ('know what to extract') mildly restates the opening sentence, a small redundancy in an otherwise tight definition.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a parameterless tool with a full output schema, the description supplies everything an agent needs: what is returned, the field-level metadata attached, and where it sits in the extraction pipeline. Nothing required to call it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters and the schema description coverage is 100%, so there is nothing for the description to clarify. Baseline of 4 applies for a parameterless tool.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Return the WHED extraction schema') and adds a scoping qualifier ('REQUIRED fields only') that tells the agent what is excluded. Combined with the workflow, an agent can distinguish it from get_db_context (allowed values) and validate_profile without opening any schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The numbered 6-step workflow places this tool precisely in sequence between scraping and get_db_context, and states the condition for its use ('know what to extract'). Explicit alternatives and their ordering are given, so nothing is left to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

save_profileA
Save a validated institution profile to disk as JSON.

Args:
    domain: Institution domain (e.g. 'www.ampa.edu.au'), used as filename
    profile_json: JSON string of the profile to save
    output_dir: Directory to save into (default: output/structured)

Returns:
    Dictionary with save status and file path
ParametersJSON Schema
NameRequiredDescriptionDefault
domainYes
output_dirNooutput/structured
profile_jsonYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A3.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does reasonably well: it discloses the side effect (writes JSON to disk), that the domain determines the filename, and that output_dir defaults to 'output/structured'. It does not state overwrite/merge behavior or error handling for a failed write, which keeps it short of a 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded one-sentence purpose followed by compact Args/Returns blocks. The parameter restatement duplicates the schema somewhat, but given 0% schema coverage it earns its place rather than being waste.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple 3-parameter write tool with an output schema already covering return values, the description supplies the missing parameter semantics and the disk-write side effect. Only edge-case behavior (overwriting an existing file, invalid JSON handling) is absent.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate, and it does: all three parameters are explained, including the example domain format, the JSON-string nature of profile_json, and the default output_dir value. Only the filename derivation detail ('used as filename') is partially specified.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (save) and resource (institution profile) with format (JSON) and destination (disk). The word 'validated' implicitly routes to the sibling validate_profile, giving partial sibling differentiation, though it never names it explicitly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The adjective 'validated' implies the profile must pass validate_profile first, which is useful implied context, but there is no explicit when-to-use statement, no failure path guidance, and no mention of what happens if an unvalidated profile is passed.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

scrape_urlA
Scrape a webpage and return its HTML content.

Args:
    url: The webpage URL to scrape
    javascript: Set to True for JavaScript-rendered sites (slower but handles dynamic content)
    wait_seconds: How long to wait for JavaScript to load (only used when javascript=True)
    save_path: Optional file path to save the result as JSON (e.g. "C:/Users/me/Desktop/result.json").
               If a directory is given, a timestamped filename is generated automatically.

Returns:
    Dictionary with html content, status code, and load time
ParametersJSON Schema
NameRequiredDescriptionDefault
urlYes
save_pathNo
javascriptNo
wait_secondsNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden and does reasonably well: it discloses the slowness tradeoff of javascript=True, the conditional coupling of wait_seconds, and the file-writing side effect of save_path including directory-to-timestamped-filename behavior. Gaps remain around failure modes, blocking/rate limits, and what happens on non-HTML responses, which keeps it below 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The purpose is front-loaded in a single sentence, followed by tightly scoped Args and Returns sections. Every line adds information about a parameter's behavior; there is no filler or restatement.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a four-parameter tool with an output schema, the description covers the input contract thoroughly and need not explain return values. The only meaningful omission is guidance about alternatives among its many sibling tools, which is a routing concern rather than an invocation one.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate, and it does: each of the four parameters is documented with meaning beyond its name, including conditional semantics ('only used when javascript=True') and non-obvious save_path behavior when a directory is passed. This is exactly what a low-coverage schema requires.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb+resource: 'Scrape a webpage and return its HTML content.' That is unambiguous about the action and the payload. It does not, however, differentiate from siblings like batch_scrape, crawl_website, or extract_data, so a 4 rather than a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is implied through the javascript flag ('Set to True for JavaScript-rendered sites'), which is useful conditional guidance. There is no explicit statement of when to prefer this tool over batch_scrape or crawl_website, no exclusions, and no prerequisites. Adequate but not routing guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

validate_profileA
Validate an extracted institution profile against the Pydantic schema
and WHED database picklists.

Args:
    profile_json: JSON string of the extracted profile
                  (must match the structure from get_extraction_schema)

Returns:
    Dictionary with validation status, cleaned data, and any warnings
ParametersJSON Schema
NameRequiredDescriptionDefault
profile_jsonYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A3.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden. It discloses the return shape (validation status, cleaned data, warnings) and that the input must be a JSON string, but says nothing about failure behavior, whether the profile is mutated, or any permission requirements.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the purpose in the first sentence, then uses a tidy Args/Returns layout. Slightly verbose with the multi-line arg wrapping, but every part conveys information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values needn't be spelled out, yet the description still summarizes them. For a single-parameter validation tool the purpose, input source, and outcome are all covered, leaving only workflow placement and error semantics unaddressed.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, so the description must compensate, and it does: it explains that profile_json is a JSON string of the extracted profile and that it must match the structure returned by get_extraction_schema. That is meaningfully more than the bare 'Profile Json' title in the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (validate) and resource (extracted institution profile) and names exactly what it validates against (Pydantic schema and WHED database picklists). This clearly separates it from siblings like save_profile and get_extraction_schema, though it stops short of explicitly contrasting them.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is implied: it references get_extraction_schema as the source of the required structure, which hints at the workflow order (extract, then validate). However, there is no explicit statement of when to call this versus save_profile, nor any when-not guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 10 tool updatesv0.2.0
    • First observedbatch_scrape
    • First observedcrawl_website
    • First observedextract_data
    • First observedextract_first
    • First observedextract_pdf_text
    • First observedget_db_context
    • First observedget_extraction_schema
    • First observedsave_profile
    • First observedscrape_url
    • First observedvalidate_profile

TDQS

A3.9/5.0

Scored across 10 tools

Disambiguation4/5

Most tools target distinct stages or actions: raw scraping, multi-selector extraction, single-element extraction, batch scraping, and crawling. Some overlap remains among the scraping tools, but descriptions clarify boundaries well enough for an agent to choose correctly.

Naming Consistency5/5

All tool names use snake_case and are action-oriented. The verb_noun pattern is consistent throughout, with only minor modifiers like batch_scrape and extract_first remaining readable and predictable.

Tool Count5/5

Ten tools is well-scoped for a scraping, extraction, validation, and persistence pipeline. Each tool has a clear role and the set avoids unnecessary redundancy.

Completeness4/5

The surface covers content acquisition via scraping, crawling, and PDF extraction, plus schema guidance, database context, validation, and saving. It lacks explicit retrieval or update operations for saved profiles, but the core WHED extraction workflow is otherwise complete.

Maintenance

ActivityInactive
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • A
    license
    A
    quality
    D
    maintenance
    Structured web context infrastructure for AI agents. Extract reliable schema-guided JSON from websites using Claude-powered parsing, Browserless fallback rendering, and MCP-native workflows.
    1
    1
    MIT
  • A
    license
    Not graded
    quality
    C
    maintenance
    A general-purpose MCP server for crawling and extracting structured data from any website. Supports tools for crawling, single-page extraction, search-and-crawl, and schema extraction.
    2
    MIT
  • A
    license
    A
    quality
    B
    maintenance
    An agent-agnostic web extraction and fetch layer that turns URLs into verified, typed data with confidence scores via MCP, REST, or SDK, orchestrating scraping engines behind a resilience ladder and supporting structured extraction against any schema.
    5
    3
    Apache 2.0