Skip to main content
Glama
Dasistaiden

whed-tools

by Dasistaiden

Server Configuration

Describes the environment variables required to run the server.

NameRequiredDescriptionDefault

No arguments

Instructions

Guidance the server publishes about itself, which clients place ahead of the tool catalog so the model reads it before choosing anything.

This server publishes no instructions, or was last inspected before Glama recorded them.

Capabilities

Features and capabilities supported by this server

Protocol revision2025-11-25

CapabilityDetails
tools
{
  "listChanged": false
}
prompts
{
  "listChanged": false
}
resources
{
  "subscribe": false,
  "listChanged": false
}
experimental
{}

Tools

Functions exposed to the LLM to take actions

NameDescription
scrape_urlA
Scrape a webpage and return its HTML content.

Args:
    url: The webpage URL to scrape
    javascript: Set to True for JavaScript-rendered sites (slower but handles dynamic content)
    wait_seconds: How long to wait for JavaScript to load (only used when javascript=True)
    save_path: Optional file path to save the result as JSON (e.g. "C:/Users/me/Desktop/result.json").
               If a directory is given, a timestamped filename is generated automatically.

Returns:
    Dictionary with html content, status code, and load time
extract_dataA
Scrape a webpage and extract specific data using CSS selectors.

Args:
    url: The webpage to scrape
    css_selectors: List of CSS selectors (e.g., ["h1", "a.link", "#content"])
    attributes: List of attributes to extract for each selector (e.g., ["text", "href", "text"])
               If not provided, defaults to "text" for all selectors
    javascript: Set to True for JavaScript-rendered sites
    save_path: Optional file path to save the result as JSON (e.g. "C:/Users/me/Desktop/result.json").
               If a directory is given, a timestamped filename is generated automatically.

Returns:
    Dictionary with extracted data for each selector

Example:
    extract_data(
        url="https://example.com",
        css_selectors=["h1", "a"],
        attributes=["text", "href"]
    )
extract_firstA
Extract the first matching element from a webpage.
Useful for getting single values like page title, main heading, etc.

Args:
    url: The webpage to scrape
    css_selector: CSS selector for the element (e.g., "h1", "title", "meta[name='description']")
    attribute: What to extract - "text" for content, or attribute name like "href", "content", "src"
    javascript: Set to True for JavaScript-rendered sites
    save_path: Optional file path to save the result as JSON (e.g. "C:/Users/me/Desktop/result.json").
               If a directory is given, a timestamped filename is generated automatically.

Returns:
    Dictionary with the extracted value

Example:
    extract_first(url="https://example.com", css_selector="title", attribute="text")
batch_scrapeA
Scrape multiple URLs efficiently.

Args:
    urls: List of URLs to scrape
    javascript: Set to True if the sites need JavaScript rendering
    save_path: Optional file path to save all results as a JSON array (e.g. "C:/Users/me/Desktop/batch.json").
               If a directory is given, a timestamped filename is generated automatically.

Returns:
    List of scraping results for each URL
crawl_websiteA
Crawl a website to discover its structure and pages.

Args:
    start_url: Starting URL
    max_pages: Maximum pages to crawl (default 50)
    max_depth: Maximum link depth (default 3)
    same_domain_only: Stay on same domain (default True)
    schema_filter: When True, only follow URLs whose path contains keywords
                   relevant to the WHED schema (about, contact, course, program,
                   faculty, etc.). Skips login pages, media, and unrelated content.
                   Recommended for institution websites.
    save_path: Optional file path to save the site map as JSON (e.g. "C:/Users/me/Desktop/sitemap.json").
               If a directory is given, a timestamped filename is generated automatically.

Returns:
    Site map with discovered pages and statistics
extract_pdf_textA
Download a PDF and extract its text content using pdfplumber.
Useful for reading course handbooks, academic calendars, prospectuses,
and other PDF documents discovered during crawling.

Args:
    url: Direct URL to a PDF file
    max_size_mb: Skip PDFs larger than this (default 5 MB)
    max_chars: Cap extracted text length (default 30000 chars)

Returns:
    Dictionary with extracted text, page count, and character count.
    Returns success=False if the PDF is too large, image-based, or unreadable.
get_extraction_schemaA
Return the WHED extraction schema (REQUIRED fields only).

The host LLM should use this template to know which fields to extract
from scraped website content. Each field includes its type and priority.

Typical workflow:
  1. crawl_website / scrape_url  → get site content
  2. get_extraction_schema       → know what to extract
  3. get_db_context(domain)      → get allowed values & reference example
  4. (Host LLM extracts data)
  5. validate_profile(json)      → check the extraction
  6. save_profile(domain, json)  → persist the result
get_db_contextA
Return WHED database reference data for a given institution domain.

Provides two types of grounding to reduce hallucination:
  1. Picklists — valid enum values (institution types, funding, divisions, etc.)
  2. Reference example — a complete record from the same country

Args:
    domain: Institution website domain (e.g. 'www.ampa.edu.au')

Returns:
    Dictionary with picklists, country code, and a reference example
validate_profileA
Validate an extracted institution profile against the Pydantic schema
and WHED database picklists.

Args:
    profile_json: JSON string of the extracted profile
                  (must match the structure from get_extraction_schema)

Returns:
    Dictionary with validation status, cleaned data, and any warnings
save_profileA
Save a validated institution profile to disk as JSON.

Args:
    domain: Institution domain (e.g. 'www.ampa.edu.au'), used as filename
    profile_json: JSON string of the profile to save
    output_dir: Directory to save into (default: output/structured)

Returns:
    Dictionary with save status and file path

Prompts

Interactive templates invoked by user choice

NameDescription

No prompts

Resources

Contextual data attached and managed by the client

NameDescription
get_helpGet help documentation for the web scraping tools

TDQS

A3.9/5.0

Scored across 10 tools

Disambiguation4/5

Most tools target distinct stages or actions: raw scraping, multi-selector extraction, single-element extraction, batch scraping, and crawling. Some overlap remains among the scraping tools, but descriptions clarify boundaries well enough for an agent to choose correctly.

Naming Consistency5/5

All tool names use snake_case and are action-oriented. The verb_noun pattern is consistent throughout, with only minor modifiers like batch_scrape and extract_first remaining readable and predictable.

Tool Count5/5

Ten tools is well-scoped for a scraping, extraction, validation, and persistence pipeline. Each tool has a clear role and the set avoids unnecessary redundancy.

Completeness4/5

The surface covers content acquisition via scraping, crawling, and PDF extraction, plus schema guidance, database context, validation, and saving. It lacks explicit retrieval or update operations for saved profiles, but the core WHED extraction workflow is otherwise complete.

Maintenance

ActivityInactive
ResponsivenessNo issues