| scrape_urlA | Scrape a webpage and return its HTML content.
Args:
url: The webpage URL to scrape
javascript: Set to True for JavaScript-rendered sites (slower but handles dynamic content)
wait_seconds: How long to wait for JavaScript to load (only used when javascript=True)
save_path: Optional file path to save the result as JSON (e.g. "C:/Users/me/Desktop/result.json").
If a directory is given, a timestamped filename is generated automatically.
Returns:
Dictionary with html content, status code, and load time
|
| extract_dataA | Scrape a webpage and extract specific data using CSS selectors.
Args:
url: The webpage to scrape
css_selectors: List of CSS selectors (e.g., ["h1", "a.link", "#content"])
attributes: List of attributes to extract for each selector (e.g., ["text", "href", "text"])
If not provided, defaults to "text" for all selectors
javascript: Set to True for JavaScript-rendered sites
save_path: Optional file path to save the result as JSON (e.g. "C:/Users/me/Desktop/result.json").
If a directory is given, a timestamped filename is generated automatically.
Returns:
Dictionary with extracted data for each selector
Example:
extract_data(
url="https://example.com",
css_selectors=["h1", "a"],
attributes=["text", "href"]
)
|
| extract_firstA | Extract the first matching element from a webpage.
Useful for getting single values like page title, main heading, etc.
Args:
url: The webpage to scrape
css_selector: CSS selector for the element (e.g., "h1", "title", "meta[name='description']")
attribute: What to extract - "text" for content, or attribute name like "href", "content", "src"
javascript: Set to True for JavaScript-rendered sites
save_path: Optional file path to save the result as JSON (e.g. "C:/Users/me/Desktop/result.json").
If a directory is given, a timestamped filename is generated automatically.
Returns:
Dictionary with the extracted value
Example:
extract_first(url="https://example.com", css_selector="title", attribute="text")
|
| batch_scrapeA | Scrape multiple URLs efficiently.
Args:
urls: List of URLs to scrape
javascript: Set to True if the sites need JavaScript rendering
save_path: Optional file path to save all results as a JSON array (e.g. "C:/Users/me/Desktop/batch.json").
If a directory is given, a timestamped filename is generated automatically.
Returns:
List of scraping results for each URL
|
| crawl_websiteA | Crawl a website to discover its structure and pages.
Args:
start_url: Starting URL
max_pages: Maximum pages to crawl (default 50)
max_depth: Maximum link depth (default 3)
same_domain_only: Stay on same domain (default True)
schema_filter: When True, only follow URLs whose path contains keywords
relevant to the WHED schema (about, contact, course, program,
faculty, etc.). Skips login pages, media, and unrelated content.
Recommended for institution websites.
save_path: Optional file path to save the site map as JSON (e.g. "C:/Users/me/Desktop/sitemap.json").
If a directory is given, a timestamped filename is generated automatically.
Returns:
Site map with discovered pages and statistics
|
| extract_pdf_textA | Download a PDF and extract its text content using pdfplumber.
Useful for reading course handbooks, academic calendars, prospectuses,
and other PDF documents discovered during crawling.
Args:
url: Direct URL to a PDF file
max_size_mb: Skip PDFs larger than this (default 5 MB)
max_chars: Cap extracted text length (default 30000 chars)
Returns:
Dictionary with extracted text, page count, and character count.
Returns success=False if the PDF is too large, image-based, or unreadable.
|
| get_extraction_schemaA | Return the WHED extraction schema (REQUIRED fields only).
The host LLM should use this template to know which fields to extract
from scraped website content. Each field includes its type and priority.
Typical workflow:
1. crawl_website / scrape_url → get site content
2. get_extraction_schema → know what to extract
3. get_db_context(domain) → get allowed values & reference example
4. (Host LLM extracts data)
5. validate_profile(json) → check the extraction
6. save_profile(domain, json) → persist the result
|
| get_db_contextA | Return WHED database reference data for a given institution domain.
Provides two types of grounding to reduce hallucination:
1. Picklists — valid enum values (institution types, funding, divisions, etc.)
2. Reference example — a complete record from the same country
Args:
domain: Institution website domain (e.g. 'www.ampa.edu.au')
Returns:
Dictionary with picklists, country code, and a reference example
|
| validate_profileA | Validate an extracted institution profile against the Pydantic schema
and WHED database picklists.
Args:
profile_json: JSON string of the extracted profile
(must match the structure from get_extraction_schema)
Returns:
Dictionary with validation status, cleaned data, and any warnings
|
| save_profileA | Save a validated institution profile to disk as JSON.
Args:
domain: Institution domain (e.g. 'www.ampa.edu.au'), used as filename
profile_json: JSON string of the profile to save
output_dir: Directory to save into (default: output/structured)
Returns:
Dictionary with save status and file path
|