Skip to main content
Glama
paulet4a-commits

WebDataTools Domain & website intelligence MCP server

wayback_page_diff

Compare the oldest and newest Internet Archive snapshots of a URL to see visible text changes, or list all captures. Returns one scored change row per URL.

Instructions

Wayback Machine Snapshot & Page Change Tracker lists every Internet Archive capture of a URL and diffs the visible text of the oldest vs. newest snapshot in range — one scored change row per URL, or a full snapshot list. Billed to your own Apify account: ~$0.003 per result (Apify free-plan price, lower on paid plans).

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
toNoTo date — Only consider captures on or before this date, e.g. 2024-06-30 or 20240630. Leave empty for no upper bound.
fromNoFrom date — Only consider captures on or after this date, e.g. 2023-01-01 or 20230101. Leave empty for no lower bound.
modeNoMode — "Diff" compares the oldest and newest archived snapshot in range and returns what changed. "Snapshots" lists every capture Wayback has for the URL, newest first, with no comparison. Options: diff = Diff (compare oldest vs. newest capture); snapshots = Snapshots (list all captures).diff
urlsYesURLs — Enter the page URLs to look up in the Wayback Machine, e.g. https://apify.com/pricing. Bare domains are accepted too. One row (diff mode) or up to Max snapshots per URL rows (snapshots mode) comes back per URL. Example: ["https://apify.com/pricing"].
maxSnapshotsNoMax snapshots per URL — Snapshots mode only: the most captures returned per URL (newest first). Ignored in diff mode, which always emits exactly one row per URL.

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observedv0.1.0

TDQS

A3.9/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does meaningfully discharge it: it discloses that the run is billed to the caller's own Apify account at roughly $0.003 per result, and that diff mode always emits exactly one row per URL while snapshots mode emits many. It does not cover rate limits, failure handling for URLs with no captures, or what the change score represents.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The purpose is front-loaded in the first clause, and the second sentence carries only the cost disclosure, which is genuinely decision-relevant. The first sentence is long and packs three ideas (listing, diffing, output cardinality), slightly hurting scanability, but there is no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 5-parameter, no-output-schema tool, the description covers what comes back per mode and the cost model, which is the main missing structured information. Remaining gaps are minor: sorting of diff rows and the meaning/range of the change score are unstated, and there is no note on behavior when Wayback has no captures.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3; the schema already documents from/to, mode, urls, and maxSnapshots in detail including format examples and enum meanings. The description adds no parameter-level detail beyond the schema, which is acceptable but earns no credit.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource set: it lists Internet Archive captures of a URL and diffs the visible text of the oldest vs. newest snapshot in range. The dual output shape ('one scored change row per URL, or a full snapshot list') makes the tool's two behaviors unambiguous, and no sibling tool overlaps with Wayback/archival work.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage context is implied rather than stated: the mention of a 'scored change row' versus a 'full snapshot list' hints at tracking changes vs. auditing history, but the description never says when to pick this tool, nor does it name any alternative or exclusion criteria. With no close siblings, the omission is low-risk but still a gap.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.