Archive Evidence
archive_evidenceWhat did a web page say on a given date? Finds the archived capture of a URL nearest that date across the Internet Archive Wayback Machine, Common Crawl and Arquivo.pt, reads it, and returns the readable text (HTML stripped) with the capture timestamp, a link to the archived copy, and the GAP between the date asked for and the capture actually used (e.g. "Captured 1.5 day(s) before the requested date"). One call replaces the find-capture, fetch-record, parse-HTML chain. If no archive holds a capture within max_gap_days (default 30) it answers status "no_capture" and names the nearest capture before and after the date with their gaps, so a distant capture is never passed off as the page on that date. Every archive's outcome is listed in sources[]; one that timed out or errored says so with its upstream status, and if all of them failed the answer is error "archives_unavailable", not no_capture. Captures are immutable and cached by capture id; served_from says whether the text came from cache or upstream. Use archive_compare to see what changed between two dates.
Input Schema
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | The page address, e.g. "https://www.sec.gov" or "https://www.ftc.gov/policy-notices/privacy-policy". | |
| date | Yes | The date to look up, as YYYY-MM-DD (e.g. "2023-06-01"), YYYYMMDD, or a full ISO timestamp. | |
| max_chars | No | Characters of page text to return. Default 12000, max 60000. | |
| max_gap_days | No | Furthest a capture may be from the date and still be used, in days. Default 30, max 3650. |