arquivo-pt-mcp
Server Configuration
Describes the environment variables required to run the server.
| Name | Required | Description | Default |
|---|---|---|---|
| HOST | No | Host to bind to (used with http transport) | 127.0.0.1 |
| PORT | No | Port to bind to (used with http transport) | 8000 |
| TRANSPORT | No | Transport mode: 'stdio' or 'http' | stdio |
| ALLOWED_HOST | No | Allowed host for CORS (used with http transport when exposed publicly) |
Instructions
Guidance the server publishes about itself, which clients place ahead of the tool catalog so the model reads it before choosing anything.
This server publishes no instructions, or was last inspected before Glama recorded them.
Capabilities
Features and capabilities supported by this server
Protocol revision2025-11-25
| Capability | Details |
|---|---|
| tools | {
"listChanged": false
} |
| experimental | {} |
Tools
Functions exposed to the LLM to take actions
| Name | Description |
|---|---|
| searchA | Full-text search across the Portuguese Web Archive (Arquivo.pt). Use for finding pages that ever contained given terms, optionally scoped by date range or site. |
| image_searchB | Search 1.8B+ archived images on Arquivo.pt (Dionisius image search). Find historical photos, logos, and graphics from the Portuguese web. |
| list_versionsA | List every archived capture of a specific URL (CDX query). Use to see how a page changed over time. |
| get_snapshotA | Get the archive URL for a specific snapshot of a page. Omit timestamp for the latest capture. |
| extract_textA | Fetch an archived snapshot and return its readable text content (HTML stripped). |
| get_screenshotA | Get the Arquivo.pt PNG render of an archived page. By default returns the screenshot URL; pass inline=true to embed the PNG (useful for letting the model see the page). Omit timestamp for the latest capture. |
Prompts
Interactive templates invoked by user choice
| Name | Description |
|---|---|
No prompts | |
Resources
Contextual data attached and managed by the client
| Name | Description |
|---|---|
No resources | |
TDQS
Scored across 6 tools
Each tool targets a clearly distinct function: text search, image search, version history, snapshot URL retrieval, text extraction, and screenshot rendering. Even the snapshot-related tools are separable by their output (URL vs. text vs. PNG). An agent can reliably choose the right tool based on the desired result.
Most tool names follow a verb_noun pattern in snake_case: list_versions, get_snapshot, extract_text, get_screenshot. The exception is image_search, which inverts the pattern to noun_verb, and bare search is less descriptive, but the overall style remains predictable and readable.
Six tools is well-scoped for a web archive server. Each tool serves a distinct retrieval or discovery need without redundancy or bloat. This is an appropriate size for the domain.
The server covers the core archive workflow: finding pages and images, viewing capture history, retrieving snapshot URLs, and extracting text or screenshots. Minor gaps exist around snapshot metadata or downloadable raw content, but agents can still accomplish most archive research tasks.