Skip to main content
Glama
harrywesterman

geldersarchief-mcp

Gelders Archief scan core

Lokale MCP-server, Python-core en CLI voor het zoeken van Gelders Archief-aktes, het ontdekken van alle registerscans en downloaden met provenance. De MCP-server draait via stdio met de officiële SDK. OCR is optioneel en standaard uitgeschakeld in MCP; zoeken, inspecteren en expliciet downloaden werken zonder OCR.

uv sync --frozen
uv run --frozen geldersarchief-mcp

Verbind een MCP-client met de clientconfiguratie. Zie MCP-tools en gebruik voor de volledige interface en downloadflow.

De voorbeeldpermalink is live getest op 5 oktober 2026: 192 unieke scans, met volgorde 1–192. De viewer laadt aanvankelijk 25 referenties. De core ontdekt alle scans rechtstreeks via HTTP, zonder Chromium te starten. Scan 176 is via de expliciete viewer-downloadlink opgeslagen als JPEG, 3019 × 4170 pixels, 1.372.341 bytes. SHA256:

c9488c8abf3aa2c97921c7f04e080cc2299a52af79156f606640f36b04091611

Dit is de door de viewer aangeboden downloadrepresentatie. Het archiefmaster- bestand is niet onafhankelijk beschikbaar gesteld of op byte-identiteit getest. Scan 176 is een technisch testobject; er is niet vastgesteld dat deze scan de akte uit de voorbeeldpermalink bevat.

Installatie

Vereist: Python 3.12+ en uv.

uv sync --frozen
# Alleen nodig voor vieweronderzoek en browserfallbacks:
uv run playwright install chromium

Related MCP server: Nationaal Archief MCP Server

CLI

uv run geldersarchief inspect \
  'https://permalink.geldersarchief.nl/E42035012A364A3E9190403665F46E77'

uv run geldersarchief download-scan \
  'https://permalink.geldersarchief.nl/E42035012A364A3E9190403665F46E77' \
  --sequence 176

JSON-resultaten staan op stdout; geredigeerde JSON-logs op stderr. Gebruik geldersarchief --debug inspect URL voor batchdiagnostiek. Downloads komen standaard in data/downloads/; --output DIRECTORY overschrijft dat per opdracht.

Bestanden heten <register>_scan-<volgnummer>.<ext>. De sidecar bevat bronpermalink, register, inventaris, scan-ID, oorspronkelijke download-URL, UTC-downloadtijd, SHA256, dimensies, discoverymethode en representatietype. Aktenummer en confidence blijven bij een expliciet gekozen CLI-scan null. MCP-downloads bewaren daarnaast de resolutiestatus en het eventuele bewijs van een aktekoppeling. Afbeeldingen worden zonder hercompressie opgeslagen.

Vieweronderzoek

uv run python scripts/inspect_viewer.py --load-all --output debug/research

Dit legt consolemeldingen, requests, redirects, relevante JSON/JavaScript, DOM, scriptbronnen, window-keys en stores vast. --load-all laadt de batches van het geselecteerde register opeenvolgend met 350 ms vertraging. --headed opent een zichtbaar venster. Een expliciete viewer-URL kan als positioneel argument worden meegegeven om ook de iframe-/tileviewer te onderzoeken.

Op de ontwikkelmachine sluit het lokale netwerkverkeer van Chromium verbindingen met deze site af, terwijl HTTP met OS-certificaatvertrouwen werkt. Hiervoor is er een expliciete transportoptie:

uv run python scripts/inspect_viewer.py --http-transport --load-all
GA_BROWSER_HTTP_TRANSPORT=true uv run pytest -m live

Hierbij voert Chromium nog steeds de viewer-JavaScript uit en traceert Playwright requests; HTTPX verzorgt het transport met gecontroleerde TLS via de OS-truststore. TLS-verificatie wordt niet uitgeschakeld. Deze optie is geen site-authenticatie- of CAPTCHA-omzeiling. Normale CLI-discovery heeft deze optie niet nodig.

Python-interface

import asyncio
from geldersarchief_mcp.archives.geldersarchief import GeldersArchiefAdapter

async def main():
    async with GeldersArchiefAdapter() as archive:
        register = await archive.inspect_register(
            'https://permalink.geldersarchief.nl/E42035012A364A3E9190403665F46E77'
        )
        image = await archive.get_scan(register.scans[175])
        # image bevat de originele bytes van de viewer-downloadrepresentatie.

asyncio.run(main())

inspect_register geeft alleen een volledig register terug. Ontbrekende batches, conflicterende posities of een onbekend totaal geven een expliciete exception. download_scan geeft naast bytes ook de gevalideerde URL, MIME en dimensies terug. Een verlopen URL vereist verse registerdiscovery; langdurige tokencaching is nog niet geïmplementeerd.

Configuratie

Variabele

Default

GA_DATA_DIR

data

GA_DOWNLOAD_DIR

<GA_DATA_DIR>/downloads

GA_BROWSER_HEADLESS

true

GA_BROWSER_HTTP_TRANSPORT

false

GA_MAX_CONCURRENT_REQUESTS

2

GA_MAX_IMAGE_DOWNLOADS

2

Eén adapter beheert één herbruikbare browser, maximaal twee contexts, maximaal twee registeroperaties en twee downloads met de defaults. Bulk-batches zijn sequentieel met 350 ms vertraging. HTTP 429/502/503/504 krijgt beperkte retries met back-off en Retry-After. Alleen openbare bronnen worden gebruikt; geen externe AI-services. De resolver gebruikt uitsluitend lokale OCR en een lokale observatiecache.

Debugoutput en lokale downloads zijn uitgesloten via .gitignore. Debugoutput redigeert sessies en toegangstokens, ook in base64-HTML en ge-escapete JSON. Lokale provenance bevat de volledige publieke download-URL voor reproduceerbaarheid; publiceer die sidecars niet zonder de URL-tokens te redigeren.

Tests

uv run pytest            # offline; live tests uitgesloten
uv run pytest -m live    # bewuste publieke netwerkcontrole

Offline tests dekken parsing, volgorde, ontdubbeling, storeselectie, ontbrekende batches, downloadvalidatie, provenance, concurrency, redactie en back-off. Live tests herhalen discovery/download met verse HTTP-clients en browserinstances. CI draait alleen offline tests, op Python 3.12 en 3.14.

De bewezen HTTP-route is primair. Browserstore-, DOM- en navigatiefallbacks zijn voor toekomstige wijzigingen; hun algemene werking buiten het geteste register is niet aangetoond. Zie het onderzoeksrapport voor waarnemingen, beperkingen en de tien onderzoeksvragen.

Docker is nog niet geïmplementeerd. De observatiecache is geen volledige registerindex.

MIT-licentie. Het project is niet verbonden aan het Gelders Archief of DE REE.

Open Archieven-integratie

Zoeken en het normaliseren van records zijn geïmplementeerd via de officiële Open Archieven 1.1 API. Zoekresultaten worden verrijkt met de volledige A2A-records, zodat aktenummer, inventarisnummer en archiefbron beschikbaar zijn.

uv run geldersarchief search-acts --name 'Derk Jan van Brink' \
  --place Terwolde --year 1921 --record-type geboorte

uv run geldersarchief get-record \
  'gld:E4203501-2A36-4A3E-9190-403665F46E77'

uv run geldersarchief inspect \
  --record 'gld:E4203501-2A36-4A3E-9190-403665F46E77'

download-scan accepteert eveneens --record in plaats van een URL, maar vereist nog steeds een expliciet --sequence. Dat is geen automatische akte→scan-match. De sidecar bewaart het aangevraagde record apart van de ongeverifieerde aktevelden.

Typen: birth/geboorte, death/overlijden, marriage/huwelijk. Een --date filtert op de gebeurtenisdatum. --limit (1–100) en --start bedienen paginering; number_found is het aantal persoonstreffers van de API, vóór lokale filtering en ontdubbeling op record-ID. next_start verwijst naar de volgende API-pagina.

De bekende geboorteakte heeft event_date=1921-12-30, act_date=1922-01-02 en act_number="1". De resolver gebruikt de aktedatum voor nummerpositionering. Onvolledige datums worden niet aangevuld met verzonnen maanden of dagen.

De client begrenst opeenvolgende calls tot maximaal één per 350 ms. Ontbrekende records, HTTP-fouten en API-foutcodes worden expliciet afgehandeld. Er is geen API-key of AI-provider nodig. Configureer GA_USER_AGENT met je project-URL of contactadres bij publiek gebruik, zoals Open Archieven vraagt.

Details en bronnen: Open Archieven-integratie.

Conservatieve akte→scan-resolver

Installeer voor resolve lokaal Tesseract en Nederlandse taaldata, bijvoorbeeld op macOS met brew install tesseract tesseract-lang.

uv run geldersarchief resolve \
  --record 'gld:E4203501-2A36-4A3E-9190-403665F46E77' --max-scans 12

uv run geldersarchief resolve 'https://permalink.geldersarchief.nl/E42035012A364A3E9190403665F46E77' \
  --act 1 --date 1922-01-02 --name 'Derk Jan van Brink'

Bij resolve betekent --date de aktedatum. De zoekstrategie controleert expliciete scanmetadata, bemonstert het register en gebruikt interpolatie alleen bij voldoende monotone nummerankers. Het scanbudget voorkomt een onbeperkte OCR-doorloop. Kandidaten en buurscans zijn verwijzingen, geen automatische download van een als juist bevestigde akte.

Live proef op het voorbeeldregister: na 12 bekeken scans blijft het resultaat AMBIGUOUS. Tesseract las losse nummers, maar bevestigde geen bijbehorende datum of naam. Vervolgonderzoek heeft scan 3 visueel bevestigd als akte 1. De optionele lokale Kraken-reader herkent de aktedatum en stelt scan 3 voor als kandidaat (AMBIGUOUS), maar bevestigt de akte nog niet automatisch. Zie lokale HTR voor installatie en downloadbewijs. Algorithmische tests met gecontroleerde observaties bewijzen de zoeklogica, niet de leeskwaliteit van deze handgeschreven bron. Zie resolver en live beperkingen.

Variabele

Default

GA_CACHE_DB

<GA_DATA_DIR>/cache.sqlite

GA_RESOLVER_MAX_SCANS

24

GA_OCR_EXECUTABLE

tesseract

GA_OCR_LANGUAGE

nld

De lokale SQLite-cache bewaart OCR-tekst met afbeeldings-SHA256 en readeridentiteit. Verwijder de cache na het vervangen van taalmodelbestanden; hun inhoud is geen onderdeel van de readeridentiteit. Cache en scans blijven buiten Git.

Available Tools

6 tools
download_actB

Resolve and download an act; original means unchanged viewer-download bytes.

FOUND downloads its bound scans. Uncertain results return files=[]. A supplied scan_sequence explicitly selects a scan (status SELECTED). allow_probable permits a PROBABLE download while preserving that status. No archive-master claim or silent first-record/candidate selection is made.

ParametersJSON Schema
NameRequiredDescriptionDefault
readerNodisabled
max_scansNo
record_idYes
output_formatNooriginal
scan_sequenceNo
allow_probableNo

TDQS

B3.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only give the safety profile (readOnlyHint=false, destructiveHint=false, openWorldHint=true). The description adds real behavioral content beyond that: 'original means unchanged viewer-download bytes', uncertain results return files=[], a supplied scan_sequence yields SELECTED status, allow_probable permits a PROBABLE download while preserving that status, and no silent first-record/candidate selection is made. This is unusually rich status/semantics disclosure.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The purpose is front-loaded and the text is short, so size is appropriate. However, the terse, jargon-dense fragments ('FOUND downloads its bound scans', 'archive-master claim') are cryptic and require domain knowledge to parse, which hurts structure more than length does.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 6-parameter mutation-capable tool with no output schema, the description usefully covers the uncertain/failed result shape (files=[]) and status semantics. It still leaves several parameters and the reader/output pipeline unexplained, so an agent cannot fully predict invocation behavior.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must carry parameter meaning. It does explain scan_sequence, allow_probable, and output_format ('original means unchanged viewer-download bytes'), but leaves reader (an enum), max_scans, and record_id completely unaddressed. Partial compensation for a 6-param, zero-coverage schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The first sentence gives a specific verb+resource ('Resolve and download an act') and the surrounding sentences describe what resolution entails. It does not explicitly distinguish itself from the sibling 'download_scan', which would be the most likely point of confusion for an agent.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It hints at behavior ('Uncertain results return files=[]') and explains what scan_sequence and allow_probable do, but never states when to use this tool versus download_scan or resolve_act_scan. There are no explicit when/when-not conditions or named alternatives.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

download_scanA

Download an explicitly selected one-based scan with SHA256 and local provenance.

Provide exactly one record_id/source_url. Selection does not verify act identity. Files are exposed through resource_uri/provenance_uri, not inline base64 JSON.

ParametersJSON Schema
NameRequiredDescriptionDefault
sequenceYes
record_idNo
source_urlNo

TDQS

A3.6/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations cover the safety profile (readOnlyHint=false, openWorldHint=true, destructiveHint=false), and the description adds value beyond them: the caveat that selection does not verify act identity, and the critical detail that files come back as resource_uri/provenance_uri rather than inline base64. That return mechanism is non-obvious and materially changes how an agent consumes the result.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three tight sentences, front-loaded with the core action and constraint, with no filler. The telegraphic style is efficient but borders on terse, leaving a little interpretation to the reader.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description usefully explains the resource_uri/provenance_uri return path, and it establishes the selection constraint. But the required sequence parameter is never really explained, and with openWorldHint=true there is no note on access or permission prerequisites, leaving gaps for a network-fetching tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description carries the burden. It explains the mutual exclusivity of record_id/source_url ('exactly one'), and 'one-based' loosely clarifies the required sequence parameter. Neither parameter's format is fully pinned down, so compensation is only partial.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (download) and resource (scan) and adds meaningful qualifiers: one-based selection, SHA256 verification, and local provenance. However, it does not distinguish this scan-download tool from siblings like download_act or resolve_act_scan, so an agent must infer the split.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives a hard constraint ('Provide exactly one record_id/source_url') that shapes invocation, and 'explicitly selected' hints a prior resolution step. But it names no alternatives and offers no when-to-use/when-not guidance relative to download_act or resolve_act_scan.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_recordC
Read-only

Get a gld: record with act number, act/event dates, source and relatives.

ParametersJSON Schema
NameRequiredDescriptionDefault
record_idYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, destructiveHint=false and openWorldHint=true, so safety is covered externally. The description's only added content is a field list, which duplicates the output schema rather than disclosing anything new (auth needs, not-found behavior, rate limits).

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single compact sentence, front-loaded with the action and resource. No filler, though the trailing field enumeration is partly redundant with the output schema.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be re-explained, and annotations cover the safety profile. Still, the description is thin on the surrounding context that matters for a lookup tool, such as what happens when the ID does not resolve.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% and record_id has no documentation in the schema, so the description is the only source of parameter meaning. It partially compensates by revealing the expected 'gld:<UUID>' identifier format, but says nothing about validity, casing, or failure modes.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (Get) and resource (a gld:<UUID> record), and enumerates the payload the caller gets back (act number, act/event dates, source, relatives). It is reasonably distinguishable from search_acts, though the cryptic 'gld:<UUID>' token is never explained.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to use this versus search_acts or resolve_act_scan, and no stated prerequisite that a known record ID is required. The agent must infer the retrieval-by-ID use case on its own.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

inspect_registerB
Read-only

Discover all scan references; return the proven total and a paginated scan list.

start is one-based. Missing batches/unknown totals are errors, never complete results.

ParametersJSON Schema
NameRequiredDescriptionDefault
urlYes
limitNo
startNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

B3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, openWorldHint=true, and destructiveHint=false, so the safety profile is covered. The description adds two genuinely useful behavioral facts beyond that: pagination is one-based, and partial results (missing batches/unknown totals) surface as errors rather than partial successes. This is real added context, though 'proven total' and 'missing batches' remain unexplained jargon.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences, front-loaded with the tool's purpose and return behavior. Minimal waste, though the trailing one-based note and the error sentence are slightly clipped and could be integrated more cleanly.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be fully explained, and the description appropriately gestures at the result shape. However, with 0% parameter coverage, the meaning of 'limit' and 'url' is missing, and the error condition is stated without saying what a caller should do about it, leaving the definition only minimally complete for a 3-parameter tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must carry the semantics, but it only addresses one parameter: 'start is one-based,' which disambiguates indexing beyond the bare minimum=1 in the schema. The 'limit' and 'url' parameters are given no meaning, so most of the burden is unmet.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Discover all scan references') plus the return shape ('proven total and a paginated scan list'), which tells an agent this is a listing/discovery operation. It does not, however, name or distinguish itself from plausible siblings like search_acts or get_record, so the agent must infer the boundary.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no explicit when-to-use or when-not-to-use guidance, and no sibling is named as an alternative. The closest thing to guidance is the error note ('missing batches/unknown totals are errors'), which is a behavioral constraint rather than a routing rule.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

resolve_act_scanA

Resolve an act using exactly one of record_id/source_url; date means act date.

Default: explicit archive metadata only, with no OCR or image downloads. Optional installed local readers return bounded evidence and candidates. AMBIGUOUS/PROBABLE must not be treated as a confirmed act.

ParametersJSON Schema
NameRequiredDescriptionDefault
dateNo
nameNo
readerNodisabled
max_scansNo
record_idNo
act_numberNo
source_urlNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.6/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Adds real behavioral context beyond the annotations: no OCR or image downloads by default, optional local readers producing 'bounded evidence and candidates', and the rule that AMBIGUOUS/PROBABLE results are not confirmed acts. This tells the agent what the tool will and won't do operationally. The readOnlyHint=false annotation is not explained by the description, though it is not directly contradicted either (local index/cache writes are plausible).

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Four short sentences, front-loaded with the selector constraint and default mode; little waste. The telegraphic phrasing ('bounded evidence and candidates') is compact but slightly cryptic.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The selector rule, default mode, and ambiguity caveat are covered, and an output schema exists so return shape need not be described. However, for a 7-parameter tool with zero schema documentation, the semantics of name, act_number, and max_scans are absent, leaving the agent guessing on non-trivial inputs.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With 0% schema description coverage across 7 parameters, the description must carry the load; it usefully disambiguates 'date means act date' and states the mutual exclusivity of record_id/source_url, and loosely characterizes reader as an 'installed local reader'. It says nothing about name, act_number, or max_scans, so several parameters remain undocumented anywhere.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a clear verb+resource ('Resolve an act') and immediately scopes it with the input selector rule ('exactly one of record_id/source_url'). It implicitly separates the tool from download_act/download_scan by declaring metadata-only default, but it never names or contrasts a sibling explicitly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives a hard usage constraint (exactly one of record_id/source_url) and states the default operating mode, which implies when to leave readers off. It never says when to prefer resolve_act_scan over search_acts, get_record, or download_act, so routing is left to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

search_actsB
Read-only

Search Gelders Archief acts in Open Archieven; date filters the event date.

Multiple records remain separate; number_found counts API person hits. Use record_id for subsequent calls. next_start enables pagination.

ParametersJSON Schema
NameRequiredDescriptionDefault
dateNo
nameYes
yearNo
limitNo
placeNo
startNo
record_typeNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

B3.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover readOnly/openWorld/non-destructive, so the bar is lower, and the description adds real context beyond them: date filters the event date, records stay separate, number_found counts API person hits, and next_start drives pagination. These are non-obvious behavioral traits, not restatements of the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Terse, front-loaded, no filler sentences. The clipped, note-like phrasing (semicolons, fragments) is efficient but slightly telegraphic, costing a point on readability.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 7-parameter search tool with 0% schema documentation, the description covers only date and pagination. It omits how name/place/year/record_type interact, whether date and year can combine, and what limit bounds mean. The presence of an output schema excuses explaining return shape but not the parameter gaps.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% across 7 parameters, so the description must compensate and largely doesn't. Only 'date' gets semantics ('filters the event date'); place, year, record_type, limit, start, and the required name are left entirely to bare schema titles.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource+corpus: 'Search Gelders Archief acts in Open Archieven'. It implicitly separates itself from retrieval siblings (download_act, get_record) by noting records are returned separately and record_id feeds follow-up calls, though it never names a sibling outright.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives downstream usage ('Use record_id for subsequent calls', 'next_start enables pagination') but no explicit when-to-use-this-vs-alternatives guidance and no exclusions. The agent must infer that get_record/download_act are the follow-on tools.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 6 tool updatesv0.1.0
    • First observeddownload_act
    • First observeddownload_scan
    • First observedget_record
    • First observedinspect_register
    • First observedresolve_act_scan
    • First observedsearch_acts

TDQS

A3.5/5.0

Scored across 6 tools

Disambiguation4/5

Most tools have distinct purposes (search, get, inspect, resolve, download), but resolve_act_scan and download_act overlap since download_act also resolves, and download_scan vs download_act both download scans. Descriptions provide enough guidance to differentiate them, keeping this a minor issue.

Naming Consistency5/5

All six tools follow a consistent snake_case verb_noun pattern: get_record, search_acts, inspect_register, resolve_act_scan, download_scan, download_act. The convention is predictable and readable throughout.

Tool Count5/5

Six tools is well-scoped for a focused archival/genealogy resolution workflow. Each tool targets a distinct stage (search, retrieve, inspect, resolve, download) without redundancy bloat.

Completeness4/5

The surface covers the full discovery-to-download lifecycle: searching acts, fetching records, inspecting registers, resolving acts to scans, and downloading scans/acts. Minor gap in that there's no explicit listing of collections/registers beyond inspect_register, but core workflows are covered.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • A
    license
    A
    quality
    A
    maintenance
    Enables users to search and access digital collections from the Swedish National Archives (Riksarkivet) through multiple APIs. Supports searching records by keywords, exploring collections, and downloading historical images and documents.
    2
    26
    Apache 2.0
  • F
    license
    Not graded
    quality
    D
    maintenance
    Provides access to the Netherlands National Archive collections through OAI-PMH and SPARQL for searching, browsing, and retrieving archival records. It features specialized tools for exploring Dutch business history and colonial records using natural language queries.
    -
  • A
    license
    A
    quality
    C
    maintenance
    Enables accessing and querying Dutch government open datasets from CBS and data.overheid.nl, with tools for searching, filtering, downloading, and analyzing data using CSV, Parquet, DuckDB, or Pandas.
    9
    MIT