snac-archives-mcp
Builds ArchiveGrid search and record links for OCLC's ArchiveGrid service, allowing users to open fielded searches or collection records for archival materials that SNAC may not find; it does not automate ArchiveGrid access due to OCLC's terms of use.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@snac-archives-mcpwhich archive holds the papers of Jane Addams?"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
snac-archives-mcp
An MCP server for finding which archive holds the papers: a family's letters, a church's registers, a store's ledgers, a county office's loose papers. Most of that material has never been digitised. It exists online only as a catalogue record or a finding aid in some library, and the hard part of the research is knowing which one.
The server searches the SNAC Cooperative's index of people, families and organisations and the archival collections that hold their papers, across thousands of repositories. The data is CC0, and the API needs no key and no account. It also builds ArchiveGrid searches for you to open, because ArchiveGrid often finds what SNAC cannot, and OCLC does not permit automated access to it.
It works the way a careful genealogist does. Everything it returns is a finding aid: evidence of where to look and roughly what is there, never of what a document says. The tools say so, and they point you at the holding repository's own finding aid, which is the description to cite.
Nothing here writes anywhere, and nothing here keeps a family tree. The server finds collections; what you conclude from the papers belongs in your genealogy software, or in a family-tree MCP server running alongside this one. It sits well beside nara-catalog-mcp (federal records) and familysearch-mcp (indexed records and images).
This is an independent project. It is not affiliated with, endorsed by, or supported by the SNAC Cooperative, the University of Virginia, or OCLC.
Tools
The server publishes eight tools, all read-only and annotated so for the client. Two make no network call at all.
Finding collections
Tool | Purpose |
| Find collections by words in their title. Each hit names its holding repository with a postal address, its OCLC number and its link, and flags the likely duplicates (the same papers catalogued two or three times). |
| One collection in full: the whole abstract, extent, repository, links, and the names SNAC connects to it. |
| What SNAC links to one repository, filtered by title and paged. |
Finding people, families and organisations
Tool | Purpose |
| Find SNAC name records by name, optionally by kind: person, family, or corporate body (churches, businesses, societies, offices). |
| One name record: headings, dates, places, biographical note and identifier links (VIAF, Library of Congress and others). With |
| Collections linked to both of two names: where to look for letters between intermarried families or an ancestor's associates. |
Where SNAC stops
Tool | Purpose |
| Build an ArchiveGrid search, in its fielded syntax, for you to open; or a record link, given an OCLC number. Makes no request. |
| This session's live calls and cache hits. Makes no request. |
Related MCP server: Community Archive MCP Server
Setup
You need Python 3.11 or later and uv. There is no key to request.
Without cloning. uvx fetches it from PyPI and runs it in one step:
uvx snac-archives-mcpFrom a clone, which is what you want if you will change it:
git clone https://github.com/ianderso/snac-archives-mcp
cd snac-archives-mcp
uv sync
uv run snac-archives-mcp # stdio server, usually launched by the clientEither way the server speaks MCP over stdio, so you will normally let an MCP client start it rather than run it by hand.
Claude Desktop
{
"mcpServers": {
"snac": {
"command": "uvx",
"args": ["snac-archives-mcp"]
}
}
}A desktop app does not always inherit your shell's PATH. If the server fails
to start because uvx cannot be found, give the full path that which uvx
prints as the command.
Claude Code
claude mcp add snac -- uvx snac-archives-mcpConfiguration
Nothing is required. A .env file in the directory the server starts in
supplies anything the environment does not; only that directory is read.
Variable | Meaning |
| The SNAC REST endpoint. Default |
| Response cache directory. Default |
| HTTP timeout in seconds. Default 60. Holdings lists get 120. |
| An email address or URL added to the User-Agent, so the SNAC team can reach you if your use causes trouble. Optional, and courteous. |
An unusable value is reported on the first tool call as a not_configured
result naming the variable.
Being a good guest
SNAC is a free service run by a small cooperative, and it publishes no rate
limit. The server sends one request at a time, at least a second apart; two
identical calls in flight share one request; and every answer is cached on
disk for 30 days (a collection record for good). A 429 or a 5xx is retried
three times with back-off, then reported as rate_limited or
upstream_error, which is never the same as "nothing found". Pass
refresh=true to ask again, and refresh a cached empty result before
concluding anything is absent.
How to read what comes back
A finding aid is not the record. A collection description says the papers exist and roughly what is in them. Do not attach a citation to a fact on its strength; record a research task (request copies, plan a visit) instead.
Cite the repository's own finding aid, by the repository and its collection number, for example "Wilder and Anderson Family Papers #01255, Southern Historical Collection, Wilson Library, UNC-Chapel Hill". Not SNAC, and not ArchiveGrid: they are indexes to it. When
link_kindisfinding_aid, the link goes there; otherwise look the collection up in the repository's own catalogue.Record the identifiers. The OCLC number ties a collection together across WorldCat, ArchiveGrid and SNAC; the SNAC ARK (
ark:/99166/...) is the stable id of a name record. Record the repository's collection number too: WorldCat merges records, and a number can come to redirect to another.Description depth varies enormously, from a one-line catalogue record ("Papers, 1881-1958", 4 boxes) to a folder-level container list. A terse description does not mean a person is absent from the papers.
Title search is title search.
search_collectionsneeds every word in the collection's title. "Davenport family papers" finds ten collections; "Davenport family papers Lincoln County" finds none, because the county is only in the abstract. Usesearch_names, or an ArchiveGrid link withplace.Duplicates are normal. One collection often appears as a WorldCat record and as one or two harvested finding aids, sometimes under different forms of the repository's name.
possible_duplicate_ofgroups them by title words and years; it is a hint, not a merge.A name record is not an identification. SNAC holds hundreds of records headed "Anderson family." with nothing to tell them apart but the collections they link to. In a name's collections,
creatorOfmeans these are the name's own papers;referencedInmeans the name is an index term on the collection, not that any document concerns the person.maybe_same_countabove 0 means SNAC suspects a conflation.The papers of slaveholding families are a primary route to enslaved ancestors. Many descriptions name enslaved people only as a category. Search the papers of the family that held them, not only the ancestor's name.
The catalogue ages. SNAC's collection index was largely built from 2010s extracts. Collections get reprocessed, renumbered and transferred, and survey "repositories" such as a state historical documents inventory record what a town clerk or church held when surveyed decades ago. Confirm the current call number before writing to a repository.
ArchiveGrid's indexes have edges.
locationis where the repository is;placeis a place the papers are about. Its name, place and subject indexes cover catalogue records and EAD finding aids only; HTML and PDF finding aids match keywords alone. It leaves out records held by more than one library, so microfilm of county or church records is usually missing: use WorldCat or the FamilySearch Catalog for that.
Deliberately not here
Writing to SNAC. Its edit commands need an account and an API key. The client refuses to send any command but its six read ones, so no argument can make it edit anything.
Fetching ArchiveGrid. OCLC's terms of use forbid robots and automated copying, and the site blocks automated clients. The server builds addresses; a person opens them.
Fetching finding aids from repositories' own sites. That would widen the server from one API to thousands of hosts, with a parser for each and a much larger surface for injected text. It may come later, behind an option.
Security
One host. A request hook refuses any request not for the configured API host, so a value a model passes in cannot make the server fetch another site.
SNAC_API_URLmust be https.Identifiers are validated (numeric ids as ASCII digits, ARKs against SNAC's pattern) before they reach a request.
Catalogue text is untrusted. Abstracts and biographical notes are written by cataloguers and contributors and reach the model verbatim. The server's instructions tell the model to treat that text as material to weigh, never as instructions; the model still decides, so review what it proposes to do.
To report a vulnerability, see SECURITY.md.
Development
uv sync --extra dev
uv run pytest # mocked with respx; never touches SNAC
uv run ruff check .
uv run ruff format --check .
uv run python -m tests.live_check # eight paced calls to the live APIThe live check asks SNAC what the recorded fixtures cannot: whether its answers still have the shape the server reads. See CONTRIBUTING.md for how the suite is organised, docs/API-NOTES.md for what was observed of the API and when, and docs/DESIGN.md for why the server is shaped this way.
Credits
The data is the SNAC Cooperative's, released under CC0, with collection records contributed by its member institutions and drawn from WorldCat and finding aids. ArchiveGrid is a project of OCLC Research.
License
MIT.
Available Tools
8 toolsarchivegrid_search_linkARead-only
Build an ArchiveGrid search address for the user to open. Makes no network call.
ArchiveGrid often finds what SNAC cannot: collections known only by place or subject, and HTML or PDF finding aids. OCLC forbids automated access, so give the link to the user; never fetch it. Fielded indexes cover catalogue records and EAD only; HTML and PDF finding aids match keywords only. Over 90% of hits are collection-level WorldCat records, and microfilm held by many libraries (county or church records) is usually missing.
| Name | Required | Description | Default |
|---|---|---|---|
| event | No | A named event or meeting. | |
| place | No | A place the papers are ABOUT, e.g. 'Lincoln County (N.C.)'. | |
| title | No | Words in the collection title. | |
| topic | No | A subject heading. | |
| family | No | A family name heading, e.g. 'Anderson family'. | |
| person | No | A person's name, e.g. 'Crothers, Robert'. | |
| archive | No | The holding repository's name. | |
| exclude | No | Words or phrases to exclude. | |
| keywords | No | Free words, as typed into ArchiveGrid. | |
| location | No | Where the REPOSITORY is, e.g. 'North Carolina'. | |
| has_links | No | Only records linking to something online. | |
| oclc_number | No | Instead of a search, link one record by its OCLC number. | |
| source_type | No | Only one kind of record. | |
| organization | No | A church, firm, society or office. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and openWorldHint=false; the description reinforces that with 'makes no network call' and adds genuinely new behavioral context: OCLC forbids automated access, fielded indexes only cover catalogue/EAD records, HTML/PDF matches keywords only, ~90% of hits are collection-level WorldCat records, and microfilm holdings are usually absent.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core behavior (build a link, no network call) followed by coverage caveats. Five sentences all carry information, though the result-coverage caveats could be tightened slightly.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description still makes the return clear (an address to open, not fetched content) and sets expectations about what searches will and will not surface. For a 14-parameter URL builder with full schema coverage, nothing essential is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so baseline is 3. The description adds real semantic value on top: the distinction that fielded indexes cover catalogue and EAD records only while HTML/PDF finding aids match keywords only tells the agent when to use fielded params (person, place, topic, archive) versus the free-text keywords param.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: 'Build an ArchiveGrid search address for the user to open,' plus the key qualifier 'Makes no network call.' It also positions itself against SNAC ('ArchiveGrid often finds what SNAC cannot'), though it does not name any of the actual sibling tools (search_collections, search_names) for direct disambiguation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear operational guidance — hand the link to the user, never fetch it — and explains when this source is preferable (place/subject-known collections, HTML/PDF finding aids). It does not explicitly say which sibling tool to choose instead when SNAC is the better fit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
cache_statusARead-only
Report this session's SNAC calls and cache use. Makes no network call.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and openWorldHint=false, so the safety profile is covered. The description adds real value by clarifying scope ('this session's') and reinforcing that no network call is made, telling the agent this is a purely local, side-effect-free diagnostic.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences with zero redundancy; the primary purpose is front-loaded and the clarifying constraint follows immediately. Every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a parameterless, read-only status tool with no output schema, the description covers what is being reported and the key behavioral constraint. It is adequate, though a hint about the shape of the reported data would make it fully self-sufficient without an output schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so the baseline is 4. There is nothing for the description to disambiguate about inputs, and it does not waste words inventing any.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a concrete verb (report) and resource (this session's SNAC calls and cache use), which is specific and unambiguous. It does not, however, name or contrast itself with any sibling tool, so sibling differentiation is left to the agent.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied by the diagnostic nature of the tool, and 'Makes no network call' hints that it is cheap to invoke at any time. There is no explicit statement of when to call it versus alternatives, nor any exclusion criteria.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
collections_in_commonARead-only
List the collections linked to both of two names.
Useful for intermarried families and the associates around an ancestor: papers that mention both are where to look for letters between them. Two names on one collection do not show the two people knew each other.
| Name | Required | Description | Default |
|---|---|---|---|
| first | Yes | One name: a SNAC ARK or constellation id. | |
| second | Yes | The other name: a SNAC ARK or constellation id. | |
| refresh | No | True asks SNAC again instead of using the cache, and replaces the cached copy. Answers are cached for 30 days; refresh a cached empty result before concluding something is absent. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint and openWorldHint, so the safety profile is covered. The description adds interpretive context beyond the annotations — that co-occurrence on a collection is not evidence the two people knew each other — which meaningfully shapes how the agent reports results.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, the core action front-loaded, followed by usage and caveat. No filler or repetition.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple two-name lookup with no output schema, the description states what is returned (the collections linked to both names) and the important interpretation limit. Minor gap: no hint about the shape of returned entries.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so first/second/refresh are fully documented in the schema, including caching and refresh semantics. The description adds nothing about parameters, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('List the collections') with a precise scope: the intersection of two names. This implicitly separates it from sibling search_collections/get_collection, though no sibling is named explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives a concrete usage context (intermarried families, associates around an ancestor) and a meaningful interpretive caveat about what a shared collection does not prove. It stops short of naming alternatives or stating when not to use it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_collectionARead-only
Read one collection description in full, with the names SNAC links to it.
Depth varies: a one-line MARC record with no box or folder list is normal,
and a terse description does not mean a person is absent from the papers.
Follow link (when link_kind is finding_aid) or the repository's own
catalogue to the finding aid; that is the description to cite, by its
collection number. Record oclc_number. "[microform]" in a title means a
copy; the original is elsewhere.
| Name | Required | Description | Default |
|---|---|---|---|
| refresh | No | True asks SNAC again instead of using the cache, and replaces the cached copy. Answers are cached for 30 days; refresh a cached empty result before concluding something is absent. | |
| resource_id | Yes | The collection's SNAC resource id, e.g. '7252207'. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover readOnly/openWorld safety, but the description adds genuinely non-obvious behavior: depth varies, a terse record does not imply absence, and '[microform]' denotes a copy. Caching behavior is disclosed in the schema rather than the description, and no error/auth/rate-limit context is given.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Purpose is front-loaded in the first sentence, but the body drifts into citation lore (microform copies, oclc_number) that is useful yet peripheral to invoking the tool, and the multi-clause sentences are not tightly structured.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description must carry the return-value burden, and it does: it describes linked names, variable depth, the `link`/`link_kind` field, and `oclc_number`. Nothing critical for a correct call is missing, though it never states what to do on an empty result.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so both `resource_id` and `refresh` (including the 30-day cache and refresh-before-concluding-absence rule) are already documented in the schema. The description adds no parameter-level meaning beyond that, so the baseline of 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource — 'Read one collection description in full, with the names SNAC links to it' — and the singular framing implicitly contrasts with the sibling search_collections. It stops short of naming an alternative explicitly, so it is clear but not maximally differentiated.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives concrete downstream guidance: follow `link` when `link_kind` is finding_aid, cite the finding aid by collection number, record `oclc_number`. It does not say when to prefer this tool over search_collections or get_name, so the selection guidance is implicit rather than explicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_nameARead-only
Read one SNAC name record: who or what it is, and the collections linked to it.
In collections, role creatorOf means these are the name's own papers;
referencedIn means the name is an index term on the collection, not proof
the person appears in any document. The biography is copied from a finding
aid: secondary information to follow up, not to cite for a fact.
maybe_same_count above 0 means SNAC suspects this record and another
describe one entity. Record the ark, not the numeric id.
| Name | Required | Description | Default |
|---|---|---|---|
| detail | No | summary: headings, dates, places, biography. full adds every linked collection and related name. | summary |
| name_id | Yes | A SNAC ARK (ark:/99166/w6xj0ds0, or just w6xj0ds0) or a numeric constellation id. | |
| refresh | No | True asks SNAC again instead of using the cache, and replaces the cached copy. Answers are cached for 30 days; refresh a cached empty result before concluding something is absent. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only declare readOnly/openWorld, so the description adds substantial value beyond them: the creatorOf vs referencedIn distinction, the warning that biography is secondary/second-hand data, and that maybe_same_count>0 signals a suspected duplicate entity. The caching/refresh behavior is repeated from the schema rather than new, keeping this just below the top score.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The purpose is front-loaded in the first line, followed by compact, non-redundant interpretation notes. Every sentence carries distinct information, though the multi-clause sentence about creatorOf/referencedIn is dense enough to slow a quick scan.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description carries return-value burden and does explain key returned fields (collections roles, maybe_same_count) and data provenance. It doesn't describe the summary-vs-full result difference or pagination, but the schema's detail enum covers the former, leaving only minor gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so name_id, detail, and refresh are already documented in the schema with formats and defaults. The description adds emphasis ('Record the ark, not the numeric id') but no new syntax or constraint beyond what the schema provides, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The first sentence states a specific verb and resource with scope: 'Read one SNAC name record: who or what it is, and the collections linked to it.' The singular 'one ... record' plus the mention of linked collections cleanly separates it from search_names and get_collection without needing to name them.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives real operational guidance: 'refresh a cached empty result before concluding something is absent' and 'Record the ark, not the numeric id.' What's missing is explicit routing advice among siblings (e.g. use search_names first to obtain the id, or get_collection for the collection side), so it stops short of full when/when-not coverage.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
repository_holdingsARead-only
List the collections SNAC links to one repository, filtered by title.
The first call for a large repository is slow (10 to 20 seconds; it is SNAC building a list of thousands) and is then cached. The list is SNAC's link table, not the repository's catalogue: it includes printed and microform items and misses recent accessions. Survey "repositories" such as a state historical documents inventory record what a town or church held when surveyed decades ago; the material may since have moved.
| Name | Required | Description | Default |
|---|---|---|---|
| page | No | Page of results, 1-based. | |
| count | No | Collections per page (1-100). | |
| refresh | No | True asks SNAC again instead of using the cache, and replaces the cached copy. Answers are cached for 30 days; refresh a cached empty result before concluding something is absent. | |
| repository | Yes | The holding repository: its SNAC ARK or constellation id, as get_collection's repository field gives it. | |
| title_contains | No | Keep only collections whose title contains this text. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With readOnlyHint and openWorldHint already covering the safety profile, the description still adds meaningful behavior: the first call for a large repository takes 10-20 seconds while SNAC builds the list, results are then cached, and the underlying data source has known coverage gaps. This is exactly the kind of operational and provenance context annotations cannot express.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The core purpose is front-loaded in one sentence, and the following sentences justify themselves by qualifying result interpretation and latency. The phrasing is slightly discursive (paragraph on survey repositories) but nothing is redundant with the schema or annotations.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema, so the description must carry the burden of telling the agent what it gets back; it does this well by characterizing the list's provenance, coverage, and latency. It stops short of describing the record shape or pagination behavior, though pagination is covered by the input schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and each parameter is already documented, including refresh's 30-day cache semantics. The description only reinforces the title filter ('filtered by title') and the single-repository scope, adding little syntax or semantics beyond the schema. Baseline 3 applies when the schema carries the parameter burden.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The opening sentence states a precise verb, resource, and scope: it lists the collections SNAC links to a single repository, with an optional title filter. The 'one repository' scoping distinguishes it from the sibling search_collections, which performs cross-repository search. An agent can tell which tool to reach for without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives strong contextual guidance on when results are trustworthy: the list is SNAC's link table rather than the repository's catalogue, it may miss recent accessions, and 'survey repositories' reflect historical holdings that may have moved. It does not, however, explicitly name an alternative (e.g. search_collections) or state when not to use this tool, so it falls short of full routing guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
search_collectionsARead-only
Find archival collections by words in their title, with who holds each.
Matches the collection TITLE only, and every word must appear: a county or
second surname named only in the abstract will not match, so use
archivegrid_search_link or search_names for those. The same papers often
appear two or three times (a WorldCat record plus harvested finding aids);
possible_duplicate_of flags them. SNAC's catalogue is largely a 2010s
snapshot. A hit tells you where papers are, not what they say: confirm the
collection number in the repository's own catalogue, and cite its finding
aid, never this result.
| Name | Required | Description | Default |
|---|---|---|---|
| page | No | Page of results, 1-based. | |
| count | No | Collections per page (1-50). | |
| refresh | No | True asks SNAC again instead of using the cache, and replaces the cached copy. Answers are cached for 30 days; refresh a cached empty result before concluding something is absent. | |
| title_words | Yes | Words that must all appear in the collection's title, e.g. 'Davenport family papers'. Places and second surnames are usually not in the title. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint and openWorldHint, so the read-only, open-world profile is covered. The description adds real value beyond that: duplicate WorldCat/finding-aid records flagged via possible_duplicate_of, the 2010s SNAC snapshot caveat, and the warning that a hit locates papers but does not describe their contents. Since annotations carry the safety profile and the description's extras are caveats rather than behavioral mechanics like caching or pagination, a 3 fits.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the purpose, then constraints, then caveats, with no filler sentences. Slight redundancy where 'Matches the collection TITLE only' restates part of the title_words schema description, but each block carries distinct information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, so the description must convey result semantics, and it does: it names the possible_duplicate_of field, explains that provenance/citation should come from the repository, and warns about catalogue staleness. Nothing needed to call or interpret it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so baseline is 3, but the description adds a nuance the schema only partially covers: a place or second surname appearing in an abstract rather than the title will not match, which sharpens the title_words contract. The page, count, and refresh parameters are left entirely to the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Find archival collections') plus what a result includes ('with who holds each'). It also distinguishes itself from siblings by naming search_names and archivegrid_search_link as the tools for names/places that won't match a title search.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly says what matches (TITLE only, all words must appear) and what does not, then routes the agent to archivegrid_search_link or search_names for those cases. It adds a concrete condition for retrying (refresh a cached empty result before concluding absence).
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
search_namesARead-only
Find SNAC name records for a person, family or organisation.
A heading such as "Anderson family." matches hundreds of unrelated families, usually one record per source collection, with no place or date to tell them apart. A SNAC record is a machine-built authority record, not an identification: narrow with get_name, reading the collections it links to and their places and dates, before treating any record as your family.
| Name | Required | Description | Default |
|---|---|---|---|
| name | Yes | A name as a heading would carry it, e.g. 'Anderson family'. | |
| page | No | Page of results, 1-based. | |
| count | No | Names per page (1-50). | |
| refresh | No | True asks SNAC again instead of using the cache, and replaces the cached copy. Answers are cached for 30 days; refresh a cached empty result before concluding something is absent. | |
| entity_type | No | Restrict to one kind of name. corporateBody covers churches, businesses, societies and government offices. | |
| search_biographies | No | Also match words in the biographical notes. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations cover readOnlyHint and openWorldHint, so safety is addressed. The description adds real behavioral context beyond that: results are ambiguous, headings can match hundreds of unrelated families, and a SNAC record is a machine-built authority record rather than an identification. It does not discuss result paging behavior, but the ambiguity warning is substantive value.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Purpose is front-loaded in the first sentence, followed by the caveat that justifies it. The prose is slightly discursive with the quoted example and a rather long closing clause, but every sentence carries information and nothing is redundant.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a search tool with no output schema and a fully documented input schema, the description supplies the crucial missing context: that results are ambiguous authority records and must be validated via get_name with linked collections, places and dates. It could say a bit more about what a result row contains, but the interpretive warning is the main need and is covered.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all six parameters already carry descriptions. The description only echoes the heading-input convention ('Anderson family') that the schema's name property already documents, adding no syntax or format meaning beyond the structured fields. Baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Find SNAC name records') and immediately scopes it to person, family or organisation. It also distinguishes itself from the sibling get_name by naming it as the narrowing step, so an agent can route correctly without opening either schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear context for use (discovery of name records) and names the alternative workflow (narrow with get_name before treating any record as your family). It stops short of explicit when-not-to-use guidance or pointing at search_collections as the contrasting entry point, so it is strong but not complete.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
8 tool updates
v0.1.0- First observed
archivegrid_search_link - First observed
cache_status - First observed
collections_in_common - First observed
get_collection - First observed
get_name - First observed
repository_holdings - First observed
search_collections - First observed
search_names
TDQS
Scored across 8 tools
Each tool targets a distinct object or action: collection search vs. name search, collection retrieval vs. name retrieval, cross-name collection overlap, repository holdings, external link building, and cache reporting. The descriptions explicitly clarify boundaries and limitations, so an agent should not confuse them.
All names use snake_case and are readable, with consistent search_/get_ pairs for the core lookup tools. However, several tools use noun phrases (collections_in_common, repository_holdings, cache_status) rather than the verb_noun pattern, which is a minor deviation.
Eight tools are well-scoped for a read-only archival discovery server. Each tool earns its place by covering a distinct part of the SNAC research workflow without obvious redundancy.
The surface covers core discovery and retrieval: searching and reading collections and names, finding shared collections, browsing repository holdings, and generating an external ArchiveGrid link. Minor gaps exist, such as direct subject/place search within SNAC or citation export, but agents can work around these via the provided tools and link builder.
Maintenance
Related MCP Connectors
Read-only OpenHeritage search for genealogy and cultural heritage records.
Search LOC digital collections, Chronicling America newspapers (full OCR), and LC Subject Headings.
Machine-readable entity discovery with provenance, trust and verified source evidence.
Search 14.5M Smithsonian Open Access objects, get CC0 images, find cross-collection connections.
Related MCP Servers
- AlicenseAqualityAmaintenanceEnables users to search and access digital collections from the Swedish National Archives (Riksarkivet) through multiple APIs. Supports searching records by keywords, exploring collections, and downloading historical images and documents.226Apache 2.0
- FlicenseNot gradedqualityDmaintenanceEnables searching and retrieving preserved Twitter data from the Community Archive, including user profiles, tweets, and keyword searches across archived content.-
- FlicenseNot gradedqualityBmaintenanceEnables querying genealogical records, archive statistics, historical weather, and full-text transcriptions from Open Archives via natural language.12 npm4-
- FlicenseAqualityCmaintenanceEnables searching and retrieving historical Swedish portraits, biographies, and printed source references from Svenskt Porträttarkiv, including direct image links and genealogical citations.3-