football-docs
football-docs is an MCP server that gives AI agents searchable, provenance-backed documentation for ~30 football data providers and tools, plus entity resolution, paper/web-source research, and a local library.
Search provider docs (
search_docs): full-text search across event types, qualifier IDs, API endpoints, coordinate systems and data models, with optional provider filter, partial-match marking, and notices for unmentioned terms or non-indexed providers.Explore the index (
list_providers,resolve_provider_id,get_provider_docs,compare_providers): list indexed providers and coverage, resolve names/aliases to canonical keys, fetch one provider's docs by topic/category, and compare how providers handle the same concept.Request doc changes (
request_update): queue requests for a new provider, recrawl, flag outdated docs, or suggest a better source.Resolve entities (
resolve_entity): map players, coaches, referees, teams, competitions, seasons, stages and matches to IDs across providers via the Reep register (local DuckDB copy or Reep API).Research papers (
search_papers,get_paper): search OpenAlex, arXiv, SportRxiv and optionally Zotero; look up one paper by DOI, arXiv ID, OpenAlex ID or Zotero item.Read and verify sources (
get_web_source,read_paper,match_quote): read public web pages/PDFs and open-access papers by section with licences, and check quotes as exact, normalised, close or none.Manage a local library (
add_local_paper,forget_paper,purge_cache): add a PDF you own, remove one paper, or delete the whole cached library.Stay current: provider docs auto-update from a signed daily index (no restart needed) and work offline or with a bundled/pinned index; paper tools can be disabled via settings.
Searches arXiv for scholarly papers on football analytics and sport science (by title, abstract and authors), and resolves individual arXiv IDs/URLs to paper metadata — title, authors, date, venue, licence and open full-text copies that can be read by section.
Searches the user's own Zotero library for papers (using Zotero on the local machine, or the Zotero web API with an API key), and accepts zotero: item keys as paper IDs for metadata lookup, reading the outline/short passages of user-supplied papers, and quote matching.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@football-docsWhat is Opta qualifier 76?"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
football-docs
Searchable football data provider and tooling documentation for AI coding agents. Like Context7 for football data.
Who it's for: Developers and analysts who use AI coding tools (Claude Code, Cursor, VS Code Copilot) to work with football data. Works with any tool that supports MCP.
What it does: Gives your AI agent a searchable index of documentation for 30 football data providers and tools — event types, qualifier IDs, coordinate systems, API endpoints, data models, identity surfaces, and cross-provider comparisons for the data providers (StatsBomb, Opta, Wyscout, Impect, SkillCorner, Sportradar, TheSportsDB, FMDB Pro, TransferRoom, and more), the open-source libraries people build with (kloppy, mplsoccer, socceraction, soccerdata, floodlight, fast-forward, unravelsports, and more), and the APIs of wearable and sports-science vendors (STATSports, Firstbeat, Hawkin Dynamics, VALD). Your agent looks up the real docs instead of guessing from training data.
Why not just let the AI figure it out? LLMs get football data specifics wrong constantly — Opta qualifier IDs, StatsBomb coordinate ranges, API endpoint URLs, library method signatures. These are mutable facts that change across versions. football-docs gives the agent verified, sourced documentation with provenance tracking so you know where every answer came from.
Strategy
football-docs is intended to be a community-owned, source-transparent Context7 for football data. The public operating contract is in STRATEGY.md: what belongs here, what must stay out, how we handle public-safe provider facts, and how contributors should prove retrieval quality.
Related MCP server: Unified Docs Hub
Provider identity facts
football-docs is the public source for provider identity-surface facts: access shape, ID schemes, matching fields, provider quirks, and provenance rules. Curated provider identity notes belong here when they can be stated as public facts about the provider. They should say whether a fact comes from public docs, public page evidence, licensed feed shape, or a reviewed public-safe observation, and must not include credentials, local paths, internal tooling details, or restricted payloads from any private project.
MCP (Model Context Protocol) is a standard for connecting AI coding tools to external data sources.
Quick start
The server needs Node.js 22.13 or newer. It uses Node's built-in SQLite, so no native module is compiled at install time.
Claude Code
claude mcp add football-docs -- npx -y football-docsCursor
Settings → MCP → Add server. Use this config:
{
"mcpServers": {
"football-docs": {
"command": "npx",
"args": ["-y", "football-docs"]
}
}
}VS Code / Copilot
Add to .vscode/mcp.json:
{
"servers": {
"football-docs": {
"command": "npx",
"args": ["-y", "football-docs"]
}
}
}Claude Desktop
Add to claude_desktop_config.json:
{
"mcpServers": {
"football-docs": {
"command": "npx",
"args": ["-y", "football-docs"]
}
}
}After an update
npx updates its cached copy of football-docs in place. A server that was
already running keeps the old version until it restarts, and its replies then
start with a note: "football-docs on disk is now vX, but this server is still
running vY". Reconnect the server to fix it (in Claude Code: /mcp, then
reconnect football-docs).
Tools
Tool | Description |
| Full-text search across all provider docs. Filter by provider. Results include provenance (source URL, version), mark partial matches, name query terms that no indexed doc mentions, and say when a query names a provider that is not indexed (the |
| Resolve provider names and aliases to canonical indexed provider keys before searching. |
| Retrieve docs for a resolved provider, optionally filtered by topic or category. |
| List all indexed providers and their doc coverage. |
| Compare how different providers handle the same concept. |
| Request a new provider, flag outdated docs, or suggest a better doc source. Queues locally and points to the matching public GitHub issue template. |
| Map a player, coach, referee, team, competition, season, stage or match to its IDs at every provider through the Reep register. Uses a local copy of the free register ( |
| Search scholarly papers on football analytics and sport science in OpenAlex (title, abstract and full text), arXiv and SportRxiv, and, when asked, your Zotero library. See Papers and web sources. |
| Look up one paper by DOI, arXiv ID, OpenAlex ID or Zotero item: authors, date, venue, all IDs, licence, open copies with their licences, abstract and a citation line. |
| Read a public web page or PDF (a blog post, newsletter, club or vendor article, or an author's copy of a paper) as text, with its author, date, licence and Wayback Machine snapshots. Long pages come back by section. |
| Read a paper: the full text of an open copy, by section, with its licence; or the outline and short passages of a paper you supplied. Kept in your library, so a second read sends no request. |
| Check that a quote appears in a paper or web page: exact, normalised, close (with a score) or none, with the section, page and a W3C text quote selector. |
| Add a PDF you have to your library, for papers with no open copy. |
| Remove one paper from your library. |
| Delete your whole paper library (needs |
Provider filters use the indexed provider keys shown by list_providers, but common aliases are accepted. Examples: fbref, understat, ClubElo, football-data.co.uk, and engsoccerdata search free-sources; Sofascore searches soccerdata; ESPN, ESPN FC, and espn-soccer search espn; FMDB searches fmdb-pro; Transfer Room searches transferroom; Hudl Wyscout searches wyscout; Stats Perform / Opta F24 / WhoScored search opta; Metrica, Sportec / DFL, and TRACAB search databallpy; Second Spectrum searches kloppy; Hawk-Eye, SciSports, Signality, Respovision, GradientSports and OptaVision search fast-forward; unravel searches unravelsports; SportRadar API / Soccer Extended search sportradar; Sonra / Apex search statsports; Hawkin searches hawkin-dynamics; ForceDecks / NordBord / ForceFrame search vald; The Sports DB / TSDB search thesportsdb; StatsBomb Open Data searches statsbomb.
How the index stays current
The package ships a docs index, but doc fixes do not wait for a new package
version. Once a day at most, the server checks the
data-latest
release of this repository for a newer index built from main:
It reads a small signed manifest (
manifest-v1-signed.json) and, only when that names a newer build, downloads the index (about 6 MB).It accepts the manifest only if its ed25519 signature verifies against a public key shipped in the package (
src/data-signing.ts). The manifest carries the index's size and SHA-256, so the signature covers the index too.It checks the download's size and SHA-256, its SQLite integrity, its exact schema and its metadata before using it, and keeps the one it had on any failure.
It stores downloads in
$XDG_DATA_HOME/football-docs/data/(by default~/.local/share/football-docs/data/). The next tool call uses the new file; no restart is needed.It runs in the background after start-up and never delays a tool call.
list_providersends with the build time of the index in use and whether it is bundled or downloaded.
Privacy: the check is one HTTPS request a day to GitHub, plus the download when there is one. No search queries or other usage data are sent.
Settings, as environment variables in the server's MCP configuration:
FOOTBALL_DOCS_DATA=bundled: use only the index inside the package. No update check, no downloads. (This does not turn off the paper tools; seeFOOTBALL_DOCS_PAPERSunder Papers and web sources.)FOOTBALL_DOCS_DATA=auto: the default for an installed package. Check daily and use the newest valid index.FOOTBALL_DOCS_DB_PATH=<file>: use exactly this index file.FOOTBALL_DOCS_DATA_BASE_URL=<url>: read the manifest and index from a mirror (HTTPS only).
Offline, or behind a proxy that Node's built-in fetch does not use, the check
fails quietly and the server keeps the index it has. Under CI (CI set) the check
does not run. A server running from a git checkout defaults to bundled, so
development and tests always use the working tree's docs.
Papers and web sources
Many football methods come from papers, and some from blog posts: Karun Singh
introduced expected threat (xT) in a blog post, not a paper. search_papers
finds papers and the works that cite an idea; a web search finds the original
post, and get_web_source reads it. get_paper checks that a reference exists,
read_paper reads it, and match_quote checks that a quote is in the source.
Open and paid papers
Paper | How it gets in | What the tools return |
Open copy: arXiv, an open repository, an open-access publisher, SportRxiv, a public web page |
| The full text, by section, with its licence |
A paper you have through a subscription or purchase | You download the PDF and call | The outline and passages of at most 200 characters ( |
football-docs never logs in to a publisher or a library, never holds your credentials or cookies, and does not use publisher text-mining APIs (their terms exclude tools like this). When a site answers with a bot check, the tool stops; download the paper yourself instead.
Zotero: search_papers with sources: ["zotero"] searches your library, and
read_paper reads an item's PDF. Two ways to connect, tried in this order:
Zotero on this computer. In Zotero 7, open Settings > Advanced and turn on "Allow other applications on this computer to communicate with Zotero". The tools then use its local API on port 23119; nothing leaves the machine.
The Zotero web API, for when Zotero is not running here (another machine, a cloud session). Create a read-only key at zotero.org/settings/keys with access to your library and its files, and set
ZOTERO_API_KEY(and optionallyZOTERO_USER_ID; without it the tools ask the key). It reads only files kept in Zotero's own storage, not linked files; for those it uses Zotero's full-text index. The key is sent only to api.zotero.org, never to the file storage host.
An item saved without its PDF (metadata only) is read by its DOI instead, so
read_paper still returns the open copy when there is one.
Services
The tools call public services at run time. Each reply ends with the services it asked.
Service | Used for | Limits |
OpenAlex | Search (title, abstract and full text); DOI and OpenAlex ID lookups; open copies | 1000 credits a day without a key: a search costs 10, a lookup costs nothing. A free key from openalex.org/settings/api raises it. |
arXiv | Search (title, abstract, authors); arXiv ID lookups with the paper's licence; the paper's HTML or PDF | One request every three seconds to the API; the tool waits its turn. |
SportRxiv | Search in a local copy of its OAI feed; preprint PDFs | The first search downloads the feed (under 1000 records); after a week, the next search asks only for changes. |
Crossref | DOI lookups that OpenAlex does not know | |
Wayback Machine | The earliest and latest snapshot of a web page; the archived copy when the live page fails | The tool never asks it to save a page. |
The host of an open copy or web page | The text | Bot checks stop the tool. |
Your library
Text the tools read is kept in
$XDG_DATA_HOME/football-docs/papers/(by default~/.local/share/football-docs/papers/). On macOS and Linux the folder and its files are readable only by you (modes 0700 and 0600); on Windows they have the access rules of your user folder. A second read sends no request.forget_paperremoves one paper;purge_cachewithconfirm: truedeletes the whole library and the SportRxiv copy. Neither touches your own files or Zotero.Nothing in the library goes into this repository,
data/docs.dbor a data release. A test fails if a cached paper or any PDF is added to the repository.get_web_source,read_paperandmatch_quoteread only publichttpandhttpsaddresses. They refuse names that resolve to loopback, private or link-local addresses, also after a redirect.add_local_paperreads only PDF files, given by their full path.
Settings
As environment variables in the server's MCP configuration:
FOOTBALL_DOCS_PAPERS=off: send no paper or web request, including to the Zotero web API. Your library, files you add and Zotero on this computer still work. The provider-doc tools never call these services.ZOTERO_API_KEY=<key>andZOTERO_USER_ID=<number>: a read-only Zotero web API key, for reading Zotero when the app is not running here. The key can also go in the keychain, as for OpenAlex below.FOOTBALL_DOCS_PAPERS_PASSAGE_CHARS=<n>: the longest passage returned from a paper you supplied, 50 to 1000 characters (default 200).OPENALEX_API_KEY=<key>: your OpenAlex key. On macOS and Linux you can keep it in the system keychain instead of the environment:macOS:
security add-generic-password -s football-docs -a OPENALEX_API_KEY -w <key>Linux (needs
secret-tool, from libsecret):secret-tool store --label=football-docs service football-docs key OPENALEX_API_KEYWindows: use the environment variable; the tools do not read Windows Credential Manager.
Example queries
"What is Opta qualifier 214?" (big chance)
"How does StatsBomb represent shot events?"
"Compare Opta and Wyscout coordinate systems"
"What player ID fields does Transfermarkt expose?"
"Does SportMonks have xG data?"
"What event types does kloppy map to GenericEvent?"
"How does SPADL represent a tackle?"
"Who introduced expected threat (xT)? Cite the original." (a web search finds the post,
get_web_sourcereads it,match_quotechecks the quote,search_papersfinds the papers that cite it)"How does VAEP define the value of an action? Quote the paper." (
read_paperon arXiv 1802.07127, thenmatch_quote)
Indexed providers
Provider | Chunks | Categories |
fast-forward | 250 | overview, getting-started, data-model, coordinate-system, orientations, layouts, transformations, distributed-compute, api-reference, 12 provider format pages |
StatsBomb | 244 | event-types, data-model, coordinate-system, api-access, api-endpoints, charting-lineups, xg-model, iq-metrics, player/team stats, player-mapping, identity-surfaces, data-provenance |
unravelsports | 202 | overview, installation, quickstart, concepts, graph converters, pressing intensity, formation detection, models, utils, american-football |
Wyscout | 165 | event-types, data-model, coordinate-system, api-access, api-endpoints, charting-analysis-metrics, glossary, identity-surfaces, data-provenance |
kloppy | 126 | data-model, usage, provider-mapping, tracking-rendering, event-derived-metrics |
floodlight | 144 | core data objects, io parsers (Tracab, DFL, Kinexon, Opta, SkillCorner, StatsBomb, StatsPerform, Second Spectrum), transforms, metrics, models, visualisation, guides |
SportMonks | 568 | full v3 endpoint reference (fixtures, livescores, leagues, seasons, states, types, statistics, brackets), syntax and includes, filtering, rate limits, error codes, changelog, plus curated event-types, data-model, api-access, charting-season-stories, identity-surfaces, data-provenance |
databallpy | 63 | data-model, overview, usage |
mplsoccer | 65 | overview, pitch-types, visualizations |
Impect | 79 | overview, data-model, event-types, coordinate-system, concepts, kpi-definitions, identity-surfaces, data-provenance |
SkillCorner | 51 | api-access, api-endpoints, data-model, physical-data, coordinate-system, concepts, identity-surfaces, data-provenance |
Free sources | 69 | overview, fbref, understat, football-data-columns, contextual-story-joins, xg-timelines, data-provenance |
soccerdata | 40 | overview, data-sources, usage |
TransferRoom | 45 | api-access, api-endpoints, charting-availability, data-model, identity-surfaces, data-provenance |
Opta | 73 | event-types, qualifiers, coordinate-system, api-access, charting-game-state, charting-lineups, charting-passmaps, charting-set-pieces, charting-shot-placement, identity-surfaces, data-provenance |
FMDB Pro | 37 | api-access, api-endpoints, data-model, identity-surfaces, data-provenance |
Sportradar | 481 | integration guide (API basics, coverage tiers, ID handling, match status, update frequencies, push, historical data), Soccer Extended v4 endpoint reference with data-point tables, FAQ, plus curated api-access, api-endpoints, data-model, charting-and-stories, integration-notes, data-provenance |
socceraction | 34 | SPADL format, VAEP, Expected Threat |
BeSoccer | 16 | api-access, api-endpoints, data-provenance |
Driblab | 32 | api-access, api-endpoints, data-model, data-provenance |
ESPN | 21 | api-access, scoreboard, match-summary, teams-and-standings, identity-and-coverage |
Reep | 29 | overview, identity-and-ids, download-duckdb-csv, api, data-provenance |
TheSportsDB | 20 | api-access, api-endpoints, livescore, identity-surfaces, data-provenance |
FotMob | 5 | identity-surfaces, data-provenance |
Soccerdonna | 5 | identity-surfaces, data-provenance |
Transfermarkt | 5 | identity-surfaces, data-provenance |
STATSports | 70 | api-access, api-endpoints, data-model, drill-kpi-metrics (all 319 DrillKpiV7 fields), identity-surfaces, data-provenance |
Firstbeat | 79 | api-access, api-endpoints, data-model, variables (100 scalars, 17 time series), identity-surfaces, data-provenance |
Hawkin Dynamics | 69 | api-access, api-endpoints, data-model, test-metrics (535 metrics across 13 test types), identity-surfaces, data-provenance |
VALD | 318 | api-access, api-endpoints, per-product endpoints and schemas (tenants, profiles, forcedecks, nordbord, forceframe, smartspeed, dynamo, humantrak), identity-surfaces, data-provenance |
3,405 searchable chunks across 30 providers and tools.
STATSports, Firstbeat, Hawkin Dynamics and VALD sell wearables and testing devices. Their APIs return a customer's own athlete data, which includes personal and health data. The docs describe the API surface only: endpoints, field names, types and units as each vendor's public spec or reference page states them. They contain no athlete data, and none of these vendors' IDs join to a football data provider's IDs.
ESPN coverage consists of curated, dated observations of ESPN-hosted soccer endpoints, checked for eng.1 and esp.1. These observations are not an official API contract or an open-data licence. See ESPN access notes for source status and coverage notes for the tested requests and limitations.
Impect documentation is built solely from the public ImpectAPI/open-data repository — a static Bundesliga 2023/24 snapshot, representative of Impect's structure and metric definitions rather than a complete or current mirror. Impect's commercial API is deliberately not documented here. Every enum value, KPI name and field name in
docs/impect/is validated against that repository in CI (pnpm impect:truth,src/__tests__/impect-open-data-validation.test.ts). Data source: Impect; use is subject to the repository's own Terms of Use.
Documentation validation
Docs for AI agents are only useful if they are correct, and prose about an API is exactly the kind of thing that drifts or gets invented. Where a machine-readable source of truth exists, this repo checks the docs against it in CI rather than trusting them.
Providers | Ground truth | Checked by |
kloppy, socceraction, soccerdata, mplsoccer, floodlight, databallpy, skillcorner, fast-forward, unravelsports | The installed package itself — enum members, importable symbols, class constants, |
|
Wyscout, SkillCorner, FMDB Pro, Sportradar | The vendor's own publicly published OpenAPI spec — endpoint paths and methods |
|
BeSoccer | The vendor's published Postman collection — request vocabulary and parameters |
|
Driblab | The vendor's published API guide — endpoint paths and methods. Field and group names are not in CI: the guide disagrees with the live API on them, so the docs follow the live API, checked by hand on 2026-09-22 |
|
Impect | The public open-data repository |
|
StatsBomb | ID and name pairs observed in a sample of the public open data (three matches per competition-season, at a pinned commit). IDs the sample lacks must be listed in both the Open Data Events specification and kloppy's parser. Regenerate with |
|
Opta | Stats Perform's F24 appendices, which have no machine-readable form. |
|
SportMonks | Type and state IDs as SportMonks publishes them on its definitions pages and in the types spreadsheet linked from its Types page ( |
|
Reep | The public OpenAPI spec (endpoint paths and methods), and one release's manifest and column schema (CSV table list, columns used in the SQL examples, licence and exclusions) |
|
STATSports, Firstbeat, Hawkin Dynamics, VALD | Each vendor's public OpenAPI spec (STATSports v5 to v7, Firstbeat's |
|
ESPN has a separate observation check in src/__tests__/espn.test.ts. It validates
documented endpoint paths, field-table names and types, and source URLs against
data/espn-observations.json. These are selected structural observations, not a
published specification. CI reads them offline and does not contact ESPN.
Refresh manually with python3 scripts/observe_espn.py --scheduled-date YYYYMMDD,
choosing a future fixture date and using permitted access. The script stores no
raw responses. Review the diff, update the docs and dates, then rebuild the index.
Reep's snapshots age weekly, as a new release is cut each week. Run
python3 scripts/check_reep_live.py to check every download link and list what
has changed since the snapshots; the refresh steps are in specs/README.md.
pnpm check:upstream runs every upstream check in one report: pinned package
versions against PyPI, npm and GitHub, each spec snapshot against its public URL,
BeSoccer's and Driblab's truth regenerated from their Postman collection and Notion
page, and the free-source and Reep live checks below. It runs by hand, never in CI, and
lists what to review.
Free sources have no spec, and they tend to fail quietly: a page still answers
200 after the data has gone. Run python3 scripts/check_free_sources_live.py to
request each access path documented in docs/free-sources/ and check that the
response still carries the data the doc describes. When a doc's access recipe
changes, change its check in the same commit.
Truth files live in data/provider-truth/ and are generated, not hand-written:
pnpm provider:truth # rebuild every package truth file (needs python3.11)
pnpm openapi:truth # rebuild every spec-derived truth fileThe specs those derive from are snapshots of publicly published, unauthenticated
vendor documentation. Source URLs, fetch dates and refresh instructions are in
specs/README.md. Wyscout's v3 and v4 specifications merge into
one truth file, because its docs span both. The v2 legacy specification is no
longer mirrored: the docs describe v3 and v4, and no documented fact derives from
the legacy surface.
Each package gets its own pinned venv — co-installing them makes pip silently
downgrade conflicting versions, which would produce truth that disagrees with the
docs. Bump a pin in scripts/gen_all_truth.sh and the matching version in
providers.json together, then re-run and fix whatever the tests flag.
A doc that names an enum member or importable symbol which does not exist in the
real package fails the build. scripts/gen_openapi_truth.py derives the same kind
of facts from a vendor OpenAPI spec, for providers documented that way.
Not every vocabulary is an enum. fast-forward's coordinate systems, orientations
and layouts are lowercase strings on Literal-annotated parameters, so the truth
files also record what each parameter accepts, and a doc writing
coordinates="statsbomb" fails the same way an invented enum member would.
Contributing
Contributions are welcome from everyone. There are three ways to help:
Open an issue — request a new provider, flag outdated docs, or suggest a better doc source
Use the
request_updatetool — AI agents can flag outdated or missing docs directly via the MCP server, which queues requests locally and points to the matching public GitHub issue templateOpen a PR — fix errors, add new providers, or improve existing docs
You don't need to be an expert. See CONTRIBUTING.md for the full guide.
For maintainers
Crawl pipeline
Provider doc sources are tracked in providers.json. The crawl pipeline discovers the best doc source (llms.txt > ReadTheDocs > GitHub README) and writes markdown with provenance frontmatter.
Two registry fields narrow a crawl. llms_indexes lists llms.txt files that link to pages rather than contain them; the crawler follows those links (and nested indexes) to each page's markdown copy instead of running discovery. Sportradar uses it for its integration guide and Soccer Extended reference. exclude_categories skips whole pages, and exclude_sections drops named ##/### sections from pages that are otherwise kept, for content INCLUSION.md keeps out, such as odds. Pages crawled through an index are not byte-for-byte copies: the crawler drops the OpenAPI definition ReadMe appends and SVG diagrams, cuts example payloads over 3 KB with a note, replaces a data-point table repeated from an earlier page with a line naming that page, and marks a split long section "(continued)". It adds no other text.
npm run discover # probe sources without crawling
npm run crawl # crawl all providers with sources
npm run crawl -- --provider kloppy # crawl one provider
npm run ingest # rebuild search index from docs/
npm run ingest -- --provider kloppy # re-ingest one provider (incremental)Each crawled doc carries provenance metadata (source URL, source type, upstream version, crawl timestamp) that is surfaced in search results, so agents can distinguish between curated content and upstream documentation.
How doc changes reach users
A merge to main that touches docs/ or providers.json runs
.github/workflows/data.yml, which rebuilds the index, runs the tests against it,
checks that the published server can use it (scripts/check-data-compat.mjs),
and publishes it to the data-latest release. Installed servers pick it up within
about a day, including servers that have been running for days (they check again
every six hours, at most once a day). Merging a doc PR is therefore also shipping it; there is no later step at
which to stop it. To undo a bad doc change, revert it on main, which publishes a
newer build.
What the signature guarantees. The publish job signs the manifest with a key
held only in the data-publish environment, whose deployment policy allows only
main. A manifest that verifies was therefore produced by data.yml running on
main. Replacing release assets by hand, or from a workflow on another branch,
cannot produce one. It does not mean anyone reviewed the change: main
accepts a PR with no approvals, so anyone who can merge to main can still ship
data. Requiring approvals on main, or required reviewers on the data-publish
environment (which makes every doc publish wait for a person), would close that
too.
Rotating the key: generate a new ed25519 pair, add its public key to
TRUSTED_KEYS in src/data-signing.ts, release, then replace the
DATA_SIGNING_KEY environment secret and DATA_SIGNING_PUBLIC_KEY in data.yml.
Drop the old key in a later release. A lost private key is handled the same way.
The index carries its own metadata (meta table, src/data-format.ts): a schema
version, the oldest server version that can read it, the build stamp and the
provider registry. Two rules follow:
Changing the schema (
SCHEMA_SQL) needsDATA_SCHEMA_VERSIONraised and an npm release. Servers read onlymanifest-v<their schema>.json, so older servers keep their last compatible data rather than breaking.Data that relies on new server code (for example a new
providers.jsonfield the tools must read) needsMIN_SERVER_VERSIONraised, and that server released first. Until then the compatibility check fails, and no installed server would accept the build anyway.
Code changes, including changes to src/ingest.ts, still ship only through a
release.
Cutting a release
Pushing the tag is the release. .github/workflows/release.yml runs on any
v* tag and does the rest: it re-runs the full check suite, creates the GitHub
Release, and publishes to npm.
Before a release, run pnpm check:upstream (see "Documentation validation"). It
reports any provider whose package, spec or live access path has moved since the
docs were written, so a release does not ship docs that are already stale. Fix or
note what it lists; a known, tracked failure (an open issue) need not block.
# 1. Bump the version in package.json and server.json (three fields in total).
# Land it on main through a pull request, as `chore: release vX.Y.Z`.
#
# 2. Tag the merge commit and push the tag.
git tag -a v0.11.0 -m "v0.11.0"
git push origin v0.11.0The release job attaches
site-stats.json(version, chunk and provider counts, tools, per-provider chunks;pnpm site:statsprints the same) to the GitHub Release. In nutmeg-site,pnpm football-docs:releasereads it from the latest release, rebuilds the football-docs section and saves the announcement card ascards/football-docs-v<version>.png. Update its "What's new" text, thenpnpm deploythere.
Release notes come from the body of the chore: release vX.Y.Z commit, so write
that message as the release notes you want readers to see. The workflow reads it
from the second parent when the tag sits on a merge commit, strips the commit
trailers, and falls back to GitHub's generated notes if the body is empty.
Three properties worth knowing, because each one exists to stop a specific failure:
The tag must match
package.json. A mismatch fails the job before anything is created or published, so a version can never ship under another version's name.The full suite runs again. A tag can be pushed to any commit, including one that never went through a pull request, so the release path cannot assume CI already passed on that tree.
Re-running is safe. An existing Release is left alone and an already-published version is skipped, so a failed job can simply be re-run.
There is no npm token in this repository. Publishing uses npm
trusted publishing: the registry
authenticates the workflow itself over OIDC, against a trust configuration on the
package that names this repository and this workflow file. Publishing rights are
bound to release.yml rather than to a secret that would work from anywhere it
leaked to, and there is nothing to rotate.
That trust was configured once, with the npm CLI:
npm trust github football-docs \
--repo withqwerty/football-docs \
--file release.yml \
--allow-publishnpm trust list football-docs shows it; npm trust revoke removes it. Renaming
release.yml, or publishing from a different workflow, breaks the match — by
design. Re-point it with npm trust github ... --file <new-name> if the file
ever moves.
Trusted publishing generates provenance automatically, so the tarball on npm is attested to this repository and this workflow run without the workflow asking for it.
Two consequences worth knowing:
The release job runs on Node 24. Trusted publishing needs npm >= 11.5.1 and Node 22 still ships npm 10. Node 24 is in the CI matrix so that a release is not the first time the suite meets it.
The release job does not cache dependencies. This is the tree that gets published, so it is resolved fresh from the lockfile rather than rehydrated from a cache that earlier runs could have poisoned.
Note that prepublishOnly runs pnpm build && pnpm ingest, which rebuilds
data/docs.db. Publishing by hand therefore leaves that file dirty in the
working tree; the content is unchanged, only SQLite's page layout differs, so
git checkout data/docs.db clears it.
License
MIT
Available Tools
15 toolsadd_local_paperA
Add a PDF the user has (for example a paper from their library's subscription) to their football-docs library. Only PDF files are read; the text stays on this computer and is never sent anywhere. Returns a local: ID for read_paper and match_quote, which give only the outline and short passages of such papers.
| Name | Required | Description | Default |
|---|---|---|---|
| id | No | The paper's DOI or arXiv ID, to fill in its title and authors. Found in the PDF when omitted. | |
| path | Yes | The full path to the PDF, for example /Users/me/Downloads/paper.pdf or ~/Downloads/paper.pdf. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations cover the safety profile (readOnlyHint=false, destructiveHint=false, openWorldHint=true), and the description adds real context beyond them: only PDF files are read, the text stays local and is never transmitted, and it returns a local: ID whose downstream tools only expose an outline and short passages. This is meaningful behavioral disclosure not present in the structured fields.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, front-loaded with the core action, then the privacy constraint, then the return/downstream contract. No filler, though the downstream-tool sentence front-loads the return value slightly before it is needed.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description correctly compensates by describing the return value (a local: ID) and how it feeds read_paper/match_quote. Privacy, file-type limits, and downstream behavior are covered; only edge cases such as duplicate additions or metadata-failure handling are left unaddressed.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents both 'path' and the optional 'id' (DOI/arXiv ID with auto-fill behavior). The description adds a constraint (only PDF files are accepted) that bears on the path parameter, but no format or syntax detail beyond the schema, so baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: add a user's local PDF to their football-docs library, and clarifies the scope is a PDF the user already has (e.g. from a subscription). This clearly separates it from remote-fetch siblings like search_papers/get_paper, though it doesn't explicitly name a sibling to contrast against.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives a concrete usage context ('a paper from their library's subscription') and names the downstream consumers (read_paper, match_quote), which implies when the tool is useful. However, it never states when NOT to use it or how it relates to alternative ingestion paths, leaving routing to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
compare_providersARead-only
Compare what two or more providers offer for a specific data type or concept. For example: 'How do Opta and StatsBomb represent shot events differently?'
| Name | Required | Description | Default |
|---|---|---|---|
| topic | Yes | The concept to compare across providers. Examples: 'shot events', 'coordinate systems', 'xG', 'pass types' | |
| providers | No | Providers to compare. If omitted, compares all indexed providers. Use list_providers for indexed keys; common aliases such as ClubElo, football-data.co.uk, engsoccerdata, Sofascore, ESPN, StatsBomb Open Data, Opta F24, WhoScored, SkillCorner, Metrica, Sportec/DFL, TRACAB, Second Spectrum, SportRadar API, Soccer Extended, TheSportsDB, and TSDB are accepted. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations (readOnlyHint=true, destructiveHint=false) indicate safe read operation. The description adds context by specifying 'compare' and 'specific data type or concept', aligning with annotations. No behavioral traits like response format or limitations are disclosed, but annotations cover safety.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences with an example, no wasted words. Efficiently communicates purpose and usage.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool has 2 parameters and no output schema. The description is minimal and lacks any mention of return format or behavior, which is a notable gap given no output schema. It is adequate but not fully complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, with descriptions for both parameters (topic and providers) including examples. The description's example ('shot events') adds minimal extra meaning; baseline 3 is appropriate as schema does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool compares providers for a specific data type/concept, with an example. The name 'compare_providers' aligns with this purpose, and it is distinct from sibling tools like list_providers and get_provider_docs.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage for comparing representations across providers via an example, but lacks explicit when-to-use or when-not-to-use guidance, such as mention of alternative tools like search_docs for finding docs. It is clear but not exhaustive.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
forget_paperADestructiveIdempotent
Remove one paper's text from the user's football-docs library. The user's own file and Zotero are not touched.
| Name | Required | Description | Default |
|---|---|---|---|
| id | Yes | A DOI, arXiv ID, OpenAlex ID, zotero:KEY from search_papers, or local:… from add_local_paper. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already flag destructiveHint=true, idempotentHint=true and readOnlyHint=false, but the description adds genuinely useful bounds: only the library text copy is removed, while the user's own file and Zotero are untouched. That materially narrows the blast radius beyond what annotations convey, though it omits what happens on an unknown or already-removed ID.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences, zero filler, with the action and its scope boundary front-loaded. Every clause earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a single-parameter destructive tool with no output schema, the description covers the essential unknown — what is and isn't deleted — and annotations cover reversibility and safety. Minor gaps remain around failure behavior and whether confirmation is required.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Only one parameter, and schema description coverage is 100% — the schema itself enumerates the accepted ID forms (DOI, arXiv, OpenAlex, zotero:KEY, local:…). The description adds no further semantics, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Remove one paper's text') plus the target collection ('football-docs library'), and immediately scopes what is *not* affected. It distinguishes itself from source-preserving siblings, though it doesn't name a counterpart tool by name.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is only implied: remove a paper's indexed text when you no longer want it in the library. There is no explicit when-to-use versus alternatives (e.g., re-indexing, add_local_paper, search_docs), and no stated preconditions or confirmation requirement.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_paperARead-only
Look up one paper by DOI, arXiv ID, OpenAlex ID, zotero: ID or local: ID: title, authors, date, venue, all IDs, licence, open copies with their licences, abstract and a citation line. arXiv IDs come from arXiv with the paper's licence; DOIs from OpenAlex, then SportRxiv or Crossref. Use it to check that a reference exists and says what is claimed.
| Name | Required | Description | Default |
|---|---|---|---|
| id | Yes | A DOI (10.1145/3292500.3330758 or https://doi.org/...), an arXiv ID or URL (1802.07127, arxiv.org/abs/1802.07127), or an OpenAlex ID (W4288278931). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint and openWorldHint, so safety is covered. The description adds genuinely useful provenance behavior that annotations cannot convey: arXiv IDs resolve via arXiv with the paper's licence, DOIs via OpenAlex then SportRxiv or Crossref. It does not mention failure behavior (e.g. what happens if the ID is unknown) or rate limits.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, front-loaded with the core action and identifier types before the use case. The returned-field list is dense but each item is meaningful for a lookup tool; nothing is redundant with the schema.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description carries the return-value burden and does so by enumerating title, authors, dates, venue, IDs, licences, open copies, abstract and citation line. Combined with the read-only annotations and full param coverage, an agent has enough to call it correctly, though error/empty-result behavior is unstated.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% and the single parameter is fully documented with examples, so the baseline is 3. The description goes beyond the schema by naming two additional accepted ID forms (zotero: ID and local: ID) that the schema does not list, adding real value.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('look up one paper') and enumerates the exact identifiers accepted (DOI, arXiv ID, OpenAlex ID, zotero: ID, local: ID) plus the fields returned. The singular 'one paper' clearly separates it from the sibling search_papers.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives an explicit use case ('check that a reference exists and says what is claimed'), which tells the agent when this tool is the right choice. It stops short of naming an alternative or stating when not to use it, so it does not reach a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_provider_docsARead-only
Retrieve documentation for a resolved provider, optionally filtered by topic or indexed category. Use after resolve_provider_id when you know which provider to inspect and want provenance-bearing docs.
| Name | Required | Description | Default |
|---|---|---|---|
| topic | No | Optional topic to search within this provider's docs. | |
| category | No | Optional indexed category to restrict results, such as api-endpoints, qualifiers, identity-surfaces, or tracking-rendering. | |
| provider | Yes | Provider key or alias. Use resolve_provider_id first when the provider name is ambiguous. | |
| max_results | No | Maximum number of provider docs to return (default 10). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already indicate readOnlyHint=true and destructiveHint=false, so the description doesn't need to restate safety. The description adds 'provenance-bearing docs' implying reliability, but doesn't expand on behavior like error handling or data format.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences: one for core function, one for usage context. No wasted words. Front-loaded with the action and resource.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description covers purpose, usage order, and filtering. No output schema exists, but the description doesn't explain return format. However, for a straightforward retrieval tool with good annotations, it is largely complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% with descriptions for each parameter. The description reiterates filtering options (topic, category) and emphasizes 'resolved provider,' adding context but not compensating for low coverage since coverage is already high.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the verb 'retrieve' and resource 'provider docs', specifying filtering by topic or category. It distinguishes from siblings by mentioning 'resolved provider' and referencing resolve_provider_id as a prerequisite.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly says 'Use after resolve_provider_id when you know which provider to inspect and want provenance-bearing docs.' This gives a clear usage context and prerequisite. It does not explicitly state when not to use, but the context is sufficient.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_web_sourceARead-only
Read a public web page (blog post, newsletter, club or vendor article) as text, with its author, date, licence and Wayback Machine snapshots. Use it for methods first published on the web, such as Karun Singh's xT post, after finding the page with a web search. A page with no date gets the date of its earliest snapshot as an upper bound. Reads HTML and PDF. Long pages come back by section. The tool stops at bot checks and refuses local addresses.
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | The page's http(s) URL. | |
| section | No | For a long page, the section number from the outline of an earlier call. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only cover the safety profile (readOnly, openWorld, non-destructive), while the description adds substantial behavior: returns author/date/licence/snapshots, an earliest-snapshot date fallback for undated pages, HTML+PDF support, sectioned returns for long pages, stopping at bot checks, and refusing local addresses. That is far beyond what the structured metadata conveys.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with purpose, then usage, then edge-case behaviors in short declarative sentences. Slightly dense and long, but every sentence carries distinct information about the return payload or edge cases rather than filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema, so the description must carry the return burden, and it does: text plus author, date, licence, snapshots, sectioned long pages. It also documents the key failure modes (bot checks, local addresses) needed to call it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3. The description reinforces the 'section' semantics ('Long pages come back by section'), matching the schema's note that it comes from an earlier call's outline, but adds no new syntax or constraints beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Read) and resource (public web page), then concretizes it with examples (blog post, newsletter, club or vendor article). An agent knows exactly what comes back: page text plus author, date, licence and Wayback snapshots. No sibling tool overlaps this scope, so differentiation is implicit.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives a clear when-to-use rule: 'for methods first published on the web ... after finding the page with a web search,' which also sequences it after a search step. It does not state explicit exclusions or name an alternative tool, but none of the paper/provider siblings competes for this job.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_providersARead-only
List all indexed football data providers, their document count, and coverage categories. Use to understand what documentation is available. Call this first to see what providers are indexed before searching.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and destructiveHint=false, so the agent knows this is a safe read operation. Description adds that it lists providers and their metadata, which is consistent and sufficient. No additional behavioral traits needed beyond what's disclosed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, both front-loaded with purpose and usage. Every word earns its place with no redundancy or filler. Highly efficient.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given zero parameters, no output schema, and a simple list operation, the description fully covers what the tool does and how to use it. No gaps remain.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
No parameters defined (0 params), so baseline is 4 per scoring guidelines. Description does not need to add param info since schema coverage is 100% and there is nothing to describe.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Description uses specific verb 'List' and clearly identifies resource as 'all indexed football data providers' with included attributes (document count, coverage categories). This distinguishes it from sibling tools like search_docs and get_provider_docs.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly states when to use: 'Use to understand what documentation is available. Call this first to see what providers are indexed before searching.' This provides clear context and implies alternatives (e.g., search_docs for searching).
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
match_quoteARead-only
Check that a quote appears in its source: a paper (any ID read_paper takes) or a web page URL. Reports exact, normalised (same words; case, spacing, quote marks, ligatures or hyphens differ), close (with a similarity score: quote the source's own words instead) or none, with the section, page and a W3C TextQuoteSelector. Use it before citing a definition or a claim.
| Name | Required | Description | Default |
|---|---|---|---|
| quote | Yes | The quote to check, at least 10 characters. | |
| source | Yes | A paper ID or an http(s) URL. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly, non-destructive, open-world behavior, so safety is covered. The description adds substantive return semantics the annotations cannot convey: the four outcome classes (exact, normalised, close, none), the similarity score, the section/page locator, and the W3C TextQuoteSelector, plus advice to quote the source's own words on a 'close' result.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, no filler. The core purpose and accepted sources come first, with the return categories and the 'before citing' directive layered after.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema, so the description's enumeration of result types, similarity score, locator and TextQuoteSelector is exactly what an agent needs to interpret the response. Combined with usage and source semantics, nothing material is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 100% schema coverage, the baseline is 3. The description goes beyond it by expanding what 'source' accepts – 'a paper (any ID read_paper takes) or a web page URL' – tying the paper-ID namespace to a sibling tool, and by framing the quote as the thing being verified.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (check) and resource (quote against its source), and names the two source kinds it accepts. It is clearly distinguishable from siblings like get_paper or read_paper, which retrieve rather than verify.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly says to use it 'before citing a definition or a claim', giving a concrete decision point. It does not name an alternative tool or state when not to use it, so it stops short of full when/when-not guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
purge_cacheADestructiveIdempotent
Delete the user's whole football-docs paper library and the SportRxiv copy. Call with confirm: true only when the user asked for it.
| Name | Required | Description | Default |
|---|---|---|---|
| confirm | Yes | Must be true to delete. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true, readOnlyHint=false, idempotentHint=true. The description adds real value beyond that: it names exactly what is destroyed (whole library + SportRxiv copy) and adds a consent precondition. It does not discuss reversibility or downstream effects (e.g. SportRxiv side effects), hence not a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, front-loaded with the destructive scope and then the consent gate. No filler; every clause earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a destructive, irreversible-sounding single-param tool, the description covers scope, consent gating, and shares the safety burden with annotations. It leaves some ambiguity about whether the SportRxiv deletion is propagated externally, but is otherwise complete enough to call correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% and the 'confirm' parameter is already documented as 'Must be true to delete.' The description reinforces intent ('only when the user asked for it') but adds no format/syntax beyond the schema. Baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific destructive verb ('Delete') and a precisely scoped resource (the user's whole football-docs paper library and the SportRxiv copy), including the destructive scope detail that distinguishes it from sibling forget_paper.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly gives the gating condition for invocation: 'Call with confirm: true only when the user asked for it.' This is clear when-to-use guidance plus an exclusion, tying directly to the required parameter.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
read_paperARead-only
Read a paper's text. Finds an open copy (arXiv, open repositories, open-access publishers, SportRxiv) and returns it in full, by section, with its licence. For a paper the user supplied (add_local_paper or zotero:), returns only the outline and passages of at most 200 characters (FOOTBALL_DOCS_PAPERS_PASSAGE_CHARS). The text is kept in the user's library, so a second call sends no request. It never logs in anywhere and stops at bot checks; when no open copy can be read it says so and suggests add_local_paper.
| Name | Required | Description | Default |
|---|---|---|---|
| id | Yes | A DOI, arXiv ID, OpenAlex ID, zotero:KEY from search_papers, or local:… from add_local_paper. | |
| query | No | Words to find: returns passages that hold all of them, with their section and page. | |
| section | No | Section number from the outline of an earlier call (open copies only). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Goes well beyond the readOnly/openWorld annotations: discloses caching behavior ('second call sends no request'), the 200-character passage cap, the licence returned, that it never logs in and halts at bot checks, and what happens on failure. No contradiction with annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Five tight sentences, front-loaded with the primary action and scope; each sentence carries distinct operational information (modes, caching, truncation, failure routing) with no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description fully covers return shape (full text by section with licence, or outline plus passages) and error behavior, which is what an agent needs to call and interpret this tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so baseline is 3; the description adds real meaning by tying `query` to passage/section/page results and `section` to the outline from an earlier call, and by explaining the passage-length constraint applied to returned text.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Read a paper's text') and immediately distinguishes two operating modes: fetching an open copy in full versus returning only an outline with short passages for user-supplied papers. This differentiates it from metadata siblings like get_paper and search_papers.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear triggering context and explicitly routes the agent to add_local_paper when no open copy can be read, and notes the zotero:/local: identifier paths. It lacks explicit 'use X instead when Y' statements against all siblings (e.g. get_paper), so it stops short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
request_updateA
Request that a provider's documentation be added, updated, or recrawled. Use when you notice docs are outdated, a provider is missing, or you know of a better documentation source. Requests are queued for review.
| Name | Required | Description | Default |
|---|---|---|---|
| type | Yes | Type of request: new_provider (add a new tool/library), recrawl (refresh existing docs), flag_outdated (mark docs as stale), suggest_source (recommend a better doc source like llms.txt) | |
| reason | Yes | Why this update is needed. Be specific: version bump, missing event types, new API endpoints, etc. | |
| provider | Yes | Provider name (existing or proposed). Examples: 'statsbomb', 'mplsoccer', 'floodlight' | |
| suggested_urls | No | URLs for documentation sources (readthedocs, GitHub, llms.txt, etc.) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations indicate non-read-only and non-destructive. Description adds that it's a queued request, not an immediate modification, which is crucial behavioral context. No contradiction with annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, each purposeful: states action, gives usage context, and adds behavioral note. No unnecessary words, front-loaded with key information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the simplicity (4 params, no output schema), the description covers purpose, usage scenarios, and queuing behavior. Lacks detail on return value or confirmation, but sufficient for agent decision-making.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Input schema has 100% coverage with clear descriptions for each parameter. The description does not add new parameter information beyond what the schema provides, meeting baseline but not exceeding.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the verb 'Request' and the resource 'provider's documentation', specifying actions (add, update, recrawl). It distinguishes from sibling tools like get_provider_docs or search_docs which are read-only.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly states when to use: 'when you notice docs are outdated, a provider is missing, or you know of a better documentation source.' It also notes requests are queued, implying delayed action. Missing explicit when-not-to-use or alternatives among siblings.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
resolve_entityARead-only
Map a football entity (player, coach, referee, team, competition, season, stage or match) to its IDs at every provider through the Reep register: Opta, Transfermarkt, Wyscout, SkillCorner, StatsBomb, FotMob, API-Football and more. Look up by provider + id (pass namespace too, e.g. transfermarkt 'spieler', opta 'person'), by reep_id, or by name. Sources, in order: a local copy of the free register (REEP_DUCKDB_PATH), then the Reep API (REEP_API_KEY; keys are issued by hand on request, never self-service). With neither set, it returns setup steps and a DuckDB query. Keeping the local file current: Reep releases weekly. Every local answer ends with the file's release stamp checked against https://data.reep.football/releases/latest.json. If it says the file is out of date, tell the user and offer to run the curl command it gives, which downloads https://reep.football/downloads/duckdb over the same path; the next call uses the new file without a restart. Before bulk matching, make sure the file is current.
| Name | Required | Description | Default |
|---|---|---|---|
| id | No | The provider's own ID, used with provider | |
| name | No | Name to search, at least 3 characters (e.g. 'Declan Rice'). A name match is a shortlist, not an answer. | |
| type | No | Restrict results to one entity type | |
| reep_id | No | A Reep ID (e.g. 'rp1b829f1d3468c4') to list every provider ID for | |
| provider | No | Provider key for an ID lookup, e.g. 'transfermarkt', 'opta', 'wyscout', 'skillcorner', 'fotmob' | |
| namespace | No | The provider's namespace for that ID, e.g. 'spieler' or 'verein' for Transfermarkt, 'person' or 'team' for Opta, 'player' for Wyscout. Recommended: some providers reuse numbers across entity types. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only cover the read-only, open-world, non-destructive profile, but the description adds substantial behavior: the source fallback order (local DuckDB, then API), the exact degraded-mode return (setup steps plus a DuckDB query) when unconfigured, API key issuance policy, the release-stamp check against a live URL, the offer to run a curl download, and that the next call picks up the new file without a restart.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Purpose is front-loaded in the first sentence and the mode/precedence information follows logically. The tail about weekly releases, the release-stamp URL, the curl command and the download URL is operational detail that is useful but dense; it could be tightened without losing meaning.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a six-parameter, all-optional resolver with no output schema, the description covers inputs, sources, auth, degraded mode and freshness workflow well. The main gap is that it never sketches the shape of a successful result (e.g. one ID per provider), so an agent knows how to call it but not quite what to expect back.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is already 100%, so a 3 is the baseline, but the description adds value by explaining how parameters combine — 'look up by provider + id (pass namespace too, e.g. transfermarkt \'spieler\', opta \'person\')' — and by clarifying that a name match is a shortlist rather than an answer. It does not document the type enum or the id/reep_id fields beyond what the schema already says.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific verb ('Map... to its IDs') and resource (football entities in the Reep register), enumerates the entity types and target providers, and lists the three lookup modes. It is clearly distinguishable from siblings like resolve_provider_id because it scopes itself to cross-provider ID resolution via the Reep register.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives concrete conditional guidance: look up by provider+id (with namespace), by reep_id, or by name; and it states prerequisites such as REEP_DUCKDB_PATH / REEP_API_KEY, that keys are issued by hand, and that the file should be current before bulk matching. What it never does is name a sibling alternative (e.g. resolve_provider_id) or state when NOT to use this tool, so it falls short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
resolve_provider_idARead-only
Resolve a football data provider name or alias to the canonical football-docs provider key before searching. Use when users mention brands, vendors, products, or aliases such as Stats Perform, Opta F24, Hudl Wyscout, Second Spectrum, FMDB, Transfer Room, FBref, Sofascore, or TheSportsDB.
| Name | Required | Description | Default |
|---|---|---|---|
| query | Yes | Provider name, brand, product, or alias to resolve to a canonical provider key. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and destructiveHint=false, so the description's additional behavioral context is minimal. It adds that resolution maps to a canonical key, which is consistent. No contradictions, but no extra behavioral details like error handling or idempotency.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences: first states the action concisely, second lists usage examples. No wasted words. Front-loaded with the core purpose.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's simplicity (one param, no output schema, read-only), the description fully covers what an agent needs: purpose, when to use, and input examples. No gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% and the schema description already defines the query parameter. The tool description adds concrete examples (e.g., Stats Perform, Opta F24) that help clarify acceptable input values, enhancing semantic understanding beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose: resolving a provider name/alias to a canonical key. It specifies the verb 'resolve', the resource (provider name to key), and lists numerous examples, leaving no ambiguity.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly says 'Use when users mention brands, vendors, products, or aliases' with specific examples. Implies it's a preprocessing step for searching. Lacks explicit exclusion criteria but provides clear guidance on when to invoke.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
search_docsARead-only
Search football data provider documentation. Use for finding event types, qualifier IDs, API endpoints, coordinate systems, data models, and cross-provider mappings. Returns the most relevant documentation chunks. Results that do not contain every query term are marked "partial", and the reply names any query term that no indexed doc mentions: if the question is about that term, it is not indexed. It also says when the query names a provider that is not indexed, and why.
| Name | Required | Description | Default |
|---|---|---|---|
| query | Yes | Search query. Examples: 'Opta goal qualifier', 'StatsBomb shot event type', 'coordinate system differences', 'xG qualifier ID', 'SportMonks fixture endpoint', 'FMDB Pro players endpoint' | |
| provider | No | Optional provider filter. Use list_providers for indexed provider keys. Common aliases such as fbref, understat, ClubElo, football-data.co.uk, engsoccerdata, Sofascore, ESPN, FMDB, TransferRoom, Hudl Wyscout, Stats Perform, Opta F24, WhoScored, Metrica, Sportec/DFL, TRACAB, Second Spectrum, SportRadar API, Soccer Extended, TheSportsDB, and TSDB are accepted. | |
| max_results | No | Maximum number of results to return (default 10) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover the safety profile (readOnly, non-destructive, closed-world), and the description adds real behavioral context beyond them: partial-result marking, disclosure of unmatched query terms, and an explicit explanation when a provider is not indexed. It does not address ranking depth or result-count behavior beyond max_results, so it is strong but not exhaustive.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core purpose in the first sentence and then the use cases, which is the right ordering. The middle clause about partial results, unnamed terms, and unindexed providers is dense and slightly repetitive, but each clause carries actionable information for interpreting results.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description does the work of explaining return semantics (most relevant chunks, partial flagging, diagnostic messages), which is exactly what an agent needs here. Remaining gaps are minor: how results are ranked and whether there is any result-count ceiling.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all three parameters are already documented in the schema, including query examples and the provider filter. The description only indirectly supports the 'provider' parameter via its note about unindexed providers and adds no syntax or format detail beyond the schema. Baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Search football data provider documentation') and then enumerates the exact content types it retrieves (event types, qualifier IDs, endpoints, coordinate systems, data models, cross-provider mappings). That enumeration separates it cleanly from siblings like get_provider_docs and list_providers without needing to name them.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
"Use for finding event types, qualifier IDs, API endpoints..." gives clear positive context for when to invoke it, which is more than most search tools offer. It stops short of explicit when-not guidance or naming the alternative (e.g., get_provider_docs when the provider/topic is already known), so it misses the top band.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
search_papersARead-only
Search scholarly papers on football analytics and sport science: OpenAlex (title, abstract and full text), arXiv (title, abstract, authors) and SportRxiv (title, abstract, keywords). Use it to find the paper behind a method (xG, VAEP, EPV, pitch control) or the works that cite an idea. Use words and "quoted phrases"; OpenAlex matches full text, so a hit may cite the idea rather than introduce it. Many methods first appeared in blog posts or conference papers without a DOI (xT, for example): search the web for those and read them with get_web_source. The reply names the services asked. FOOTBALL_DOCS_PAPERS=off turns paper lookups off.
| Name | Required | Description | Default |
|---|---|---|---|
| query | Yes | Words and "quoted phrases", optionally with AND / OR / NOT. Examples: '"expected threat" soccer', '"pitch control" Spearman', 'VAEP action values' | |
| sources | No | Sources to ask. Default: openalex, arxiv and sportrxiv (searched in a local copy of its feed). Add zotero to search the user's own Zotero library (Zotero on this computer, else the Zotero web API with ZOTERO_API_KEY). | |
| year_to | No | Only papers published in or before this year. | |
| year_from | No | Only papers published in or after this year. | |
| max_results | No | Results per source, 1 to 25 (default 10). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover the safety profile (readOnly, openWorld, non-destructive), so the bar is lower. The description still adds real context: OpenAlex matches full text so a hit may merely cite the idea, the reply names which services were queried, and FOOTBALL_DOCS_PAPERS=off disables lookups entirely.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with purpose and sources, then usage, then caveats. Every sentence carries information, though the block is dense and the full-text caveat is woven in mid-paragraph rather than isolated.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 5-parameter, read-only, no-output-schema tool the description covers sources, matching behavior, routing to siblings and an environment gate. It only lightly touches what the reply looks like beyond naming the services queried, but nothing essential for correct invocation is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents query, sources, year_from/year_to and max_results with examples and defaults. The description's query-syntax hints (quoted phrases, AND/OR/NOT) largely restate the schema, and only the full-text matching nuance goes beyond it. Baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (search) and resource (scholarly papers) with the exact scope: OpenAlex, arXiv and SportRxiv, and which fields each covers (title, abstract, full text). An agent can distinguish it from get_paper, read_paper and search_docs without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives concrete use cases (find the paper behind xG, VAEP, EPV, pitch control, or works that cite an idea) and an explicit alternative path for material without a DOI: search the web and read it with get_web_source. Both when-to-use and when-to-route-elsewhere are stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
8 tool updates
v0.16.2- Added
add_local_paper - Added
forget_paper - Added
get_paper - Added
get_web_source - Added
match_quote - Added
purge_cache - Added
read_paper - Added
search_papers
1 tool update
v0.12.2- Changed
resolve_entity8 fields changed- changed
Input schema / properties / id / descriptionPrevious value: -"ID from the source provider to resolve to all other IDs"New value: +"The provider's own ID, used with provider" - changed
Input schema / properties / name / descriptionPrevious value: -"Entity name to search for (e.g. 'Cole Palmer', 'Arsenal'). Fuzzy match on name and aliases."New value: +"Name to search, at least 3 characters (e.g. 'Declan Rice'). A name match is a shortlist, not an answer." - added
Input schema / properties / namespaceAdded value: +{ + "description": "The provider's namespace for that ID, e.g. 'spieler' or 'verein' for Transfermarkt, 'person' or 'team' for Opta, 'player' for Wyscout. Recommended: some providers reuse numbers across entity types.", + "type": "string" +} - changed
Input schema / properties / provider / descriptionPrevious value: -"Source provider for ID resolution (e.g. 'transfermarkt', 'fbref', 'sofascore', 'opta', 'soccerway')"New value: +"Provider key for an ID lookup, e.g. 'transfermarkt', 'opta', 'wyscout', 'skillcorner', 'fotmob'" - removed
Input schema / properties / qidRemoved value: -{ - "description": "Wikidata QID for direct lookup (e.g. 'Q99760796')", - "type": "string" -} - added
Input schema / properties / reep_idAdded value: +{ + "description": "A Reep ID (e.g. 'rp1b829f1d3468c4') to list every provider ID for", + "type": "string" +} - changed
Input schema / properties / type / descriptionPrevious value: -"Filter results by entity type"New value: +"Restrict results to one entity type" - changed
Input schema / properties / type / enumPrevious value: -[ - "player", - "team", - "coach" -]New value: +[ + "player", + "coach", + "referee", + "team", + "competition", + "season", + "stage", + "match" +]
7 tool updates
v0.6.4- First observed
compare_providers - First observed
get_provider_docs - First observed
list_providers - First observed
request_update - First observed
resolve_entity - First observed
resolve_provider_id - First observed
search_docs
TDQS
Scored across 15 tools
Most tools have clearly distinct purposes: paper search/read/quote tools, provider documentation tools, and entity/provider resolution tools are separated by their descriptions. Minor overlap exists between search_docs and get_provider_docs, and between list_providers and resolve_provider_id, but the descriptions provide enough context to select correctly.
All tool names use consistent snake_case and follow a predictable verb_noun or verb_object pattern (e.g., search_papers, get_paper, read_paper, resolve_entity, compare_providers). Even less conventional names like purge_cache and forget_paper fit the same structural convention, making the set easy to scan.
The 15 tools are reasonable for a documentation and research server spanning papers, provider docs, web sources, local PDFs, and entity resolution. The count sits at the upper end of the comfortable range but does not feel excessive given the breadth of supported workflows.
The surface covers the core research and documentation lifecycle well: searching, retrieving, reading, quoting, adding local papers, resolving entities/providers, and requesting doc updates. A minor gap is the lack of a tool to list or search the user's own saved local paper library, but other tools mostly compensate through local IDs and deletion/purge operations.
Maintenance
Related MCP Connectors
Versioned documentation registry and semantic search for AI tools and coding assistants.
Football fixtures, standings, and odds intelligence for AI agents.
Provide your AI coding tools with token-efficient access to up-to-date technical documentation for…
Public social-data API and live docs for AI coding agents.
Related MCP Servers
- FlicenseNot gradedqualityDmaintenanceProvides AI agents with professional coding standards, development best practices, and context-aware guidance through static documentation and AI-powered custom recommendations. Enables agents to access comprehensive development guidelines including coding rules, debugging techniques, and AI steering instructions.-
- AlicenseNot gradedqualityDmaintenanceProvides AI assistants with searchable access to documentation from 170+ curated repositories and 1000+ popular GitHub projects across 20+ categories including trading, AI/ML, DevOps, and web development.3MIT
- FlicenseNot gradedqualityDmaintenanceProvides intelligent access to OPTA football API documentation with automatic authentication, allowing users to query specific endpoint documentation and get answers about soccer match data, events, fixtures, and statistics through natural language.7-
- AlicenseNot gradedqualityCmaintenanceProvides real-time, up-to-date documentation for major LLM providers (OpenAI, Anthropic, Google Gemini) to prevent hallucinations and outdated code patterns in AI agents.9 npm6MIT