Wikipedia MCP
This server gives MCP clients read-only access to Wikipedia and Wikimedia Commons content, metadata, and community signals across 10 language editions — no API key needed.
Find articles:
searchfor ranked matches by query;disambiguationto resolve ambiguous titles like 'Mercury' or 'Apple'.Read content:
summary(gist + thumbnail),article_extract(full plain text),article_sections+section_text(targeted section reading),simple_summary(Simple English),infobox(structured fact table), andrandomfor surprise topics.Daily/curated feeds:
featured_article,picture_of_the_day,media_of_the_day,on_this_day,deaths_on_this_day,births_on_this_day,news,did_you_know,dino_fact,quote.Discovery & navigation:
categories/category_members(taxonomy),links/backlinks/external_links(reference network),related_articles(semantic similarity),nearby(geographic),translations(language editions).Media:
image(lead image URLs),media_list(all media in an article),media_search(keyword search of Commons).Analytics & trends:
pageviews(daily views),top_reads(most-read articles by date),recent_changes(live edit stream).Provenance & trust signals:
revisionsandrevision_diff(edit history and diffs),contributors/user_contribs(who edits what),article_quality(WikiProject grades),article_flags,article_protection,article_pulse,talk(discussion pages),references/citation_sources/citation_needed(source audits and unsourced claims).Language support: most tools accept a
langparameter (en, de, es, fr, ja, zh, pt, it, ru, nl);media_searchis language-independent andsimple_summaryis always Simple English.
Provides a tool to retrieve Wikimedia Commons' Picture of the Day for today or a specified date.
Provides tools for accessing Wikipedia content via its free REST API, including article search, summaries, full extracts, section listings, categories, links/backlinks, external links, nearby articles, translations, revision history, pageviews, top reads, random articles, featured articles, quotes, on-this-day events, notable deaths, current news, recent changes, and media inventories, with support for multiple languages.
Wikipedia MCP
A Model Context Protocol (MCP) server that provides access to Wikipedia via the free REST API. No API key required.
⭐ If you find this useful, please star the repo — it helps others discover it.
Contributing
Contributions are welcome — new tools especially. See CONTRIBUTING.md for the how-to: one file, one tool per PR, smoke tests included.
Related MCP server: wikipedia-recent-changes-mcp
Tools
Tool | Description |
| Search Wikipedia for articles matching a query |
| Get a Wikipedia article summary + thumbnail by title |
| Get a random Wikipedia article summary |
| Get the Simple English Wikipedia version of a topic — plain-language explanation |
| Get a random "Did You Know" style fact |
| Get a dino/prehistory-specific fact (specific species or random) |
| Get today's Wikipedia Featured Article |
| Get Wikimedia Commons' Picture of the Day (today, or a YYYYMMDD date) |
| Get Wikimedia Commons' Media of the Day — the curated daily video/audio clip (today, or a YYYYMMDD date) |
| Get a full plain-text extract of an article (longer than |
| Get the table of contents (section headings) for an article — useful for navigating long articles before reading the full body |
| Read one section of an article as plain text — by number, hierarchical number (e.g. |
| Get historical events that happened on today's date |
| Get notable deaths that happened on today's date (companion to |
| Get notable births that happened on today's date (companion to |
| List Wikipedia categories an article belongs to |
| List outgoing Wikipedia links from an article (the article's reference network) |
| List incoming Wikipedia links to an article (what links here / referrer pages — inverse of |
| List external (off-wiki) links from an article — citations, references, and primary sources the article points to (outbound complement to |
| List Wikipedia articles geographically near a location — anchor by article title (e.g. 'Eiffel Tower') or lat/lon, with distances |
| List all language versions of an article (langlinks) — discover what languages it exists in |
| Show an article's recent edit history (who edited it, when, edit summaries, size deltas) with diff links |
| Get daily view counts for an article (popularity, trending, historical interest) |
| Get current events from Wikipedia's Main Page "In the news" section |
| Get the most-read articles on Wikipedia for a given date |
| Get just the lead image (thumbnail + original URLs) for an article, no summary text |
| List all media (images, videos, audio) used in an article — full inventory with type, caption, and thumbnail |
| Search Wikimedia Commons for freely-licensed media by keyword (topic-based discovery — |
| Get a random notable quote from a curated list of famous authors |
| Live window into Wikipedia right now — most recent edits, with kind filter ('all', 'edit', 'new', 'categorize', 'log') for breaking-news edits, newly published articles, and more |
| List articles filed under a category — taxonomy-based discovery, the reverse of |
| Extract an article's structured fact box (infobox) as a field/value table — dates, people, places, statistics; the fastest path to a concrete fact without reading prose |
| Wikipedia's quality assessments for an article — WikiProject grades (FA, GA, B, C, Start, Stub) + importance ratings; the encyclopedia's own trust signal before relying on an article |
| Find articles semantically similar to a given article ("what should I read next") — Wikipedia's own MoreLikeThis search ranking, each with short description + thumbnail |
| Who writes and maintains an article — most active recent editors ranked by edit count (up to 500 sampled revisions), with user-page links + anonymous (IP) edit share; the provenance companion to |
| The sources an article cites — its bibliography: each citation's text plus the off-wiki URLs it points to (DOI, publisher, archive, primary-source links); the verification companion to |
| Compare two revisions of an article — plain-text unified diff of what a specific edit changed, each side labelled with timestamp, editor, and edit summary |
| Resolve a disambiguation page into its candidate articles — title + one-line description, grouped by section; pick the right one, then fetch it |
| What a Wikipedia editor has been doing — latest contributions by a username or IP (timestamp, byte delta, edit comment, new-page/minor/current flags), with registration date + total edit count; profile contributors or audit anonymous IPs |
| Find statements Wikipedia has flagged as needing a source — pass |
| Read an article's talk page — the most recently active editor discussion threads, each with heading, last-activity timestamp, signed-comment count, and an excerpt of the latest comment; automated notices filtered out — the behind-the-scenes companion to |
| Show the maintenance banners editors placed on an article — {{POV}}, {{Original research}}, {{Unreferenced}}, {{Cleanup}}, {{Disputed}} and more, each with tag date, location (article top or section), and a plain-language meaning; the trust-signal companion to |
| Is this Wikipedia article locked — which actions are restricted (editing, moving/renaming, creating), at what level (semi-protection, extended-confirmed, administrators-only), and when each restriction expires; follows redirects; the lockdown companion to the other trust signals |
| Vital signs of an article — creation date + creator, page length, watcher count, most recent edit, and edit velocity over the last 30 days, with a plain-language activity verdict (buzzing / active / quiet / dormant); the liveliness companion to |
| Which publishers an article's evidence comes from — its references aggregated by source domain into a ranked publisher map with per-domain share, plus archived-copy and DOI-linked scholarly shares and a plain-language verdict (diverse / balanced / top-heavy / lopsided / thin); the diversity companion to |
All tools accept an optional lang parameter (one of: en, de, es, fr, ja, zh, pt, it, ru, nl), except media_search — Wikimedia Commons is language-independent, so it takes no lang — and simple_summary, which always reads Simple English Wikipedia. Note: quote accepts the parameter for API consistency but is currently English-only (curated list).
Setup
This is a standard MCP server — it works with any MCP client (Claude Code, Cursor, Windsurf, etc.).
Any MCP client
Add to your client's MCP config (e.g. ~/.claude.json, or Cursor's mcp.json):
{
"mcpServers": {
"wikipedia": {
"command": "python3",
"args": ["/path/to/wikipedia-mcp/src/server.py"]
}
}
}Then restart your client (or reload its MCP servers) to pick it up.
OpenClaw (via mcporter)
Add the same mcpServers block to ~/.openclaw/workspace/config/mcporter.json, then restart the gateway:
openclaw gateway restartUsage
# Search
mcporter call wikipedia search --args '{"query": "velociraptor", "limit": 5}'
# Article summary
mcporter call wikipedia summary --args '{"title": "Tyrannosaurus"}'
# Random article
mcporter call wikipedia random
# Simple English explanation ("explain it simply")
mcporter call wikipedia simple_summary --args '{"title": "Photosynthesis"}'
# Random dino fact
mcporter call wikipedia dino_fact
# Specific species
mcporter call wikipedia dino_fact --args '{"species": "Spinosaurus"}'
# Today's featured article
mcporter call wikipedia featured_article
# Picture of the Day (curated daily image from Wikimedia Commons)
mcporter call wikipedia picture_of_the_day
mcporter call wikipedia picture_of_the_day --args '{"date": "20260901"}'
# Media of the Day (curated daily video/audio clip from Wikimedia Commons)
mcporter call wikipedia media_of_the_day
mcporter call wikipedia media_of_the_day --args '{"date": "20260830"}'
# Full plain-text article extract (vs summary)
mcporter call wikipedia article_extract --args '{"title": "Tyrannosaurus"}'
# Article table of contents — section headings (navigate before reading full body)
mcporter call wikipedia article_sections --args '{"title": "Tyrannosaurus"}'
# Read one section of an article (by number, "2.1", or heading name — 0 = lead/intro)
mcporter call wikipedia section_text --args '{"title": "Tyrannosaurus", "section": "Description"}'
mcporter call wikipedia section_text --args '{"title": "Tyrannosaurus", "section": 3}'
# On this day (historical events for today)
mcporter call wikipedia on_this_day
mcporter call wikipedia on_this_day --args '{"count": 8}'
# Deaths on this day (notable deaths for today — in memoriam content hooks)
mcporter call wikipedia deaths_on_this_day
mcporter call wikipedia deaths_on_this_day --args '{"count": 6}'
# Births on this day (notable births for today — "born on this day" content hooks)
mcporter call wikipedia births_on_this_day
mcporter call wikipedia births_on_this_day --args '{"count": 6}'
# Categories for an article (taxonomy-based discovery)
mcporter call wikipedia categories --args '{"title": "Tyrannosaurus"}'
mcporter call wikipedia categories --args '{"title": "Tyrannosaurus", "limit": 10}'
# Outgoing links from an article (graph-style discovery)
mcporter call wikipedia links --args '{"title": "Tyrannosaurus"}'
mcporter call wikipedia links --args '{"title": "Tyrannosaurus", "limit": 30}'
# Translations — list all language editions of an article
mcporter call wikipedia translations --args '{"title": "Tyrannosaurus"}'
mcporter call wikipedia translations --args '{"title": "Tyrannosaurus", "limit": 10}'
# Revision history — recent edits, editors, summaries, diffs
mcporter call wikipedia revisions --args '{"title": "Tyrannosaurus"}'
mcporter call wikipedia revisions --args '{"title": "Tyrannosaurus", "limit": 20}'
# Daily view counts (popularity research, trending topics)
mcporter call wikipedia pageviews --args '{"title": "Tyrannosaurus"}'
mcporter call wikipedia pageviews --args '{"title": "Python_(programming_language)", "start": "20250101", "end": "20250107"}'
# Current events from Wikipedia's Main Page (today's "In the news")
mcporter call wikipedia news
mcporter call wikipedia news --args '{"limit": 8}'
# Top reads — most-viewed articles on a given date
mcporter call wikipedia top_reads
mcporter call wikipedia top_reads --args '{"date": "20260101", "limit": 15}'
# Lead image — thumbnail + original URLs for an article (no text)
mcporter call wikipedia image --args '{"title": "Tyrannosaurus"}'
mcporter call wikipedia image --args '{"title": "Tyrannosaurus", "lang": "de"}'
# Media inventory — all images/videos/audio in an article (full list, not just lead)
mcporter call wikipedia media_list --args '{"title": "Tyrannosaurus"}'
mcporter call wikipedia media_list --args '{"title": "Tyrannosaurus", "limit": 50}'
mcporter call wikipedia media_list --args '{"title": "Berlin", "lang": "de"}'
# Media search — freely-licensed Commons media by keyword (topic-based, not article-based)
mcporter call wikipedia media_search --args '{"query": "aurora borealis"}'
mcporter call wikipedia media_search --args '{"query": "volcano eruption", "filetype": "video", "limit": 5}'
# Related articles — semantically similar articles ("what should I read next")
mcporter call wikipedia related_articles --args '{"title": "Velociraptor"}'
mcporter call wikipedia related_articles --args '{"title": "Velociraptor", "limit": 10}'
# Contributors — who writes and maintains an article
mcporter call wikipedia contributors --args '{"title": "Albert Einstein"}'
mcporter call wikipedia contributors --args '{"title": "Velociraptor", "limit": 5}'
# References — the sources an article cites (its bibliography)
mcporter call wikipedia references --args '{"title": "Albert Einstein"}'
mcporter call wikipedia references --args '{"title": "Velociraptor", "limit": 10}'
# Random notable quote (curated list of famous authors)
mcporter call wikipedia quote
mcporter call wikipedia quote --args '{"lang": "de"}' # lang accepted, currently English-only
# Non-English Wikipedia
mcporter call wikipedia summary --args '{"title": "Berlin", "lang": "de"}'Requirements
Python 3.10+
requests>=2.28.0
API
Uses Wikipedia's free REST API:
Search: MediaWiki Action API (
/w/api.php)Summary / Random / Featured / Media-list: REST API v1 (
/api/rest_v1/...)Pageviews / Top reads: Wikimedia cross-wiki metrics API (
https://wikimedia.org/api/rest_v1/metrics/pageviews/...)Related articles: MediaWiki Action API (search generator with
morelike:scoring)
No API key required. Respects Wikipedia's User-Agent policy.
Development
Run the smoke tests:
python3 tests/test_server.pyLicense
MIT
Available Tools
40 toolsarticle_extractA
Get a Wikipedia article's full plain-text extract by title — much longer than summary (typically several paragraphs). Returns plain text (no HTML). Complements summary: use it when the summary is too brief and you want a fuller reading.
| Name | Required | Description | Default |
|---|---|---|---|
| lang | No | Wikipedia language code (default 'en') | en |
| title | Yes | Article title (e.g. 'Tyrannosaurus' or 'Albert_Einstein') |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations and no output schema, the description carries the full disclosure burden. It usefully discloses the return format ('plain text (no HTML)') and rough size ('typically several paragraphs'), which is real behavioral value. It does not cover failure behavior for missing titles, disambiguation pages, redirects, or truncation limits, so it is only partially transparent.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, zero filler, with the primary purpose and the sibling contrast front-loaded. Every clause earns its place and nothing is repeated from the schema.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 2-parameter read tool with no annotations or output schema, the description covers purpose, output format, and size — enough to call it correctly. Edge-case behavior (nonexistent titles, disambiguation, language fallback) is unaddressed, which keeps it short of fully complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% — both `lang` (enum plus default) and `title` (with examples) are fully documented in the schema. The description adds nothing parameter-specific beyond what the schema already provides, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Get a Wikipedia article's full plain-text extract by title') and immediately distinguishes itself from the `summary` sibling by contrast in length. An agent can tell exactly what this returns and how it differs from the other extract-like tools without opening a schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly names the alternative ('Complements `summary`: use it when the summary is too brief') and gives the selecting condition. The only omission is a when-not case (e.g. use `summary` when you need one sentence), which is easily inferred and does not undermine the routing.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
article_qualityA
Get Wikipedia's quality assessments for an article — the WikiProject grades (FA, GA, B, C, Start, Stub) and importance ratings assigned by editors. The encyclopedia's own trust signal: use it before relying on an article (a GA/FA passed formal review; a Stub is a skeleton). Reports an overall class plus the per-project breakdown with a quality-ladder legend. Assessment is enabled per language edition (en works; some editions like de report no data). Read-only via the pageassessments action API — GET only, no new dependencies.
| Name | Required | Description | Default |
|---|---|---|---|
| lang | No | Wikipedia language code (default 'en') | en |
| title | Yes | Article title (e.g. 'Albert Einstein' or 'Paris') |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations provided, so description carries the burden. It discloses read-only via pageassessments API, GET only, no new dependencies, and outlines the output structure (overall class plus per-project breakdown with legend). Lacks error handling details, but coverage is good.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Concise yet information-dense: purpose, usage guidance, output description, and technical details are each conveyed in a single sentence. Front-loaded with the core purpose and trust signal, no wasted words.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Complete for a simple tool with two well-documented params. Explains output (overall class + breakdown + legend), read-only nature, and language caveats. An agent has everything needed to call it correctly without an output schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% with descriptions for both lang and title. Description adds the language caveat (de reports no data) and explains the significance of grades, which helps agents interpret results. This exceeds the baseline 3.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Clearly states it retrieves Wikipedia quality assessments (WikiProject grades and importance) for an article. Distinguishes itself from content-focused siblings (e.g., summary, article_extract) by emphasizing it is a trust signal about article quality.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly recommends using it before relying on an article and notes language-specific data availability (en works, de may report no data). Does not name specific alternative tools, but the sibling list makes it obvious when to use this tool over content retrieval tools.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
article_sectionsA
Get the table of contents (section headings) for a Wikipedia article — section number, heading text, and nesting level. Useful for navigating long articles before committing to the full body via article_extract. Major articles can have 50KB+ of body text; article_sections gives the TOC in a compact numbered list so callers can pick what to read next. Pairs with summary (lead), article_sections (structure), article_extract (full body).
| Name | Required | Description | Default |
|---|---|---|---|
| lang | No | Wikipedia language code (default 'en') | en |
| title | Yes | Article title (e.g. 'Tyrannosaurus' or 'Albert_Einstein') |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations were provided, so the description carries the full burden. It discloses that the output is a compact numbered list and notes that major articles can have 50KB+ of body text, giving useful context about the cost of alternatives. However, it does not state whether the tool is read-only (implied by 'Get') or mention any pagination, rate limits, or error cases. This is a moderate gap for a tool with zero annotation coverage.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, each front-loaded with purpose: first defines what it gets, second explains why (use case), third maps the tool family. Zero waste, and the most important information (what it returns) is first.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's simplicity (2 params, no output schema, no annotations) and the complexity of the sibling ecosystem, the description is complete: it covers purpose, usage context, return format, and pairing. Nothing an agent needs to call it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the schema already documents both 'lang' (with enum and default) and 'title' (with examples). The description does not add any parameter-specific meaning beyond what the schema provides. Baseline 3 is appropriate when the schema does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Get the table of contents (section headings) for a Wikipedia article') and immediately specifies the returned fields (section number, heading text, nesting level). It clearly distinguishes itself from siblings like article_extract and summary by naming them in the pairing sentence.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly describes when to use this tool: 'before committing to the full body via article_extract' and 'for navigating long articles.' It names alternatives (article_extract, summary) and explains the workflow, making the selection condition unambiguous.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
backlinksA
List incoming Wikipedia links to an article — i.e. 'what links here' / backlinks / referrer pages. Inverse of links: given 'Velociraptor', see which other articles reference it (cultural mentions, scientific citations, comparative anatomy pages, etc.). Useful for graph-style discovery in the reverse direction — mapping an article's position in the encyclopedia network. Filters to main namespace so talk/user/etc. don't pollute the result.
| Name | Required | Description | Default |
|---|---|---|---|
| lang | No | Wikipedia language code (default 'en') | en |
| limit | No | Max backlinks to return (default 20, max 50) | |
| title | Yes | Article title (e.g. 'Tyrannosaurus' or 'Albert_Einstein') |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the burden. It usefully discloses the main-namespace filtering constraint ('talk/user/etc. don't pollute the result'), which is genuine behavioral context. However, it says nothing about result ordering, pagination beyond the limit param, or rate limits.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the core definition ('List incoming Wikipedia links to an article') before elaborating with examples and the namespace filter. Slightly verbose with the parenthetical examples, but each sentence adds routing or scope value.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no annotations and no output schema, the description needs to stand alone, and it covers purpose, the inverse relationship to `links`, use case, and result scoping. It omits return-format/ordering details, which is a minor gap for a simple list tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so lang, limit, and title are already documented in the schema. The description adds no parameter syntax or format detail beyond that, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (list) and resource (incoming Wikipedia links/backlinks), and explicitly names the sibling it inverts: 'Inverse of `links`'. An agent can distinguish it from external_links, links, and search without opening a schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear usage context ('graph-style discovery in the reverse direction — mapping an article's position in the encyclopedia network') and the inverse-relationship cue to `links`. It stops short of explicit when-not conditions or naming a preferred alternative for overlapping needs, so it lands just under a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
births_on_this_dayA
Get notable births that happened on today's date (UTC) from Wikipedia's 'On This Day' feed — the births companion to on_this_day (events) and deaths_on_this_day (deaths). Useful for 'born on this day' content hooks, birthday round-ups, and newsletter intros. Pairs with on_this_day (events) and featured_article (today's long-form pick) for a full daily 'today in Wikipedia' digest.
| Name | Required | Description | Default |
|---|---|---|---|
| lang | No | Wikipedia language code (default 'en') | en |
| count | No | Number of births to return (default 5, max 10) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It usefully discloses the data source (Wikipedia 'On This Day'), the temporal scope (today's date, UTC), and the content type (notable births). However, it does not mention behavior around empty results, count handling beyond the schema, or whether the API always uses the current date rather than a user-supplied date.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with the core action and scope, followed by relevant use cases and sibling integration. It is slightly generous with marketing-style phrases, but each sentence earns its place by clarifying purpose or usage.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's low complexity, two optional parameters, and 100% schema coverage, the description is largely complete. It explains source, scope, use cases, and relationship to siblings. The lack of an output schema means return shape is not specified, but 'notable births' plus a count parameter gives enough context for most agents to infer a list of birth entries.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so both parameters (lang and count) are already fully documented with defaults and constraints. The description adds no additional parameter-level meaning beyond the overall resource context, so the baseline score of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'Get notable births that happened on today's date (UTC) from Wikipedia's On This Day feed.' It also explicitly differentiates this tool from its siblings by calling it the 'births companion' to on_this_day (events) and deaths_on_this_day (deaths), so an agent can immediately understand its unique role.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description names the sibling tools and clarifies which one handles births vs events vs deaths, giving clear selection context. It also provides concrete use cases like birthday round-ups and newsletter intros, though it does not explicitly state when not to use this tool or list exclusion criteria.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
categoriesA
List Wikipedia categories an article belongs to. Useful for taxonomy-based discovery — finding related topics that don't appear in text search. Hidden/maintenance categories are filtered out.
| Name | Required | Description | Default |
|---|---|---|---|
| lang | No | Wikipedia language code (default 'en') | en |
| limit | No | Max categories to return (default 20, max 50) | |
| title | Yes | Article title (e.g. 'Tyrannosaurus' or 'Albert_Einstein') |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description must carry the full behavioral burden. It discloses that hidden/maintenance categories are filtered out, which is useful. However, it omits other behavioral details like whether results are ordered, pagination behavior, or whether it requires authentication.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two concise sentences, front-loaded with the core purpose and followed by a use-case and a filtering note. Every sentence earns its place without redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given no output schema exists, the description adequately explains what the tool returns (categories) and the filtering behavior. It doesn't cover return format or pagination details, but overall provides sufficient context for an agent to invoke it correctly, though minor gaps remain.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, meaning all parameters are already documented in the schema with descriptions and defaults. The description adds no parameter-specific information beyond what the schema provides, so the baseline score of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (List) and resource (Wikipedia categories an article belongs to) with clear scope. The description distinguishes it from text-search siblings by describing it as taxonomy-based discovery, making the purpose immediately identifiable among tools like search or links.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides clear context for when to use the tool: 'taxonomy-based discovery — finding related topics that don't appear in text search.' This implicitly contrasts with search but doesn't name specific sibling tools or state explicit exclusions or prerequisites.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
category_membersA
List Wikipedia articles filed under a category — taxonomy-based discovery. The reverse direction of categories (which lists an article's categories): given 'Machine learning researchers', enumerate who's actually in it. Each entry includes a 1-2 sentence extract plus a thumbnail URL when one exists, so results are browsable at a glance. Filters to main-namespace articles so subcategories and files don't pollute the result. The 'Category:' prefix is optional. Pairs with categories (find the taxonomy) and search (find the entry point).
| Name | Required | Description | Default |
|---|---|---|---|
| lang | No | Wikipedia language code (default 'en') | en |
| limit | No | Max articles to return (default 20, max 50) | |
| category | Yes | Category name, with or without the 'Category:' prefix (e.g. 'Flightless birds' or 'Category:Flightless birds') |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the behavioral disclosure burden. It includes several meaningful behaviors: it returns a 1-2 sentence extract and a thumbnail URL per entry, filters out subcategories and files by limiting to main namespace, and treats the 'Category:' prefix as optional. It does not cover pagination, ordering, or error behavior, but these are minor for a list-read tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is concise and well-structured. It front-loads the core action, then guides to the sibling tool, then details the output, then adds a non-obvious filtering behavior, and finishes with the direct pairing. No sentence is wasted, and the flow makes it easy to scan.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the absence of an output schema, the description supplies adequate shape information: each entry contains a title (implended by 'enumerate who's actually in it') plus an extract and thumbnail URL when available. It also shares the main-namespace filter and the prefix tolerance. Notationally, it could state ordering or pagination, but for a tool with 3 parameters and understandable semantics, this is sufficiently complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3. The description adds meaningful value: it clarifies that the category parameter accepts an optional prefix, and it explains the practical meaning of the extract and thumbnail responses that depend on the query. This goes beyond the schema's bare field descriptions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb and resource ('List Wikipedia articles filed under a category') and immediately differentiates itself from siblings by calling out `categories` and `search`. It even names a concrete example ('Machine learning researchers'), so an agent knows exactly what the tool does without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear usage context: it is the reverse of `categories`, and it pairs with `categories` (to build the taxonomy) and `search` (to find an entry point). This implies when to choose this tool over the obvious alternatives, though it does not spell out an explicit 'when not to use' list beyond that alignment.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
citation_neededA
Find statements Wikipedia has flagged as needing a source. Two angles: pass article to extract every sentence in that article carrying a {{citation needed}} tag (the exact claims editors flagged, with tag dates), or pass topic (or neither) to search Wikipedia for articles with unsourced claims on that topic, each with the flagged claim's text. Use it to find sourcing work as an editor or to spot the shakiest claims in a topic you're researching. The verification companion to references: this finds what's missing a source. Tag-name matching is English-centric; lang accepted for API consistency. Read-only — GET only, no new dependencies.
| Name | Required | Description | Default |
|---|---|---|---|
| lang | No | Wikipedia language code (default 'en') | en |
| limit | No | Max claims/articles to return (default 10, max 25) | |
| topic | No | Keyword to scope the search (e.g. 'climate'); omit for a Wikipedia-wide sample | |
| article | No | Exact article title to scan for tagged claims (takes precedence over topic) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description usefully discloses read-only GET-only behavior, no new dependencies, and an English-centric tag-matching limitation. It does not cover auth or rate-limit behavior, but the safety and tagging constraints are clearly surfaced.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the purpose and mode selection, with each sentence contributing useful routing or behavior. It is slightly dense and includes minor redundancy such as 'read-only — GET only', but remains well-structured and appropriately sized.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given no output schema, the description does a reasonable job explaining what each mode returns: tagged sentences with dates for `article`, and articles with flagged claim text for `topic`. It omits some return-shape and pagination details, but the core behavior is sufficiently covered.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3. The description restates the `article` vs `topic`/neither modes and precedence, and adds that `article` returns tagged sentences with tag dates, but it adds limited parameter-level syntax or format detail beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: finding Wikipedia statements flagged as needing a source. It distinguishes itself from the sibling `references` by explicitly calling itself the verification companion that finds missing sources, so an agent can tell the tools apart without opening schemas.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear usage contexts: find sourcing work as an editor or spot shaky claims in a topic. It also explains the two primary modes via `article` vs `topic`/neither, but does not state explicit when-not conditions or fully route the agent away from alternatives.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
contributorsA
Who writes and maintains a Wikipedia article — the most active recent editors. Tallies the article's recent edit history (up to 500 revisions) into a ranked contributor table: top named editors by edit count with their share of sampled edits and user-page links, plus the anonymous (IP) edit share. A provenance companion to article_quality (the grade earned) and revisions (the raw edit log): a page tended by veteran caretakers reads differently from one mostly touched by drive-by IP edits, and the top names are the people to credit or check for conflicts of interest. Follows redirects. Read-only action API revisions query — GET only, no new dependencies.
| Name | Required | Description | Default |
|---|---|---|---|
| lang | No | Wikipedia language code (default 'en') | en |
| limit | No | Max top contributors to return (default 10, max 20) | |
| title | Yes | Article title (e.g. 'Albert Einstein' or 'Paris') |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure and does so thoroughly: it reveals that the tool follows redirects, uses a read-only action API revisions query, is GET only, has no new dependencies, samples up to 500 revisions, and outputs ranked named editors, edit shares, user-page links, and anonymous IP share. This is exceptionally complete.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is well organized and front-loaded with the core purpose. Each sentence adds value, though the provenance analogy is slightly expansive; overall there is no significant filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Despite having no output schema, the description spells out the return contents and key caveats: ranked contributors, edit shares, user-page links, anonymous IP share, sample size, redirect behavior, and read-only nature. With annotations absent, this fully equips an agent to invoke and interpret the tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema already documents all three parameters at 100% coverage, so the baseline of 3 applies. The description adds useful global context such as redirect-following and the 500-revision window, but it does not add per-parameter semantics beyond what the schema already provides.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a precise job description ('Who writes and maintains a Wikipedia article'), then specifies the mechanism: tallies recent edit history (up to 500 revisions) into a ranked contributor table. It names the output components and contrasts with sibling tools article_quality and revisions, so an agent can clearly distinguish it from related tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description positions the tool as a provenance companion to article_quality and revisions, and gives concrete use cases: crediting top editors, checking for conflicts of interest, and understanding whether a page is veteran-tended or IP-driven. It doesn't explicitly state when not to use it, but the contextual guidance is strong.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
deaths_on_this_dayA
Get notable deaths that happened on today's date (UTC) from Wikipedia's 'On This Day' feed — the deaths-only companion to on_this_day (which returns events). Useful for 'in memoriam' content hooks, obituary-style social posts, and newsletter intros. Pairs with on_this_day (events) and featured_article (today's long-form pick) for a full daily 'today in Wikipedia' digest.
| Name | Required | Description | Default |
|---|---|---|---|
| lang | No | Wikipedia language code (default 'en') | en |
| count | No | Number of deaths to return (default 5, max 10) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It adds useful behavioral context: the date is fixed to today's date in UTC, results are 'notable deaths', and the source is Wikipedia's On This Day feed. This clearly signals a read-only retrieval operation, though it does not describe return structure or pagination beyond that.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences with no filler: the first front-loads the core action and scope, the second adds use cases and sibling relationships. Every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple two-parameter read-only feed tool, the description is nearly complete: it names the source, scope, timezone, use cases, and sibling tools. Since there is no output schema, a brief note on the returned shape would have made it fully complete, but this is a minor gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, and the schema already fully documents 'lang' (with enum) and 'count' (default 5, max 10). The description adds no additional parameter-specific meaning, so baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Get notable deaths... from Wikipedia's On This Day feed') and distinguishes itself as the 'deaths-only companion to on_this_day' which returns events. An agent can immediately tell what this tool does and how it differs from its closest sibling.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly gives concrete use cases ('in memoriam' content hooks, obituary-style social posts, newsletter intros) and names complementary siblings (on_this_day, featured_article) for building a full digest. It also clarifies the boundary: deaths only, versus events from on_this_day.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
did_you_knowA
Get a random 'Did You Know' style fact from Wikipedia — great for hooks and general trivia
| Name | Required | Description | Default |
|---|---|---|---|
| lang | No | Wikipedia language code (default 'en') | en |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden. It discloses that the fact is random and sourced from Wikipedia, which is useful. However, it lacks details on rate limits, authentication requirements, error handling, or whether the fact changes on each call.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single, concise sentence that front-loads the core functionality and ends with a note on its practical use. No wasted words.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple tool with one optional parameter (language), no output schema, and no annotations, the description covers the basics but omits behavioral details like rate limits or authentication. It is adequate but has clear gaps in behavioral context.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents the 'lang' parameter fully. The description adds no additional parameter semantics beyond what the schema provides. Baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Get) and resource (random 'Did You Know' style fact from Wikipedia). It distinguishes itself from the sibling 'random' by referencing the 'Did You Know' style, and from 'dino_fact' by citing Wikipedia.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage via 'great for hooks and general trivia', but does not explicitly state when to use this tool versus alternatives like 'random' or 'dino_fact'. No exclusions or prerequisites are given.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
dino_factA
Get a 'Did You Know' style fact about dinosaurs or prehistoric life. Pass a specific species ('Tyrannosaurus', 'Spinosaurus') for a targeted fact, or call with no arguments for a random dino. Falls back to a random dino if the requested species isn't found on Wikipedia.
| Name | Required | Description | Default |
|---|---|---|---|
| lang | No | Wikipedia language code (default 'en') | en |
| species | No | Specific dinosaur name (e.g. 'Tyrannosaurus'). Empty for random. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It usefully discloses the fallback behavior (a missing species silently returns a random dino) and that facts originate from Wikipedia, but it omits any note on rate limits, network dependence, or the shape of the returned fact.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, front-loaded with purpose and then split cleanly between the targeted and random modes. Minor redundancy: the random-dino case is stated twice (no-arguments and fallback).
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a two-parameter, annotation-free tool with no output schema, the description covers purpose, both invocation modes, and fallback behavior. The only meaningful gap is the absence of any indication of what the returned fact looks like.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so both the 'species' and 'lang' parameters are already documented (including the empty-string-for-random convention). The description restates the species semantics but does not explain the 'lang' enum's effect on output, so the baseline of 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb+resource combination ('Get a Did You Know style fact about dinosaurs or prehistoric life'), which is concrete and distinguishable from the generic sibling 'did_you_know' by domain. It stops short of explicitly naming the sibling to route away from, so it is clear but not maximally differentiating.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It clearly states the two invocation modes: pass a species name for a targeted fact, or call with no arguments for a random dino. This is explicit context for how to use the tool, though it offers no exclusion guidance relative to sibling fact/article tools.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
disambiguationA
Resolve a Wikipedia disambiguation page into its candidate articles. Detects whether a title (e.g. 'Mercury', 'Apple', 'Python') is a disambiguation page and, if so, returns the structured option list — article title plus one-line description — grouped by the page's own sections. Resolves the classic dead-end where search/summary land on an ambiguous title: call this, pick the right candidate, then fetch it with summary or article_extract. Reports clearly when the title is a regular article or doesn't exist.
| Name | Required | Description | Default |
|---|---|---|---|
| lang | No | Wikipedia language code (default 'en') | en |
| limit | No | Max options to return (default 30, max 100) | |
| title | Yes | Article title (e.g. 'Mercury' or 'Apple') |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the behavioral burden, and it does well: it explains that output is a structured option list grouped by the page's own sections, and pre-empts the dead-end edge case by promising a clear report for regular or nonexistent titles. It stops short of covering pagination/limit interaction or confirming the operation is side-effect-free, but the read-only nature is strongly implied.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three dense sentences, front-loaded with the core action, then the workflow, then the edge-case behavior. No filler and no repetition of the name.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 3-parameter lookup with no output schema and no annotations, the description supplies the return shape (title plus one-line description, section-grouped), the recommended follow-up tools, and the non-disambiguation outcomes. Nothing an agent needs to call it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the schema already documents lang, limit, and title fully; baseline 3 applies. The description adds example titles that mirror the schema's own example and nothing about lang or limit semantics beyond what the schema says.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Specific verb+resource: 'Resolve a Wikipedia disambiguation page into its candidate articles,' with concrete examples ('Mercury', 'Apple', 'Python'). It is clearly distinguished from siblings like search, summary, and article_extract, which it explicitly references as the tools that produce ambiguous results.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly states the triggering condition ('search/summary land on an ambiguous title') and the full workflow: call this, pick the right candidate, then fetch it with summary or article_extract. Also clarifies the negative case — it reports when the title is a regular article or doesn't exist — so the agent knows when this tool is not the answer.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
external_linksA
List external (off-wiki) links from a Wikipedia article — citations, references, primary sources, and other off-wiki resources the article points to. Outbound complement to links (internal outgoing) and backlinks (internal incoming): the trio (links + backlinks + external_links) maps the full reference network around an article. Useful for source verification, citation audits, primary-source discovery, and fact-checking research.
| Name | Required | Description | Default |
|---|---|---|---|
| lang | No | Wikipedia language code (default 'en') | en |
| limit | No | Max external links to return (default 20, max 50) | |
| title | Yes | Article title (e.g. 'Tyrannosaurus' or 'Albert_Einstein') |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden. 'List' implies a safe read, and the scope ('off-wiki') is clarified, but nothing is said about error behavior for a missing article, pagination beyond the limit param, or return shape. Adequate but incomplete for a tool with zero annotation coverage.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core verb+resource, then the trio framing, then use cases — a logical ordering. Slightly long with two enumerations of examples, but each sentence earns its place by adding differentiating context.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema and no annotations, the description still covers purpose, sibling positioning, and use cases, and the schema fully documents all three parameters. The only real gap is behavior on failure (nonexistent/invalid title) and any rate-limit or pagination nuance.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so `lang`, `limit`, and `title` are already fully documented in the schema (including defaults and max). The description adds no format, syntax, or edge-case guidance for those parameters, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('List') and resource ('external/off-wiki links from a Wikipedia article') and explicitly distinguishes itself from siblings `links` (internal outgoing) and `backlinks` (internal incoming). An agent can immediately tell which of the trio to call.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Names the scenario set (source verification, citation audits, primary-source discovery, fact-checking) and frames the tool as the 'outbound complement' within a trio, giving the agent a clear selection rationale versus `links` and `backlinks`. No competing tool overlaps are left ambiguous.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
featured_articleA
Get today's Wikipedia Featured Article — the single article Wikipedia's editors showcase as the best of the encyclopedia that day. Returns the full long-form extract plus thumbnail: reliably high-quality, surprising, in-depth content. A strong daily source of hooks and deep dives; for the curated daily image instead use picture_of_the_day, and for today's historical events use on_this_day.
| Name | Required | Description | Default |
|---|---|---|---|
| lang | No | Wikipedia language code (default 'en') | en |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations and no output schema, the description carries the full burden and does disclose the returned payload ('full long-form extract plus thumbnail') and the qualitative nature of the content, which the structured fields do not. It stops short of describing error or edge behavior (e.g., what happens for a language with no featured article that day) or any rate/caching behavior, so it is helpful but not exhaustive.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core action and resource, then the return value, then the sibling routing — a sensible order. The middle sentence ('reliably high-quality, surprising, in-depth content') is somewhat promotional and expendable, but the piece as a whole is tight and well under the point of bloat.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-required-param, no-output-schema tool, the description supplies the retrieval semantics, the returned content shape, and the sibling alternatives an agent needs. The only remaining gap is edge-case behavior such as missing featured articles per language, which is minor given the tool's simplicity.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% for the single `lang` parameter, including its enum and default, so the baseline is 3. The description adds no language-specific meaning beyond what the schema already documents.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Get today's Wikipedia Featured Article') and immediately defines the resource as the single article Wikipedia editors showcase that day, so the agent knows exactly what is being retrieved. It also names the two nearest siblings it must not be confused with, making it unambiguous without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly routes the agent: use picture_of_the_day for the daily image and on_this_day for historical events, with a stated rationale ('a strong daily source of hooks and deep dives') for when this tool is the right pick. Both the when-to-use and when-to-use-something-else conditions are given.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
imageA
Get just the lead image for a Wikipedia article — returns both the 300px thumbnail URL and the full-resolution original URL from Wikipedia's REST summary endpoint. Useful when you want the article's image for embedding elsewhere (cards, Telegram posts, slide decks, README hero images) without the surrounding summary text. summary embeds the thumbnail inline; image exposes both URLs separately so downstream tools can fetch / display at any size. Returns a clean message if the article has no image.
| Name | Required | Description | Default |
|---|---|---|---|
| lang | No | Wikipedia language code (default 'en') | en |
| title | Yes | Article title (e.g. 'Tyrannosaurus' or 'Albert_Einstein') |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are supplied, so the description carries the disclosure burden, and it does reveal the return shape (300px thumbnail URL plus full-resolution original URL) and the no-image fallback behavior. It omits any note on rate limits, auth, or whether the URL is hotlink-safe, which keeps it short of a 5 for a zero-annotation tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core return contract, then the disambiguation from `summary`. The parenthetical use-case list ('cards, Telegram posts, slide decks, README hero images') is slightly padded and could be trimmed without losing meaning.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema, so the description's explanation of the two returned URLs and the empty-image case does real work. Combined with a fully documented two-parameter schema, an agent has everything needed to call and interpret the result.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and both params (lang enum, title) are fully documented in the schema, so the baseline of 3 applies. The description adds no syntax or format detail about `title` handling (e.g. underscores vs spaces) beyond the schema's own example.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a precise verb and resource ('Get just the lead image for a Wikipedia article') and immediately names what it returns. It explicitly differentiates itself from the sibling `summary`, which embeds the thumbnail inline, so an agent can route correctly without opening either schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives a concrete when-to-use condition ('when you want the article's image for embedding elsewhere ... without the surrounding summary text') with example scenarios, and names the alternative tool (`summary`) plus the reason to prefer this one. The contrast is explicit rather than inferred.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
infoboxA
Extract the structured fact box (infobox) from a Wikipedia article as a field/value table — dates, people, places, statistics, founders, CEOs, populations, capitals. The fastest path to a concrete fact ('who founded X?', 'what's the population of Y?') without wading through prose. Renders wikitext into clean plain text (citations stripped, links flattened, fields capped at 50). Use this for facts; use summary for the prose gist and article_extract for full text. Reports clearly when an article has no infobox.
| Name | Required | Description | Default |
|---|---|---|---|
| lang | No | Wikipedia language code (default 'en') | en |
| title | Yes | Article title (e.g. 'Albert Einstein' or 'Paris') |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It goes beyond the schema by stating that wikitext is rendered into clean plain text, citations are stripped, links are flattened, fields are capped at 50, and missing infoboxes are reported. This is rich, non-obvious behavioral context.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with the core action and output format, followed by concrete examples, behavioral details, and routing to alternatives. Every sentence earns its place with no filler. It is rich but appropriately compact.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple 2-parameter tool with 100% schema coverage and no output schema, the description fully explains what is returned, how it behaves, what happens when no infobox exists, and when to prefer sibling tools. Nothing needed for correct invocation is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents both `title` and `lang`. The description adds context about Wikipedia articles and example questions, but it does not add meaningful parameter-level detail beyond the schema. The baseline of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'Extract the structured fact box (infobox) from a Wikipedia article as a field/value table.' It names concrete content types and example questions, and it explicitly distinguishes itself from summary and article_extract. An agent can clearly tell what this tool does and how it differs from siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly says 'Use this for facts; use `summary` for the prose gist and `article_extract` for full text.' This gives direct selection guidance, names the alternatives, and states the condition for choosing this tool. No inference is required.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
linksA
List outgoing Wikipedia links from an article (the article network in raw form). Useful for graph-style discovery — given 'Tyrannosaurus', see which genera, paleontologists, formations, and anatomical terms it references. Filters to main namespace so talk/user/etc. don't pollute the result.
| Name | Required | Description | Default |
|---|---|---|---|
| lang | No | Wikipedia language code (default 'en') | en |
| limit | No | Max links to return (default 20, max 50) | |
| title | Yes | Article title (e.g. 'Tyrannosaurus' or 'Albert_Einstein') |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden. It does disclose a genuinely useful trait not in the schema: results are filtered to main namespace so talk/user pages don't pollute output, and that this is the raw article network. However it says nothing about ordering, deduplication, or truncation behavior beyond the schema's limit.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, front-loaded with the core action and scope. The middle illustrative sentence is a bit long but earns its place by grounding the use case; no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
A simple read-only list tool with 3 fully documented params and no output schema; the description covers scope, the outgoing-vs-incoming distinction, and namespace filtering. Missing only minor operational detail like ordering.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so lang, limit and title are already documented in the schema. The description repeats the 'Tyrannosaurus' title example but adds no syntax, format, or constraint meaning beyond what the schema provides. Baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a precise verb and resource ('List outgoing Wikipedia links from an article'), and the word 'outgoing' implicitly separates it from the sibling 'backlinks'. An agent can tell what it retrieves without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives a clear use case with a concrete example ('given Tyrannosaurus, see which genera, paleontologists, formations...'), which orients the agent on when this is the right tool. It stops short of naming alternatives or stating exclusions, so it does not reach a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
media_listA
List all media (images, videos, audio) used in a Wikipedia article — not just the lead thumbnail. Returns a structured markdown list: file title, type (image/video/audio), caption, and thumbnail URL. Lead media is marked so callers can skip it when they already have it via image. Uses Wikipedia's REST /page/media-list endpoint (structured JSON, no HTML parsing). Pairs with image (lead only) — use image for the headline thumbnail, media_list for the full inventory (gallery generation, fact-checking, slide decks, audits). limit clamps the number of items (default 25, max 100).
| Name | Required | Description | Default |
|---|---|---|---|
| lang | No | Wikipedia language code (default 'en') | en |
| limit | No | Max media items to return (default 25, max 100) | |
| title | Yes | Article title (e.g. 'Tyrannosaurus' or 'Albert_Einstein') |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden. It discloses the output shape (structured markdown list with file title, type, caption, thumbnail URL), notes lead-media marking, identifies the underlying endpoint and that it is structured JSON with no HTML parsing, and states the limit default/max. Remaining gaps are minor (error behavior, truncation signaling beyond limit).
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the core scope in the first clause, then layers output shape, the `image` comparison, and the `limit` semantics. A few sentences are dense but each earns its place; no restating of the tool name or boilerplate.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 3-parameter, no-output-schema read tool, the description covers scope, output format, sibling routing, and a key parameter constraint — everything an agent needs to call it correctly. It could mention error/empty cases or auth requirements, but those are minor for a public Wikipedia endpoint.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3. The description does add useful runtime context for `limit` (default 25, max 100, clamping behavior) and confirms that `title` is the article being inventoried, but it doesn't add anything beyond the schema for `lang`.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (List) and resource (all media in a Wikipedia article) and immediately disambiguates scope ('not just the lead thumbnail'). It names the sibling `image` and draws the exact boundary between the two, so an agent can route correctly without opening either schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly names the alternative (`image` for the headline thumbnail) and the conditions that select this tool ('full inventory — gallery generation, fact-checking, slide decks, audits'). It also anticipates the case where the caller already has the lead media ('marked so callers can skip it when they already have it via `image`').
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
media_of_the_dayA
Get Wikimedia Commons' Media of the Day — the curated daily video/audio clip from Commons' Media of the Day selection. Returns media kind, duration, a preview thumbnail (video) or listen link (audio), artist, license, and description. Accepts an optional YYYYMMDD date (default today UTC) to browse past picks — the motion-and-sound counterpart to picture_of_the_day for daily content hooks.
| Name | Required | Description | Default |
|---|---|---|---|
| date | No | Date in YYYYMMDD format (default: today UTC) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations supplied, the description carries the behavioral burden. It discloses what the tool returns (media kind, duration, preview/listen link, artist, license, description) and the default date behavior. It does not describe edge cases such as invalid dates or missing media, but the disclosed behavior is solid for a read-style tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single dense sentence, but every clause contributes: purpose, output fields, parameter behavior, and sibling relation. It is front-loaded and reasonably sized, though splitting into two sentences would improve scannability.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple tool with one optional parameter and no output schema, the description is complete: it names the resource, the optional input semantics, and explicitly lists all return fields. An agent has enough information to decide when to call it and what to expect from the result.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema already documents the date parameter fully (format and default), so the baseline is 3. The description adds value by emphasizing optionality, the default-today behavior, and the purpose of browsing past picks, which goes beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'Get Wikimedia Commons' Media of the Day'. It clearly identifies the tool as a video/audio counterpart to picture_of_the_day and enumerates what it returns, so an agent can distinguish it from sibling tools without inspecting schemas.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explains when the optional date parameter is useful ('to browse past picks') and frames the tool as a counterpart to picture_of_the_day for 'daily content hooks'. However, it does not explicitly contrast with media_search or media_list, leaving some ambiguity about when those siblings should be preferred.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
media_searchA
Search Wikimedia Commons for freely-licensed media by keyword — the topic-based counterpart to image (lead image of a known article) and media_list (media already used in an article). Give it a topic ('aurora borealis', 'vintage trains') and get back real Commons files with thumbnail and full-size URLs, dimensions, license, and artist — ready to embed in cards, posts, slide decks, or README hero images. filetype filters to 'image' (default: photos + diagrams / SVGs), 'video', 'audio', or 'all'. Commons is language-independent, so this tool takes no lang parameter. Everything returned is freely licensed; check the license on the file page before reuse.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | Max results to return (default 10, max 50) | |
| query | Yes | Media search keywords (e.g. 'aurora borealis' or 'steam locomotive') | |
| filetype | No | Media type filter: 'image' (photos + diagrams/SVGs), 'video', 'audio', or 'all' | image |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden of explaining behavior. It discloses what is returned (URLs, dimensions, license, artist), the default filetype behavior, the language-independent nature, and a licensing caveat. It does not mention error behavior or rate limits, but those are not essential for this kind of search tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is efficiently organized into four sentences: purpose, output detail, parameter semantics, and licensing caveat. It is somewhat dense, but every sentence carries useful information and the main purpose is front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given there is no output schema, the description adequately explains the return values and gives practical context for embedding media. It also covers filtering, language independence, and licensing, leaving no significant gap for an agent to invoke the tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Although the input schema already covers all parameters at 100%, the description adds useful semantics: it provides query examples, clarifies that `filetype` 'image' includes photos and diagrams/SVGs, and explains why there is no language parameter. This goes beyond simply restating the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'Search Wikimedia Commons for freely-licensed media by keyword.' It also explicitly differentiates from siblings by naming `image` and `media_list` and describing their distinct use cases, so an agent can immediately understand what this tool does and how it is unique.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives clear selection criteria: use this for topic-based searches, `image` for a known article's lead image, and `media_list` for media already used in an article. It also explains why no `lang` parameter exists, preventing an agent from expecting language support.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
nearbyA
List Wikipedia articles geographically near a location — location-based discovery. Anchor by article title (e.g. 'Eiffel Tower' — uses that article's coordinates, no geocoding service needed) or by explicit lat/lon. Returns nearby articles with distances in meters/km. Useful for travel research and 'what's notable around here' questions.
| Name | Required | Description | Default |
|---|---|---|---|
| lat | No | Latitude in decimal degrees (-90 to 90). Requires lon. | |
| lon | No | Longitude in decimal degrees (-180 to 180). Requires lat. | |
| lang | No | Wikipedia language code (default 'en') | en |
| limit | No | Max articles to return (default 20, max 50) | |
| title | No | Anchor article title (e.g. 'Eiffel Tower'). Use instead of lat/lon. | |
| radius | No | Search radius in meters (default 1000, max 10000) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the burden and does add real behavioral context: it is a read-only listing operation, title anchoring reuses the article's own coordinates (no geocoding service), and results include distances. It omits auth, pagination, and error behavior, which keeps it from a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core action and resource, then the two anchoring modes and the payoff. Most sentences earn their place, though 'Useful for travel research' is mild filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With 6 optional parameters, no annotations, and no output schema, the description supplies the missing glue: the anchoring modes and the return content. It never explicitly states that one of title or lat+lon is effectively required, which is the one remaining gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the description adds meaning the schema does not: the title-vs-lat/lon mutual exclusivity and the fact that title anchoring avoids geocoding. That is genuine added semantics rather than restatement.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('List Wikipedia articles geographically near a location') and immediately scopes it as location-based discovery. This clearly separates it from siblings like search, random, or featured_article without needing to name them.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explains the two anchoring modes and when to prefer each (title 'Use instead of lat/lon'), plus a concrete use case ('what's notable around here' questions). It lacks any explicit exclusion or named alternative sibling, so it falls short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
newsA
Get current events from Wikipedia's Main Page 'In the news' section — the editorially-curated list of recent notable events. Pairs with featured_article (today's long-form pick) and on_this_day (historical) — news covers the present tense. Bold-linked article titles become Markdown so the main subject of each event stands out.
| Name | Required | Description | Default |
|---|---|---|---|
| lang | No | Wikipedia language code (default 'en') | en |
| limit | No | Max events to return (default 5, max 10) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden, and it does disclose meaningful behavior: the content is editorially curated (not algorithmic/search-driven) and bold-linked article titles are converted to Markdown. It does not mention caching, freshness, or failure behavior, but for a read-only fetch the disclosure is substantive.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, front-loaded with the source and scope, then sibling differentiation, then output formatting. Every sentence carries information; the Markdown note could arguably sit in a return-value section but does earn its place given there is no output schema.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a two-parameter, no-required-inputs read tool with no output schema, the description covers source, editorial nature, sibling alternatives, and output formatting (Markdown emphasis). Nothing an agent needs to call it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so both lang and limit are already documented in the schema with defaults and the enum list. The description adds no format or semantic detail beyond that, so the baseline 3 is correct.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource (get current events from Wikipedia's Main Page 'In the news' section) and names the exact source and scope. It explicitly differentiates itself from featured_article and on_this_day by framing itself as 'the present tense,' so an agent can route without opening a schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It names the sibling tools (featured_article for long-form, on_this_day for historical) and gives the selecting condition ('news covers the present tense'). It stops short of an explicit when-not-to-use clause, but the routing rule is clear and actionable.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
on_this_dayA
Get historical events that happened on today's date (UTC) from Wikipedia's 'On This Day' feed. Returns a random sample of events with year + description + Wikipedia link — great daily content hook alongside featured_article.
| Name | Required | Description | Default |
|---|---|---|---|
| lang | No | Wikipedia language code (default 'en') | en |
| count | No | Number of events to return (default 5, max 10) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does reasonably well: it discloses the data source, timezone basis (UTC), and the key behavioral quirk that results are a 'random sample' rather than a deterministic or complete list. It does not cover auth needs or rate limits, but for a public read-only feed tool that gap is minor.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences: the retrieval action and source come first, the return shape second, with no redundant restatement of the tool name. Every clause carries information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, so the description must describe returns — and it does ('year + description + Wikipedia link', random sample). Combined with a fully documented 2-param schema, an agent has everything needed to invoke it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% (both lang and count are documented with defaults and max), so the baseline is 3. The description adds the UTC date framing but nothing further about lang semantics or count behavior beyond what the schema already states.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Get historical events ... from Wikipedia's On This Day feed') and scopes it to today's date in UTC. It also names the sibling it pairs with (featured_article), making it distinguishable from random, did_you_know, and deaths_on_this_day.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'great daily content hook alongside featured_article' gives an implied usage context, but there is no explicit when/when-not guidance and no mention of closer alternatives like deaths_on_this_day (which covers the same date concept) or random. Usage must be largely inferred.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
pageviewsA
Get daily view counts for a Wikipedia article over a date range (popularity research, trending topics, historical interest). Uses Wikimedia's pageviews REST API. Default window is the last 7 days ending yesterday UTC. Returns total + daily average + markdown table of daily views.
| Name | Required | Description | Default |
|---|---|---|---|
| end | No | End date in YYYYMMDD (default: yesterday UTC) | |
| lang | No | Wikipedia language code (default 'en') | en |
| start | No | Start date in YYYYMMDD (default: 7 days before end) | |
| title | Yes | Article title (e.g. 'Tyrannosaurus' or 'Albert_Einstein') |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the burden and does well: it discloses the upstream source (Wikimedia's pageviews REST API), the default window semantics (last 7 days ending yesterday UTC), and the return shape. It omits error behavior, rate limits, and whether the title must match exactly, which mutes a perfect score.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, each earning its place: purpose, situational use, source API/default window, and return format in that order. Front-loaded with the core verb and free of filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description correctly compensates by naming the return contents (total, daily average, markdown table). Defaults, source, and return format are all covered for a straightforward read-only lookup, leaving nothing an agent needs missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema itself documents title, start, end, and lang. The description's note about the default window duplicates the schema defaults rather than adding new syntax or format meaning, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Get daily view counts for a Wikipedia article') with scope ('over a date range'), which separates it from the sibling aggregate tools like top_reads and recent_changes. An agent can identify the resource type (per-article pageviews) without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The parenthetical '(popularity research, trending topics, historical interest)' gives clear situational context for when to reach for this tool. It stops short of naming an alternative or stating when-not to use it (e.g. vs. top_reads for site-wide rankings), so it doesn't reach the top band.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
picture_of_the_dayA
Get Wikimedia Commons' Picture of the Day — the curated daily image from Wikipedia's featured feed. Returns preview + full-size image URLs, photographer, license, and description. Accepts an optional YYYYMMDD date (default today UTC) to browse past pictures — the visual counterpart to featured_article for daily content hooks.
| Name | Required | Description | Default |
|---|---|---|---|
| date | No | Date in YYYYMMDD format (default: today UTC) | |
| lang | No | Wikipedia language code (default 'en') | en |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations, so the description carries the full burden. It usefully discloses the return payload (preview + full-size URLs, photographer, license, description) and the default date behavior, but says nothing about error cases, permission/auth needs, rate limits, or what happens for dates with no picture.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three dense sentences, front-loaded with the resource and what it returns, then the optional parameter and the sibling relationship. No filler or repetition; every clause carries information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description correctly compensates by enumerating returned fields (URLs, photographer, license, description), and it explains the date default. The lang parameter's effect on output is left to the schema, a minor gap for an otherwise self-sufficient read tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so both parameters are already documented in the schema; baseline is 3. The description adds the YYYYMMDD format and UTC-default meaning for date, but does not mention the lang parameter or its enum values at all.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Get Wikimedia Commons' Picture of the Day') with the exact scope ('curated daily image from Wikipedia's featured feed'). It explicitly self-positions against a sibling by calling itself 'the visual counterpart to featured_article,' so an agent can distinguish it from 20+ other retrieval tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear context: default is today, and passing a date lets the agent 'browse past pictures,' framed as 'daily content hooks.' It names featured_article as the related tool, but offers no explicit when-not-to-use or selection rule beyond that pairing.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
quoteA
Get a random notable quote from a curated list of famous authors (Churchill, Einstein, Twain, Gandhi, Mandela, Wilde, Angelou, Jobs, Lennon, Socrates, etc.). Returns a short, time-tested quotation with author attribution. Pairs with did_you_know (random encyclopedia fact) and dino_fact (random dino fact) for variety in 'today's trivia' outputs — great for daily content hooks, social posts, newsletter intros. Currently English-only (curated list); the lang parameter is accepted for API consistency but non-English values still return English quotes.
| Name | Required | Description | Default |
|---|---|---|---|
| lang | No | Wikipedia language code (default 'en'). Currently English-only; non-English values fall back to English. | en |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden, and it does disclose real behavior: selection is random from a curated (not open) corpus, and the output is a short quote with author attribution. It also flags a genuine gotcha — the `lang` parameter is accepted but non-English values silently fall back to English. It omits any mention of rate limits or caching, which keeps it below a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with what the tool returns, then the related-tool pairing, then the language caveat, so the critical limitation is not buried. A clause of promotional framing ('great for daily content hooks, social posts, newsletter intros') is soft filler, but the sentences otherwise earn their place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a one-optional-parameter trivia tool with no output schema, the description covers purpose, the shape of the return value, the fallback behavior, and sibling relationships — enough to call it correctly. Nothing essential is missing; only marginal detail (rate limits, determinism) is absent.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is already 100%, so the baseline is 3, but the description adds non-obvious meaning: the enum of ten language codes is effectively a no-op and returns English anyway. That directly corrects a misleading schema signal and is the single most useful thing an agent could know about this parameter.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Get a random notable quote') plus the source pool (curated list of named authors). It also names the sibling tools it complements (did_you_know, dino_fact), so an agent can separate it from the generic 'random' tool without opening a schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives concrete usage contexts ('today's trivia' outputs, daily content hooks, social posts, newsletter intros) and lists the tools it pairs with. It stops short of any exclusion or routing rule — notably it never distinguishes itself from the sibling 'random' tool — so it is clear context but not full when/when-not guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
randomA
Fetch a summary of a random Wikipedia article — serendipitous discovery across the entire encyclopedia. Returns the same title + summary + thumbnail shape as summary, but for a surprise topic. Ideal for exploration, icebreakers, trivia, and content inspiration when there's no specific subject in mind; supports other languages via lang.
| Name | Required | Description | Default |
|---|---|---|---|
| lang | No | Wikipedia language code (default 'en') | en |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden. It does disclose the return shape ('title + summary + thumbnail') and language support, which is useful. However, it says nothing about the non-deterministic nature of the result, failure/empty-result behavior, rate limits, or whether results are cached, which matters for a randomness-based tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core action and scoping, followed by the return-shape comparison and use cases. Two sentences with little waste, though 'serendipitous discovery across the entire encyclopedia' and the icebreakers/trivia list lean slightly promotional.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-required-parameter, read-only tool with no output schema, the description does the important work of summarizing the return shape and language behavior. What remains missing is only minor behavioral detail (randomness, error cases), so it is nearly complete for its complexity.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and the single `lang` parameter already documents itself with an enum and default. The description only adds 'supports other languages via `lang`,' which is marginal over the schema. Baseline 3 applies when the schema does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (fetch) and resource (a random Wikipedia article summary) and explicitly contrasts itself with sibling `summary` by noting it returns the same title + summary + thumbnail shape 'but for a surprise topic.' An agent can distinguish this from `summary`, `featured_article`, and `did_you_know` without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear usage context: 'exploration, icebreakers, trivia, and content inspiration when there's no specific subject in mind,' and implicitly routes to `summary` by referencing it as the same-shaped-but-targeted alternative. The condition for choosing this tool is stated, but the inverse case ('use summary when you do have a subject') is only implied rather than spelled out.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
recent_changesA
Show the most recent changes to Wikipedia articles — a live window into what editors are doing right now. Uses the MediaWiki recentchanges feed over the article namespace. kind filters the stream: 'edit' (text changes), 'new' (newly published articles — a discovery feed of brand-new pages), 'categorize' (category membership changes), 'log' (page moves, deletions, protections). Complements revisions (history of one article) with the reverse angle: the freshest edits everywhere. Each entry shows the change kind, article link, byte-size delta, editor, timestamp, edit comment, and a diff link — great for spotting breaking-news edits, new pages on emerging topics, and bot-maintenance sweeps. Read-only.
| Name | Required | Description | Default |
|---|---|---|---|
| kind | No | Filter the change stream: 'all' (default), 'edit', 'new', 'categorize', or 'log' | all |
| lang | No | Wikipedia language code (default 'en') | en |
| limit | No | Max changes to return (default 10, max 50) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden and does notable work: it declares 'Read-only,' identifies the underlying feed and namespace scope, and enumerates the shape of each returned entry (kind, article link, byte delta, editor, timestamp, comment, diff link). It does not cover rate limits, latency/freshness guarantees, or pagination, which keeps it out of the top band.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Purpose and source are front-loaded in the first sentence, followed by the filter semantics and the sibling comparison. It is somewhat long and includes mild promotional framing ('great for spotting breaking-news edits'), but each sentence conveys usable information rather than filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 3-param, zero-required, read-only feed tool with no output schema, the description supplies the missing return-value detail and the read-only safety profile. An agent has enough to call it correctly; only operational details like pagination or refresh behavior are absent.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline would be 3, but the description adds genuine semantic meaning beyond the enum labels — 'edit' = text changes, 'new' = newly published articles, 'categorize' = category membership changes, 'log' = page moves, deletions, protections. It does not elaborate on `lang` or `limit`, which the schema already handles.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb + resource ('show the most recent changes to Wikipedia articles') and scopes it to the article namespace via the MediaWiki recentchanges feed. It also explicitly names the sibling it differs from ('Complements `revisions`... with the reverse angle'), so an agent can disambiguate without opening either schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear selection guidance: it explains that `revisions` is for the history of one article while this is the freshest edits everywhere, and it maps each `kind` value to a use case ('new' as a discovery feed, 'log' for moves/deletions/protections). It stops short of explicit when-not-to-use rules or prerequisites, but the alternative is named with the condition that selects it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
referencesA
The sources an article cites — its bibliography. Returns the article's numbered reference list: each citation's text plus the off-wiki URLs it points to (DOI, publisher, archive, primary-source links). The verification companion to external_links (which dumps every off-wiki link on the page): references returns only the sources the article actually cites. Useful for source verification, citation audits, bibliography building, and primary-source discovery. Follows redirects. Read-only parse API — GET only, no new dependencies.
| Name | Required | Description | Default |
|---|---|---|---|
| lang | No | Wikipedia language code (default 'en') | en |
| limit | No | Max citations to return (default 20, max 50) | |
| title | Yes | Article title (e.g. 'Albert Einstein' or 'Paris') |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the behavioral burden and handles it well: it discloses 'Read-only parse API — GET only, no new dependencies' and 'Follows redirects.' It could go further by specifying handling of missing titles or empty reference lists, but the core side effects and safety profile are clear.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Every sentence earns its place: the core function, response contents, sibling comparison, use cases, and behavioral notes. It is structured to front-load purpose and then build outward without redundancy or filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given there is no output schema, the description sufficiently explains what will be returned (numbered list with citation text and off-wiki URLs). It also covers safety behavior and use context, so an agent has everything needed to invoke it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all three parameters thoroughly. The description adds no parameter-specific semantics beyond the output, correctly leaving parameter detail to the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb and resource: it 'Returns the article's numbered reference list' with details on what each citation includes. It also explicitly distinguishes itself from the sibling tool external_links, making its scope unmistakable.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It names the alternative (external_links), explains the key difference between them, and lists concrete use cases (source verification, citation audits, bibliography building, primary-source discovery). This gives an agent clear guidance on when to select this tool.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
revision_diffA
Compare two revisions of an article and show exactly what changed — a plain-text unified diff of the article's wikitext between revision rev_from and rev_to (get revision IDs from revisions). Each side is labelled with its timestamp, editor, and edit summary. The edit-auditing companion to revisions: use it to see what a specific edit added or removed, review edits before trusting a new paragraph, or spot stealth rewrites. limit clamps diff lines (default 100, max 500).
| Name | Required | Description | Default |
|---|---|---|---|
| lang | No | Wikipedia language code (default 'en') | en |
| limit | No | Max diff lines to show (default 100, max 500) | |
| title | Yes | Article title (e.g. 'Tyrannosaurus') | |
| rev_to | Yes | Newer revision ID (from `revisions`) | |
| rev_from | Yes | Older revision ID (from `revisions`) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It discloses the output format (plain-text unified diff), that each side is labelled with timestamp, editor, and edit summary, and the limit clamping behavior. It does not explicitly state it is read-only, but that is implied by 'compare'. It does not mention error cases or rate limits, but the core behavior is well covered. The description adds value beyond the schema without contradicting any annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, front-loaded with the primary purpose and output format. Every clause adds value: the diff description, the labelling detail, the companion reference, and the limit clamp. No wasted words.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Without an output schema, the description conveys what the agent can expect: a plain-text diff with labelled sides and a line limit. It covers the main inputs and their source. It does not specify error behavior (e.g., missing revisions) or pagination beyond the limit, but for a diff tool this is adequate. The description is sufficient for an agent to call the tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so parameters are already documented. The description adds extra meaning: it clarifies that rev_from is the older revision and rev_to the newer (also in schema), explains the limit default and max (also in schema), and tells the agent to obtain revision IDs from `revisions`. This goes beyond the schema's bare descriptions, so a score above the baseline is warranted.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description explicitly states the verb 'compare' with the resource 'revisions' and explains the output is a plain-text unified diff. It clearly distinguishes itself from the sibling tool `revisions` by framing itself as the edit-auditing companion. The purpose is specific and unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explicitly names the sibling `revisions` and provides concrete use cases: 'see what a specific edit added or removed, review edits before trusting a new paragraph, or spot stealth rewrites.' It also tells the agent where to get revision IDs ('get revision IDs from `revisions`'), giving clear when-to-use guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
revisionsA
Show an article's recent edit history ('View history'): revision id, timestamp, editor, edit summary, and byte-size delta vs the previous revision. Each revision links to its diff (Special:Diff/) so you can inspect exactly what changed. Useful for tracking how an article evolves, auditing edits on a topic, or spotting edit activity around current events. Complements pageviews (popularity) with provenance (who changed what, when).
| Name | Required | Description | Default |
|---|---|---|---|
| lang | No | Wikipedia language code (default 'en') | en |
| limit | No | Max revisions to return (default 10, max 50) | |
| title | Yes | Article title (e.g. 'Tyrannosaurus' or 'Albert_Einstein') |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does so well: it discloses the exact return fields, the byte-size delta semantics, and that each revision links to its diff via Special:Diff/<revid>, implying a safe read. It omits pagination/limit behavior and rate-limit or auth context, which are minor for a read-only query.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core purpose and return fields, then use cases and sibling contrast. Four sentences stay efficient, though the diff-link detail and use-case list could be trimmed slightly.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, so the description must describe return values — and it does thoroughly (id, timestamp, editor, summary, delta, diff link). An agent gets enough to call it correctly; only pagination/limit semantics and ordering are left implicit.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents lang (with enum), limit (default and max), and title. The description adds nothing beyond the schema about parameter meaning or constraints, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Show an article's recent edit history') and enumerates exactly what it returns (revision id, timestamp, editor, edit summary, byte-size delta). It also explicitly distinguishes itself from the sibling `pageviews` by contrasting provenance vs. popularity.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides clear usage contexts (tracking article evolution, auditing edits, spotting current-event activity) and names `pageviews` as the complementary tool. However, it gives no when-not guidance and doesn't distinguish itself from plausible edit-oriented siblings like `recent_changes`, leaving the nearest alternative unaddressed.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
searchA
Search Wikipedia for articles matching a query — the entry point when you don't know the exact article title. Returns a ranked, numbered list of matching articles, each with a short text snippet and its Wikipedia URL (use limit for up to 20 results). Once you've identified the right title, pass it to summary for the gist, article_extract for the full text, or links to explore outward. Reports 'No results found' for empty queries instead of erroring.
| Name | Required | Description | Default |
|---|---|---|---|
| lang | No | Wikipedia language code (default 'en') | en |
| limit | No | Max results (default 5, max 20) | |
| query | Yes | Search query |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the burden and does well: it discloses the return shape (ranked, numbered list with snippet and URL), the result cap (up to 20), and notable error-avoidance behavior ('Reports No results found for empty queries instead of erroring'). It doesn't cover cross-language behavior for `lang` or any rate/availability constraints, leaving a small gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, front-loaded with the purpose, then return shape, then routing guidance. Every sentence earns its place and nothing is repeated from the schema beyond a brief parenthetical.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, so the description must explain return values — and it does, describing the ranked numbered list, snippets, and URLs. Combined with the routing guidance to downstream tools, an agent has everything needed to call this correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the schema already documents `query`, `lang`, and `limit`. The description only echoes the `limit` cap (already in the schema as 'max 20'), adding no new semantic detail. Baseline 3 is appropriate when the schema does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Search Wikipedia for articles matching a query') and explicitly positions itself as 'the entry point when you don't know the exact article title,' which distinguishes it from siblings like `summary` and `article_extract` that require a known title.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit when-to-use ('when you don't know the exact article title') and names the downstream alternatives with their distinct purposes: `summary` for the gist, `article_extract` for full text, `links` to explore outward. The routing logic is fully spelled out.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
section_textA
Read ONE section of a Wikipedia article as plain text — the targeted-reading companion to article_sections. Pass a section number exactly as shown by article_sections (e.g. 2 or '2.1'; 0 reads the lead/intro), or a heading name (case-insensitive, e.g. 'Life and career'); misspelled names get close-match suggestions. Renders the section via the read-only parse API and returns clean text with paragraph structure — no need to pull the whole 50KB+ article via article_extract when you only want one part of it. Follows redirects; the link at the bottom deep-links to the section on the article page.
| Name | Required | Description | Default |
|---|---|---|---|
| lang | No | Wikipedia language code (default 'en') | en |
| title | Yes | Article title (e.g. 'Albert Einstein' or 'Paris') | |
| section | Yes | Section to read: number as shown by `article_sections` (e.g. 2 or '2.1'; 0 = lead/intro), or heading text (e.g. 'Life and career') |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description must carry the behavioral burden. It does this well by disclosing the read-only parse API, redirect-following, close-match suggestions for misspelled headings, return structure ('clean text with paragraph structure'), and the deep-link at the bottom. It does not mention error cases or rate limits, but the disclosed behaviors go well beyond a minimal description.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense but well-organized, front-loading the core purpose and then explaining arguments, output, and redirect behavior. Every sentence contributes useful information, though the deep-link sentence is slightly peripheral and could be considered extra detail.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given no output schema and no annotations, the description covers what the tool returns (clean text with paragraph structure), how to identify a section, what happens with misspellings, how redirects are handled, and when this tool beats its siblings. An agent has enough to select and call the tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema already documents all parameters with descriptions, giving 100% coverageوط; however, the description adds valuable semantics beyond the schema: section numbers are exactly as shown by `article_sections`, 0 reads the lead/intro, headings are case-insensitive, and misspelled headings get close-match suggestions. This meaningfully helps an agent construct correct arguments.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'Read ONE section of a Wikipedia article as plain text'. It distinguishes itself from siblings by naming `article_sections` and `article_extract`, making its scope immediately clear.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly says when to use this tool instead of `article_extract` ('no need to pull the whole 50KB+ article') and positions it as the targeted-reading companion to `article_sections`. It also gives concrete argument-passing guidance, so an agent knows exactly how to invoke it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
simple_summaryA
Get the Simple English Wikipedia version of a topic by title — a plain-language explanation written in short sentences with common words, in the same title + summary + thumbnail shape as summary. Use when you want "explain it simply" (kids, ESL readers, quick intuition) or the full article is too dense. Reports clearly when no Simple English article exists for the topic, pointing you at summary for the full version.
| Name | Required | Description | Default |
|---|---|---|---|
| title | Yes | Article title (e.g. 'Photosynthesis' or 'Quantum_mechanics') |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations exist, so the description carries the full burden, and it delivers meaningful context: the return shape (title + summary + thumbnail) and the behavior when no Simple English article exists. It stops short of return-format details, pagination, or any auth/rate-limit notes, but for a read-only lookup tool the coverage is solid.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One dense, front-loaded sentence that leads with the core action and then folds in audience, fallback, and edge-case behavior. Slightly long, but no sentence is wasted.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema and a single fully documented param, the description supplies everything needed to call correctly: the output shape by reference to `summary`, the missing-article behavior, and the alternative tool. Nothing material is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
A single parameter with 100% schema description coverage including an example ('Photosynthesis', 'Quantum_mechanics'), so the schema does the heavy lifting; the description only restates 'by title'. Baseline 3 per the high-coverage rule.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Get the Simple English Wikipedia version of a topic by title') and immediately differentiates itself from the sibling `summary` by naming the exact shape it mirrors. An agent can tell it apart from `summary` and `search` without opening a schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicit when-to-use ('explain it simply', kids/ESL readers, quick intuition, or when the full article is too dense), explicit fallback ('pointing you at `summary` for the full version'), and it discloses the no-Simple-English-article case rather than leaving the agent to guess.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
summaryA
Get a concise summary of a Wikipedia article by exact title, plus its lead thumbnail image when one exists. The fastest way to get the gist of a known topic — e.g. 'Tyrannosaurus' or 'Albert_Einstein'. If you don't know the exact title, use search first to find it. For the full article text use article_extract; for just the section outline use article_sections.
| Name | Required | Description | Default |
|---|---|---|---|
| lang | No | Wikipedia language code (default 'en') | en |
| title | Yes | Article title (e.g. 'Tyrannosaurus' or 'Albert_Einstein') |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the behavioral burden and does well: it discloses the return content (summary plus optional lead thumbnail) and the exact-title precondition, which is the key failure mode for this tool. It does not mention error behavior when a title doesn't exist, but the essential traits are covered.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, each earning its place: capability, usage hint with examples, then sibling routing. The core capability is front-loaded and nothing is padded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, so the description must convey returns — and it does (summary plus thumbnail when one exists). Combined with explicit sibling routing and full schema coverage, an agent has everything needed to select and invoke it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so both `title` and `lang` are already documented, including the enum values and default. The description only reinforces the exact-title semantics already in the schema, so baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Specific verb (Get) plus resource (concise summary of a Wikipedia article) with scope stated: by exact title, plus lead thumbnail when present. It clearly distinguishes itself from article_extract and article_sections, so an agent can pick it without opening schemas.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicit routing: exact title required, and if the title is unknown the agent is told to use `search` first. It also names the alternatives for adjacent needs (article_extract for full text, article_sections for outline), covering when-not-to-use.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
top_readsA
Get the most-read articles on Wikipedia for a given date. Uses Wikimedia's top-pageviews endpoint (all-access, daily). Default date is yesterday UTC (today's data is typically not yet finalized). Filters out non-content namespaces (Main_Page, Special:Search, Portal:Current_events, Wikipedia:*, etc.) so the result is real articles only. Pairs with pageviews (per-article over a range) for trending-vs-popular comparisons — top_reads answers 'what is everyone reading right now' while pageviews answers 'how is this specific article trending'.
| Name | Required | Description | Default |
|---|---|---|---|
| date | No | Date in YYYYMMDD (default: yesterday UTC) | |
| lang | No | Wikipedia language code (default 'en') | en |
| limit | No | Max articles to return (default 10, max 50) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden and does a good job: it discloses the data source (Wikimedia top-pageviews, all-access, daily), the default date rationale (today's data not finalized), and the namespace filtering that removes non-article noise. It does not cover rate limits, auth, or response shape, but the substantive behavior is well documented.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core purpose, followed by source, default-date caveat, filtering note, and sibling comparison. Slightly long, but each sentence earns its place; only the endpoint/routing sentence could be trimmed.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a no-annotation, no-output-schema tool with three well-documented params, the description is complete enough to invoke correctly: it covers source, date semantics, filtering, and sibling differentiation. Minor gaps (return shape, result ordering) remain, but nothing blocks correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the schema already documents date, lang, and limit. The description reinforces the date default (yesterday UTC) and explains the namespace filtering behavior, but adds no syntax or format detail beyond the schema. Baseline 3 is appropriate when the schema does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a precise verb+resource: 'Get the most-read articles on Wikipedia for a given date,' and then names the underlying endpoint. It clearly distinguishes top_reads (aggregate popularity) from the sibling pageviews (per-article trending), so an agent can route between them without opening either schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explicitly frames when to use this tool versus pageviews ('what is everyone reading right now' vs 'how is this specific article trending'), giving the agent a decision rule. It also notes the default date behavior and why finalized data matters.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
translationsA
List all language versions of a Wikipedia article (langlinks). Returns the other-language editions the article exists in — e.g. for 'Tyrannosaurus' (en), returns de/fr/es/ja/zh titles. Complements the one-way lang parameter used by other tools: every tool can query a single language, but only translations reveals the article's full language coverage so callers can pick a target language to fetch next. Useful for translation research (full coverage vs. stub languages), cross-language content sourcing, and language-coverage analysis. limit clamps the number of entries (default 30, max 100) — popular articles can have 100+ language versions.
| Name | Required | Description | Default |
|---|---|---|---|
| lang | No | Wikipedia language code to query from (default 'en') | en |
| limit | No | Max language entries to return (default 30, max 100) | |
| title | Yes | Article title (e.g. 'Tyrannosaurus' or 'Albert_Einstein') |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the disclosure burden and does a reasonable job: it reveals the result shape (other-language editions/titles), discloses the `limit` clamp with default and max, and warns that popular articles exceed 100 languages. It does not explicitly state that the operation is read-only or describe pagination/truncation behavior beyond the clamp, so a 4 rather than a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The core definition and the sibling differentiation are front-loaded, and most sentences earn their place. The use-case list ('translation research, cross-language content sourcing, language-coverage analysis') is somewhat enumerative and could be trimmed, but the description is otherwise tight.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema, so the description must convey return contents, and it does so at a useful level (other-language titles, e.g. de/fr/es/ja/zh). Combined with the coverage of all three parameters, an agent has enough to call it correctly; explicit read-only/pagination notes are the only meaningful omission.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the description adds practical meaning: it explains why `limit` exists ('popular articles can have 100+ language versions') and frames `lang` as the source-language selector that other tools expose one-way. The `limit` default/max is largely a restatement of the schema, so the added value is real but modest.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific verb and resource ('List all language versions of a Wikipedia article (langlinks)') and grounds it with a concrete example ('Tyrannosaurus' -> de/fr/es/ja/zh). It also differentiates itself from siblings by contrasting with the one-way `lang` parameter used elsewhere, so an agent can distinguish it without opening a schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It states clearly when this tool is the right choice: when the caller needs the article's full language coverage to pick a target language to fetch next, and lists concrete scenarios (translation research, cross-language sourcing, coverage analysis). It does not give explicit when-not-to-use exclusions or route to a specific alternative tool, keeping it just short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
user_contribsA
What a Wikipedia editor has been doing — their recent contributions across the encyclopedia. Shows the latest edits by a named account (or an IP address, e.g. user='192.0.2.1'): edited pages, timestamps, byte-size deltas, edit comments, and flags for new pages, minor edits, and edits still current. The header reports registration date and total edit count. The reverse angle of contributors (who edits this article): profile a top contributor, audit an anonymous IP's activity, or spot single-purpose accounts (edits confined to one topic hint at a conflict of interest). namespace scopes the search (default 0 = articles). Read-only — GET only, no new dependencies.
| Name | Required | Description | Default |
|---|---|---|---|
| lang | No | Wikipedia language code (default 'en') | en |
| user | Yes | Wikipedia username or IP address (e.g. 'Jimbo Wales') | |
| limit | No | Max contributions to return (default 10, max 50) | |
| namespace | No | Namespace to search (default 0 = articles; 3 = user talk) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations, so the description carries the burden and mostly does: it declares read-only, GET-only, no new dependencies, and enumerates the returned fields including flags and a header with registration date and edit count. It stops short of mentioning rate limits or pagination, but the behavioral profile is well covered for a read tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core purpose in the first clause, then details, then the sibling relationship, then a compact annotation-style tail. Dense but every sentence contributes; the parenthetical example list of fields is slightly heavy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description compensates by describing exactly what the response contains (edited pages, timestamps, byte deltas, comments, flags, header metadata). Combined with full param coverage and explicit read-only status, an agent has everything needed to call it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% (baseline 3), and the description still adds value: it explains that `user` accepts an IP address with a concrete example and that `namespace` scopes the search with the article default, which goes beyond the terse schema text.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource - 'recent contributions' by a named editor or IP - and explicitly positions itself as 'the reverse angle of `contributors`', letting an agent distinguish the two siblings immediately.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit use cases: profile a top contributor, audit an anonymous IP, spot single-purpose accounts. It also names the alternative tool and the axis on which the choice is made (who edits this article vs what this editor edits).
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
1 tool update
v1.1.26- Added
citation_needed
3 tool updates
v1.1.25- Added
disambiguation - Added
simple_summary - Added
user_contribs
1 tool update
v1.1.24- Added
revision_diff
1 tool update
v1.1.23- Added
section_text
1 tool update
v0.1.9- Added
references
3 tool updates
v0.1.8- Added
contributors - Added
media_of_the_day - Added
related_articles
1 tool update
v0.1.7- Added
births_on_this_day
1 tool update
v0.1.6- Added
article_quality
1 tool update
v0.1.4- Added
media_search
1 tool update
v0.1.3- Added
infobox
1 tool update
v0.1.2- Added
category_members
25 tool updates
v0.1.0- First observed
article_extract - First observed
article_sections - First observed
backlinks - First observed
categories - First observed
deaths_on_this_day - First observed
did_you_know - First observed
dino_fact - First observed
external_links - First observed
featured_article - First observed
image - First observed
links - First observed
media_list - First observed
nearby - First observed
news - First observed
on_this_day - First observed
pageviews - First observed
picture_of_the_day - First observed
quote - First observed
random - First observed
recent_changes - First observed
revisions - First observed
search - First observed
summary - First observed
top_reads - First observed
translations
TDQS
Scored across 40 tools
Tools are mostly distinct, and descriptions explicitly contrast similar ones (e.g., revisions vs recent_changes, links vs backlinks). However, with 40 tools, clusters like edit-history and daily-content tools require careful reading to avoid misselection.
All tool names use consistent snake_case with lowercase, no mixing of conventions. The naming pattern is predictable and readable throughout.
40 tools is far beyond the recommended 3-15 range, making the set heavy and potentially overwhelming for an agent. While the domain is broad, many tools could be consolidated or grouped.
The read-only surface covers article retrieval, search, media, categories, links, edit history, analytics, and daily feeds comprehensively. No obvious gaps for a Wikipedia reading server.
Maintenance
Related MCP Connectors
Wikipedia MCP — wraps Wikipedia REST API (free, no auth)
Search Wikipedia, read summaries and full text, target sections, find nearby pages, list languages.
Wikifeed MCP — wraps Wikimedia Feed API (free, no auth)
Wikidata MCP — wraps Wikidata API (wikidata.org/w/api.php)
Related MCP Servers
- AlicenseNot gradedqualityDmaintenanceProvides structured access to Wikipedia content including search, summaries, images, links, and more via MCP tools.6Apache 2.0
- FlicenseNot gradedqualityDmaintenanceMCP server providing live Wikipedia recent changes feed, page summaries, trending pages, and Wikidata entity lookup.-
- FlicenseNot gradedqualityDmaintenanceProvides comprehensive Wikipedia access for AI assistants via MCP Streamable HTTP transport, enabling search, article retrieval, summaries, section analysis, link discovery, and multi-language support.2-
- AlicenseNot gradedqualityAmaintenanceMCP server for searching and reading Wikipedia articles, including summaries, full text, targeted sections, nearby pages, and language editions.672 npm3Apache 2.0