Skip to main content
Glama
jermatic1

mcp-wikipedia

by jermatic1

mcp-wikipedia

An MCP server for searching Wikipedia's Vital Articles, the roughly 50,000 most important English articles, from a local index.

Setup

Copy .env.example to .env and set a User-Agent with your contact details, which Wikipedia asks clients to send.

Related MCP server: Wikipedia MCP Server

Run

Requires uv and Task.

task setup
task serve

The first run downloads the article list, streams the 37 GB English Wikipedia dataset keeping only the Vital Articles, and builds the index in ./data. If interrupted, it resumes where it left off. Later runs start immediately.

The server listens at http://localhost:8000/mcp.

Docker

Set UID and GID in .env to your user and group IDs so the files in ./data belong to you. Then:

mkdir data
docker compose up -d

Tools

  • search(query, limit=5): find articles by keywords or title; tolerates small typos. Returns title, description, start of the summary, Wikidata ID, Vital level, and topic.

  • get_article(title, section=None): an article's summary, infobox facts, and section headings. Pass section to also get that section's text.

Evaluate

Create data/questions.tsv with a question and the expected article title per line, separated by a tab, then run task evaluate to see how often the expected article is in the top 3 results.

Refresh the article list

The Vital list is fetched as it stood on the dataset's snapshot date, so its titles match the dataset. To move to a newer dataset, update REVISION and SNAPSHOT in src/mcp_wikipedia/download.py, then:

task vital --force
task serve

The articles are downloaded again for the new list and the index is rebuilt. With Docker: docker compose exec wikipedia task vital --force, then docker compose restart.

Development

task check    # tests, lint, formatting

License

MIT

Related MCP Connectors

Related MCP Servers

  • A
    license
    Not graded
    quality
    D
    maintenance
    Provides comprehensive access to Wikipedia content including article search, full text retrieval, summaries, categories, links, images, language versions, and external references through 9 specialized tools.
    24 npm
    MIT
  • A
    license
    Not graded
    quality
    D
    maintenance
    Provides offline search and retrieval of Wikipedia articles using Kiwix .zim files, enabling LLMs to access full Wikipedia content without internet.
    2
    MIT