@cyanheads/wikipedia-mcp-server

MCP serverSearch

Search Wikipedia, read summaries and full text, target sections, find nearby pages, list languages.

Available today. Use it from your connected AI after setup.

Connect ahel once, and every AI you use reads what you have installed.

Then ask your AI: use @cyanheads/wikipedia-mcp-server

From the project's README

As published by cyanheads/wikipedia-mcp-server in README.md.

Public Hosted Server: https://wikipedia.caseyjhand.com/mcp


Overview

Wikipedia content via the MediaWiki REST API and Action API. Search articles, read summaries or targeted sections, find geotagged pages near a coordinate, and list language editions from any MCP client. Runs as a stdio process, a local Streamable HTTP server, or the public hosted endpoint above.

Tools

ToolDescription
wikipedia_search_articlesFull-text search across Wikipedia, returning ranked results with plain-text snippets and page IDs.
wikipedia_get_summaryLead-section summary for any article — plain text, Wikidata QID, description, thumbnail URL, and page type.
wikipedia_get_articleFull article or a targeted section as clean plain text, with section markers preserved.
wikipedia_get_sectionsTable of contents with section_index values for targeted section reads.
wikipedia_search_nearbyGeotagged Wikipedia articles within a radius of a WGS 84 coordinate, sorted by distance.
wikipedia_get_languagesAll language editions available for an article, with titles and URLs.

Capability reference

wikipedia_search_articles tool

  • Free-text query, ranked by relevance; returns plain-text snippets (HTML stripped), page IDs, and word counts
  • limit capped at 50; offset pages further results — enrichment nextOffset signals more remain and is passed back as offset
  • language selects any Wikipedia edition (default en)
  • Best when the exact article title is unknown, or to discover multiple articles on a topic

wikipedia_get_summary tool

  • Returns the 2–4 paragraph lead extract, Wikidata QID (wikibase_item), short description, and thumbnail URL
  • page_type discriminates standard / disambiguation / no-extract — on disambiguation, re-query with wikipedia_search_articles for a more specific title
  • Redirect pages are followed automatically
  • Right tool for most encyclopedic "what is X?" lookups; use wikipedia_get_article for full depth

wikipedia_get_article tool

  • Without section_index: full article with == Section == markers, unless it exceeds WIKIPEDIA_ARTICLE_OVERFLOW_BYTES (default 80,000 bytes) — then returns a section outline (truncated: true) pointing to wikipedia_get_sections plus a targeted section_index read
  • With section_index (from wikipedia_get_sections): returns that section plus every nested subsection, each heading above its own body
  • Data tables are omitted from both paths — a section whose body is entirely a data table returns little beyond its heading; layout-only tables (multi-column lists, succession boxes) keep their content
  • Page furniture — maintenance banners, sister-project and library-resource boxes, portal bars, spoken-article notices — is stripped; hatnotes are kept
  • Redirect pages are followed automatically

wikipedia_get_sections tool

  • Returns section titles, heading levels, hierarchical numbering (e.g. "2.1"), and section_index values
  • section_index is the integer to pass to wikipedia_get_article for a targeted read
  • Fails with no_sections on a stub or very short article — read it with wikipedia_get_article instead
  • Redirect pages are followed automatically

wikipedia_search_nearby tool

  • Returns geotagged articles sorted ascending by distance, with coordinates and distance_meters
  • radius_meters: 10–10,000 (default 1000); limit: 1–500 (default 10) — no pagination past limit, so raise it or sweep narrower radii for full coverage
  • Only articles with a geographic coordinate in their Wikidata record are returned
  • Enrichment truncated flags when more articles matched than limit allowed

wikipedia_get_languages tool

  • Returns each edition's language_code, tool-usable edition_code (can differ, e.g. gsw vs als), article title, and URL
  • Pass edition_code — not language_code — as the language parameter on other tools
  • Fails with no_other_languages when the article has no translations
  • Redirect pages are followed automatically; source_title reports the resolved title

Features

Built on @cyanheads/mcp-ts-core: stdio and Streamable HTTP transports, pluggable auth (none / jwt / oauth), swappable storage (in-memory, filesystem, Supabase, Cloudflare KV/R2/D1), structured logging with optional OpenTelemetry tracing.

Wikipedia-specific:

  • Dual API integration — MediaWiki REST API (/api/rest_v1/) for summaries, Action API (/w/api.php) for search, full text, sections, geo search, and language links
  • Retry and backoff on all requests; User-Agent header per Wikimedia API policy
  • Both read paths render to the same plain-text shape — == Heading == markers, one list item per line — the full article from Action API extracts, a section from the parser's own HTML for that section. A section read additionally keeps code-sample indentation and the lists inside layout tables, neither of which the extract carries
  • Per-call language parameter on every tool — all Wikipedia language editions accessible in a single session
  • Language validation against a live edition registry built from the MediaWiki action=sitematrix endpoint (cached 24h) — catches structurally valid but nonexistent editions before they cause timeouts

Agent-friendly output:

  • page_type on summaries discriminates standard / disambiguation / no-extract — no string parsing needed
  • wikibase_item (Wikidata QID) on summaries enables direct cross-referencing with wikidata-mcp-server
  • section_index on table-of-contents entries links directly to the targeted-read parameter on wikipedia_get_article
  • Recovery hints on every error type — callers get actionable next steps (e.g., "use wikipedia_search_articles to find the correct title")

Getting started

Public Hosted Instance

A public instance is available at https://wikipedia.caseyjhand.com/mcp — no installation required. Point any MCP client at it via Streamable HTTP:

{
  "mcpServers": {
    "wikipedia-mcp-server": {
      "type": "streamable-http",
      "url": "https://wikipedia.caseyjhand.com/mcp"
    }
  }
}

Self-Hosted / Local

Add the following to your MCP client configuration file.

{
  "mcpServers": {
    "wikipedia-mcp-server": {
      "type": "stdio",
      "command": "bunx",
      "args": ["@cyanheads/wikipedia-mcp-server@latest"],
      "env": {
        "MCP_TRANSPORT_TYPE": "stdio",
        "MCP_LOG_LEVEL": "info"
      }
    }
  }
}

Or with npx (no Bun required):

{
  "mcpServers": {
    "wikipedia-mcp-server": {
      "type": "stdio",
      "command": "npx",
      "args": ["-y", "@cyanheads/wikipedia-mcp-server@latest"],
      "env": {
        "MCP_TRANSPORT_TYPE": "stdio",
        "MCP_LOG_LEVEL": "info"
      }
    }
  }
}

Or with Docker:

{
  "mcpServers": {
    "wikipedia-mcp-server": {
      "type": "stdio",
      "command": "docker",
      "args": [
        "run", "-i", "--rm",
        "-e", "MCP_TRANSPORT_TYPE=stdio",
        "ghcr.io/cyanheads/wikipedia-mcp-server:latest"
      ]
    }
  }
}

For Streamable HTTP, set the transport and start the server:

MCP_TRANSPORT_TYPE=http MCP_HTTP_PORT=3010 bun run start:http
# Server listens at http://localhost:3010/mcp

Prerequisites

  • Bun v1.3.0 or higher (or Node.js v24+).
  • No API keys required — Wikipedia's API is public.

Installation

  1. Clone the repository:
git clone https://github.com/cyanheads/wikipedia-mcp-server.git
  1. Navigate into the directory:
cd wikipedia-mcp-server
  1. Install dependencies:
bun install
  1. Configure environment (optional):
cp .env.example .env
# edit .env if you want to customize WIKIPEDIA_USER_AGENT or logging

Configuration

VariableDescriptionDefault
WIKIPEDIA_USER_AGENTUser-Agent header sent with every Wikimedia API request. Customize for your deployment.wikipedia-mcp-server/0.2.0 (https://github.com/cyanheads/wikipedia-mcp-server)
WIKIPEDIA_BASE_URLOptional single-instance override. Unset (default): compose per-language hosts, language selects the edition per call. Set to a full base URL (e.g. a private MediaWiki mirror): route every call at that one fixed host — language no longer varies it.(unset)
WIKIPEDIA_ARTICLE_OVERFLOW_BYTESByte budget above which a full-article read (wikipedia_get_article without section_index) returns a section outline instead of the full text. Tuned for this domain — ordinary articles stay whole; only genuine mega-articles (World War II ~86 KB, United States ~94 KB) outline. Section-targeted reads are never affected.80000
MCP_TRANSPORT_TYPETransport: stdio or http.stdio
MCP_HTTP_PORTPort for HTTP server.3010
MCP_SESSION_MODEHTTP session mode: stateless, stateful, or auto (which resolves to stateful). The Docker image ships stateless.auto
MCP_AUTH_MODEAuth mode: none, jwt, or oauth.none
MCP_LOG_LEVELLog level (RFC 5424).info
LOGS_DIRDirectory for log files (Node.js only).<project-root>/logs
OTEL_ENABLEDEnable OpenTelemetry instrumentation (spans, metrics, completion logs).false

See .env.example for the full list of optional overrides.

Running the server

Local development

  • Build and run:

    # One-time build
    bun run rebuild
    
    # Run the built server
    bun run start:stdio
    # or
    bun run start:http
    
  • Run checks and tests:

    bun run devcheck   # Lint, format, typecheck, security
    bun run test       # Vitest test suite
    bun run lint:mcp   # Validate MCP definitions against spec
    

Docker

docker build -t wikipedia-mcp-server .
docker run --rm -p 3010:3010 wikipedia-mcp-server

The Dockerfile defaults to HTTP transport, stateless session mode, and logs to /var/log/wikipedia-mcp-server. OpenTelemetry peer dependencies are installed by default — build with --build-arg OTEL_ENABLED=false to omit them.

Project structure

DirectoryPurpose
src/index.tscreateApp() entry point — registers tools and inits the Wikipedia service.
src/configServer-specific environment variable parsing and validation with Zod.
src/mcp-server/toolsTool definitions (*.tool.ts) — one file per tool.
src/services/wikipediaWikipediaService — REST API + Action API client with retry/backoff and language validation.
tests/Unit and integration tests mirroring src/.

Development guide

See CLAUDE.md for development guidelines and architectural rules. The short version:

  • Handlers throw, framework catches — no try/catch in tool logic
  • Use ctx.log for request-scoped logging, ctx.state for tenant-scoped storage
  • Register new tools in src/mcp-server/tools/definitions/index.ts
  • Wrap external API calls: validate raw → normalize to domain type → return output schema; never fabricate missing fields

Contributing

Issues are welcome. Run checks and tests before submitting:

bun run devcheck
bun run test

License

Apache-2.0 — see LICENSE for details.

Advanced
Delivery
wikipedia-mcp-server MCP server → your ahel gateway (mcp.ahel.ai) → every connected AI client.
Catalog kind
mcp-server
Gateway key
io-github-cyanheads-wikipedia-mcp-server
Source
github.com/cyanheads/wikipedia-mcp-server
Hosted endpoint
https://wikipedia.caseyjhand.com/mcp