AMIRA MCP Server

MCP serverSearch

Read-only access to AMIRA, the Africa Multiple research-data platform on Omeka S.

Use AMIRA MCP Server in Claude, ChatGPT or Ahel Desktop

Free. Sign in, add AMIRA MCP Server and connect your AI. About a minute.

Also: Claude Code · Cursor · Codex

Then ask your AI: use AMIRA MCP Server

Details

Available today. Use it from your connected AI after setup.

Add ahel to your AI once: Claude, ChatGPT, Cursor, Claude Code or Codex. Then ask it to use this.

AMIRA MCP ServerStart free

Install AMIRA MCP Server

The server’s own address, for the clients that take one directly. Or connect ahel once and every client you use reads it from one address, with the account kept on ahel rather than in each client’s config.

  • Claude Code

    claude mcp add --transport http --scope user amira-mcp-server 'https://data.africamultiple.uni-bayreuth.de/mcp'

    Run it once in your project, then open /mcp to approve any sign-in the server asks for.

  • Claude Desktop

    https://data.africamultiple.uni-bayreuth.de/mcp

    Add a custom connector in Settings, paste this address, and approve the sign-in.

  • Cursor

    cursor://anysphere.cursor-deeplink/mcp/install?name=amira-mcp-server&config=eyJ1cmwiOiJodHRwczovL2RhdGEuYWZyaWNhbXVsdGlwbGUudW5pLWJheXJldXRoLmRlL21jcCJ9

    Open the link and Cursor adds the server at that address.

  • ChatGPT

    https://data.africamultiple.uni-bayreuth.de/mcp

    In Settings, enable Developer mode, create an MCP app, and paste this address. Your plan and workspace must allow custom apps.

  • Codex

    codex mcp add amira-mcp-server --url 'https://data.africamultiple.uni-bayreuth.de/mcp'

    Run it once, then sign in with codex mcp login amira-mcp-server if the server asks for an account.

From the project's README

As published by AM-Digital-Research-Environment/amira-mcp-server in README.md.

Read-only Model Context Protocol server for the Africa Multiple Interactive Research Atlas (AMIRA), the research-data platform of the Africa Multiple Cluster of Excellence at the University of Bayreuth. AMIRA is published on the cluster's public Omeka S site at data.africamultiple.uni-bayreuth.de.

AMIRA is built and maintained by the Cluster's Digital Research Environment (DRE), its digital infrastructure unit. The DRE designs, builds, and maintains the data systems that connect researchers across the Africa Multiple Research Centres (AMRCs) and partner institutions worldwide. Curation and description are joint efforts with partners at the AMRCs in Université Joseph Ki-Zerbo, Rhodes University, the University of Lagos, and Moi University. Bayreuth hosts the coordinating DRE infrastructure and metadata layer. Federal University of Bahia is a privileged partner, not an AMRC.

Storage stays distributed by default: data remains in its local repository, while Bayreuth holds the metadata layer that points to it. Research data becomes findable without being relocated. High-quality metadata makes researchers' data more discoverable and their work more visible to a wider community of peers.

This server exposes AMIRA's projects, thematic research sections, ~4,000 digitised research items, people, institutions, groups, collections, the cluster bibliography with searchable full text (extracted from the open-access PDFs) and its journals, podcast episodes with transcripts, and the cluster's YouTube videos with searchable transcripts — as 33 core tools an LLM can query, five research-workflow prompts, and resources for citable records and bulk exports. From one MCP interface, clients can move across records and the places, languages, and subjects that connect them.

Every record carries an amira_url — its public page on the Omeka S site (…/s/amira/item/<id>) — so findings can be cited as links back to the source. AMIRA focuses on the Cluster's research data. For news, events, and general information about the Africa Multiple Cluster of Excellence, visit africamultiple.uni-bayreuth.de.

Coverage checked 6 October 2026: 3,975 research items (1,419 with digitised media), 93 projects, 555 publications (60 with extracted full text), 87 journals, 43 podcast episodes, 140 videos and 3,085 subject authorities (646 Library of Congress headings). Counts are a dated snapshot; use get_collection_overview for what your running server actually holds. See the publication guide and the roadmap for further improvements.

Two October 2026 reviews document the current design: the 5 October review and its implementation in 1.18 (research tools, apps, durable snapshots), and the 6 October review and its implementation in 1.19 (matching fixes, prompts, resources, authority identifiers, media and hardening).

How it gets its data — and why nothing else is needed

The server is self-contained and offline-first. It reads from two sources:

  1. A complete data snapshot bundled inside the .mcpb — crawled from the public Omeka S REST API at build time and transformed into compact typed records, so the server works fully offline with zero setup.
  2. (optional, on by default) the Omeka S API over HTTPS. At startup and daily, probes compare item and item-set modification signatures and totals. A changed signal triggers a complete crawl; a weekly forced crawl also catches vocabulary-only edits. A writer lock, immutable generations and an atomic pointer preserve the previous usable snapshot if refresh fails. The API origin and installation path partition the cache and must match the manifest before its records can be served under that site's citation URLs.

End users need no API key, no credentials, and no VPN — it's the same openly published data that powers the public site.

Tools

Call get_collection_overview first to scope the data, then drill in.

ToolPurpose
get_collection_overviewCounts and breakdowns across the whole collection + snapshot freshness
search_research_itemsFind items by keyword (every word must match; quote a phrase), subject, location (hierarchy-aware), country (exact, with aliases such as Côte d'Ivoire = Ivory Coast), contributor, project, section, university, resource type, format/genre, language, year, has_media, added_since / modified_since; `export=csv
get_research_itemFull metadata for one item (by Omeka id or typed id): typed dates, roles with the affiliation at the time, places with their region/country chain, provenance as linked institutions, typed identifiers, collections, related items, media files and the IIIF manifest — plus a generated citation and a BibTeX/RIS/CSL-JSON export (citation_format)
search_projects / get_projectProjects by keyword or acronym, university, section, PI, member, funder — detail with item breakdown, top subjects and sample items
list_research_sections / get_research_sectionThematic sections with funding phases (AM 1.0 / AM 2.0), PIs, counts, projects; detail by name or id
search_persons / get_personPeople (either name order works) — profile with GND and other authority identifiers, projects, items, publications and top collaborators
list_institutions / get_institution / list_cluster_partners / list_groupsOrganisations (institutions, Africa Multiple partner categories, and research groups) by name, acronym or id, their projects, people and items
list_subjectsSubject headings ranked by item frequency, each marked as a Library of Congress heading (with its id.loc.gov URI) or a free tag (`vocabulary=lcsh
list_locationsEvery place — countries and cities in one flat list (hierarchy rolled up) — ranked by item count, with coordinates and Wikidata ids; filter by country, name, near (radius) or bbox
list_collectionsCollections (item sets) ranked by research-item count, with IIIF collection links — pair with the collection filter
list_categoriesFacet values: formats/genres, languages, resource types
list_yearsDate histogram of research items by year or decade — coverage over time, most-covered year/decade
search_publications / get_publicationFilter the bibliography by author, subject, language, type, venue, year, date added and full-text availability; search abstracts in every language and extracted PDF text. Detail adds every abstract with its language, linked authorities, conference/extent/access/thesis metadata. citation_format selects BibTeX/RIS/CSL-JSON per page; export links the whole result as one file
list_publication_facetsCounts by type, year, language, subject, author/editor or venue across the complete filtered bibliography; ranked and paginated
list_journalsThe journals the cluster publishes in, ranked by publication count, with ISSN and country — pair with the venue filter
find_relatedCross-entity discovery: pivot from a subject/place/person/project to co-occurring entities (incl. publications)
search_podcasts / get_podcastCluster podcast episodes with searchable transcripts, filterable by language; detail gives the duration, audio file and the model that generated the transcript; transcript text is opt-in
search_videos / get_videoThe cluster's YouTube videos — full-text search over transcripts (match snippets; thumbnail; transcript opt-in on detail)
resolve_entityResolve names or typed IDs (either vocabulary: item:7392 = research_item:7392); return separate candidates for homonyms and unreconciled literals
get_entity_graphBounded one-hop graph with distinct explicit links/co-occurrences and paginated cited evidence
get_text_passagesFind passages in selected publication/video/podcast records, with exact original-text offsets
compare_collectionsCompare 2–4 projects or collections using common filters, denominators and missingness
get_data_qualitySnapshot coverage, missing metadata and unresolved-reference counts
get_snapshot_changesPage added, updated or deleted records across retained snapshots from the same instance

Prompts and resources

Five prompts package multi-step research workflows. Hosts show them as slash commands; their arguments autocomplete from the snapshot's own authority names, and each restates the citation rules. They are listed through prompts/list, so they add nothing to the per-turn tool payload.

PromptArgumentsWorkflow
literature_reviewtopic, year_from?, year_to?Facets, publications, full-text passages and related research data on a topic
project_dossierprojectTeam, holdings, periods, places and connections of one project
person_profilenameDisambiguation, profile, graph and publication list of one person
place_reportplaceItems, projects, subjects and periods for a country, city or region
transcript_evidencequeryQuoted passages from transcripts and open-access full texts

Resources expose data outside the tool calls:

URIContent
amira://record/{kind}/{id}Any record as citable JSON — the same projection as its get_* tool; kinds follow resolve_entity ids (research_item, publication, person, location, subject, …)
amira://export/{corpus}/{format}/{query}A filtered export (CSV/JSONL for research items; CSV/JSONL/BibTeX/RIS/CSL-JSON for publications). The search tools return these links as resource_links when called with export
amira://datasetA schema.org Dataset description of the served snapshot: counts, coverage, licence mix, funder and how to cite it

Profiles and research workflows

AMIRA_TOOL_PROFILE=full is the default: 33 core tools, or 35 over HTTP. research, discovery and visualization expose smaller documented subsets (22/12/14 core tools respectively; HTTP adds search and fetch). Profiles are fixed at startup; they reduce discovery cost, not access permissions. Tools outside a profile are removed, and so are the apps that render them. See src/toolProfiles.ts for the exact lists.

Typed ids work everywhere in either vocabulary (item: / research_item:, pub: / publication:, section:, person:, organisation:, project:, video:, podcast:), including the id of every get_* tool. Paged results stop early, with response_limited: true and a correct next_offset, when a page would exceed 40,000 characters — Claude Code moves larger tool results out of the conversation. Error results carry the error as JSON text only.

Use resolve_entity before graph traversal; pass its typed id unchanged to get_entity_graph. Graph edges count distinct source records, separate each corpus, and distinguish catalogue links from co-occurrence. Neither co-occurrence nor a shared subject establishes collaboration. Follow an edge with edge_id and the returned snapshot_id for stable evidence paging. Graphs cap at 100 nodes, 200 edges and 42,000 JSON bytes, so check truncated.

get_text_passages takes 1–10 publication:ID (or pub:ID), video:ID or podcast:ID identifiers. It returns at most 20 passages per page, original UTF-16 offsets and source citations; every match is reachable through next_offset (scanned_matches_capped reports a document with more than 10,000 matches). These are extracted-text offsets, not PDF page numbers or audio timestamps. get_snapshot_changes requires two local same-source generations; it reports available history instead of inventing a diff.

Example questions it can answer

You ask…The model uses…
"What does the collection hold on Islam?"list_subjects keyword=Islam → search_research_items subject=Islam
"Show me all the French-language items"search_research_items language=French (codes fr/fra/legacy fre work too)
"What has Ulli Beier contributed?"get_person name="Ulli Beier" (resolves to 'Beier, Ulli')
"Which items come from Nigeria?"search_research_items location=Nigeria (Lagos items count too — location walks the place hierarchy) or country=Nigeria (exact: Niger stays out)
"What does AMIRA hold from Côte d'Ivoire?"search_research_items country="Côte d'Ivoire" (the authority stores "Ivory Coast"; common aliases resolve)
"What has been added since September?"search_research_items added_since=2026-09-01 (or search_publications added_since=…)
"Which Lagos items have digitised images?"search_research_items location=Lagos has_media=true → get_research_item for the files and IIIF manifest
"Give me all Islam-related items as a spreadsheet"search_research_items subject=Islam export=csv → read the returned amira://export/… link
"What audio recordings are in the ILAM collection?"search_research_items project_id=37700 resource_type=Audio
"In which talks does anyone discuss decoloniality?"search_videos keyword=decolonial (matches inside transcripts, flagged matched_in)
"Which cluster publications discuss migration control — and what do they actually say?"search_publications keyword="migration control" (matches inside the extracted full text, flagged matched_in) → get_publication include_fulltext=true
"How many French-language publications are there, and of which types?"list_publication_facets facet=type language=fr → search_publications language=fr
"Export the French-language bibliography for my reference manager"search_publications language=fr citation_format=ris → follow next_offset; combine the complete ris records
"Which journals does the cluster publish in?"list_journals (ranked by publication count)
"What themes travel with Architecture across projects?"find_related entity_type=subject value=Architecture
"When was this photograph taken?"get_research_item → typed dates (created/collected/issued/…)
"Give me a citation for this item — and the BibTeX"get_research_item → generated_citation + bibtex (citation_format=ris/csl-json for a reference manager)
"Which decade does the collection cover most?"list_years bucket=decade sort=count

Companion skill

A research-workflow skill ships in .claude/skills/amira-mcp/: cluster context, a query workflow, a tool-by-task map, the citation discipline, and coverage caveats. Install it either way:

  • Download amira-mcp-skill.zip from the releases page and unzip it into your Claude skills directory (e.g. ~/.claude/skills/) — it expands to ~/.claude/skills/amira-mcp/.
  • Or copy the .claude/skills/amira-mcp/ folder from this repo there.

…or let the server hand it over (Skills over MCP)

Official extension; host support varies. SEP-2640 is now Final. The server follows the published Skills extension, including SHA-256 digests and byte sizes for every file. The zip above remains available for clients without extension support. The research tools and citation contract work either way.

The server also serves that same skill over the connection, following the SEP-2640 extension (io.modelcontextprotocol/skills). Nothing to download and nothing to keep in sync: a host that implements the extension discovers the skill on connect, and the copy it gets is the one built into the running server rather than a zip that may predate the tool surface it describes. It is the only route that reaches the remote HTTP surface — ChatGPT, Claude.ai connectors and the APIs have no local skills directory.

MethodWhat it returns
skills/listThe catalog: skill://amira-mcp/SKILL.md, its frontmatter, and a SHA-256 digest and byte size for every file
skills/getOne skill by URI, to refresh digests without re-listing
resources/readAny single file — SKILL.md or a reference — read on demand
resources/directory/readOne directory level at a time (skill://amira-mcp, skill://amira-mcp/references)

Cost when unused: zero. Skill text is never injected into the server instructions or into any tool description, so the tools/list payload — the part re-sent every turn — is byte-identical whether or not a host supports the extension. Disclosure timing stays a host decision; the three reference files load only when something actually reads them.

Set AMIRA_SKILLS=0 to drop the capability and the three methods entirely.

Install (end users)

Download amira-mcp-server.mcpb from the releases page and double-click it. Claude Desktop shows an install dialog; click Install. No further configuration is required. The latest tagged release carries the current data; the rolling data-latest pre-release always tracks the freshest site snapshot.

Use it from ChatGPT, the API, or any remote client

The .mcpb is the local, offline option for Claude Desktop. The same server can also run as a remote Streamable HTTP endpoint — one HTTPS URL that ChatGPT, Claude (web + desktop remote connectors), the OpenAI and Anthropic APIs, Cursor, VS Code and other clients connect to by pasting a URL (no download, always-fresh data). The remote surface serves the same 33 tools plus the OpenAI-compatible search / fetch tools for ChatGPT research integrations (35 total). Access is unauthenticated — the data is public and read-only.

search takes plain keywords (matched term-by-term, not as an exact phrase), with optional limit and types to keep the result set tight, and returns ranked hits across research items, the bibliography (reaching into publication full text), podcasts, videos, projects and research sections. The url on search results and fetch documents is always the AMIRA/Omeka public record page; DOI, YouTube/watch, and podcast/listen URLs are kept as secondary metadata/text links. fetch returns one record's full text by the id search hands back — for videos and podcasts the transcript is omitted by default (metadata + description only, since a full one can run to tens of thousands of characters), and include_transcript=true pulls it in — paged with transcript_offset / transcript_max_chars (the same names get_video / get_podcast use). Publication full text works the same way (include_fulltext=true, paged with fulltext_offset / fulltext_max_chars, matching get_publication), and max_chars caps the whole text body. The appended window is sized against what max_chars leaves after the metadata header, so *_returned_chars is exactly what landed in text and the next page starts at offset + returned_chars with no gap.

npm run build && npm run start:http     # → http://localhost:8787/mcp
# or, self-contained, via Docker:
docker build -t amira-mcp . && docker run -p 8787:8787 amira-mcp

Endpoints: POST /mcp (the MCP endpoint) and GET /healthz. Bind with PORT / HOST.

  • ChatGPT → enable Developer mode under Settings → Security and login, then add your server URL from ChatGPT Plugins. Research integrations use search + fetch; see the current OpenAI setup guide.

  • Claude (web or desktop) → Settings → Connectors → Add custom connector → https://<your-host>/mcp.

  • OpenAI API (Responses) — point the mcp tool at the endpoint:

    { "type": "mcp", "server_label": "amira",
      "server_url": "https://<your-host>/mcp",
      "allowed_tools": ["search", "fetch"], "require_approval": "never" }
    

Deploy (self-hosted, e.g. alongside the amira site)

The multi-stage Dockerfile crawls a fresh snapshot at build time and runs the self-contained bundle (no node_modules at runtime). On a Linux host you can equally run it under systemd behind a reverse proxy:

# /etc/systemd/system/amira-mcp.service
[Service]
ExecStart=/usr/bin/node /opt/amira-mcp/server/http.js
Environment=PORT=8787 HOST=127.0.0.1 AMIRA_LIVE_REFRESH=true
Restart=always
User=www-data
location /mcp { proxy_pass http://127.0.0.1:8787/mcp; proxy_buffering off; client_max_body_size 64k; }

proxy_buffering off keeps the Streamable-HTTP/SSE responses flowing. The refresh uses AMIRA_SITE_BASE (the public HTTPS API by default), even when co-located with Omeka. It performs a full crawl when changes are detected; co-location does not eliminate API load. The server caps request bodies at 64 KiB and rate-limit state at 10,000 clients. Configure connection, request-rate and concurrency limits at the reverse proxy for public deployments; the in-process limiter is a courtesy control. Shutdown cancels refresh work and gives HTTP connections a ten-second drain window.

For a build that consumes a previously reviewed data/ snapshot without calling Omeka during the build:

docker build --build-arg SNAPSHOT_STAGE=bundled -t amira-mcp .

The default SNAPSHOT_STAGE=fetch still crawls fresh data. Both stages validate the snapshot. The image pins its base digest; npm ci uses the committed lockfile.

Develop / rebuild

npm ci
npm run fetch-data    # crawl the public Omeka API -> ./data snapshot (~1 min)
npm run typecheck     # tsc --noEmit
npm run build         # esbuild -> server/{index,http,fetchCli,lib}.js
npm test              # unit tests: transform fixtures, folding, snapshot + store lifecycle,
                      # and the full tool layer against a fixture snapshot via
                      # InMemoryTransport (offline)
npm run test:live     # integration tests against the live API (network)
npm run smoke         # test/smoke/stdio.mjs: spawn the stdio server, exercise every tool offline
npm run smoke:http    # test/smoke/http.mjs: spawn the HTTP server on a free port: search/fetch +
                      # parity, CORS preflight for both protocol revisions, the rate limiter
npm run weigh         # token budget report (needs ./data — run fetch-data first)
npm run benchmark     # offline latency baseline against ./data (30 warm samples per query)

Token budgets

Shortened here. Read the whole README on GitHub.

Advanced
Delivery
amira-mcp-server MCP server → your ahel connector (mcp.ahel.ai) → your AI.
Item type
mcp-server
Key
io-github-am-digital-research-environment-amira-mcp-server
Source
github.com/AM-Digital-Research-Environment/amira-mcp-server
Hosted endpoint
https://data.africamultiple.uni-bayreuth.de/mcp