AMIRA MCP Server
MCP serverSearchRead-only access to AMIRA, the Africa Multiple research-data platform on Omeka S.
Use AMIRA MCP Server in Claude, ChatGPT or Ahel Desktop
Free. Sign in, add AMIRA MCP Server and connect your AI. About a minute.
Also: Claude Code · Cursor · Codex
Then ask your AI: use AMIRA MCP Server
No other account needed.
Details
Available today. Use it from your connected AI after setup.
No other account needed.
Add ahel to your AI once: Claude, ChatGPT, Cursor, Claude Code or Codex. Then ask it to use this.
Install AMIRA MCP Server
The server’s own address, for the clients that take one directly. Or connect ahel once and every client you use reads it from one address, with the account kept on ahel rather than in each client’s config.
Claude Code
claude mcp add --transport http --scope user amira-mcp-server 'https://data.africamultiple.uni-bayreuth.de/mcp'Run it once in your project, then open /mcp to approve any sign-in the server asks for.
Claude Desktop
https://data.africamultiple.uni-bayreuth.de/mcpAdd a custom connector in Settings, paste this address, and approve the sign-in.
Cursor
cursor://anysphere.cursor-deeplink/mcp/install?name=amira-mcp-server&config=eyJ1cmwiOiJodHRwczovL2RhdGEuYWZyaWNhbXVsdGlwbGUudW5pLWJheXJldXRoLmRlL21jcCJ9Open the link and Cursor adds the server at that address.
ChatGPT
https://data.africamultiple.uni-bayreuth.de/mcpIn Settings, enable Developer mode, create an MCP app, and paste this address. Your plan and workspace must allow custom apps.
Codex
codex mcp add amira-mcp-server --url 'https://data.africamultiple.uni-bayreuth.de/mcp'Run it once, then sign in with codex mcp login amira-mcp-server if the server asks for an account.
From the project's README
As published by AM-Digital-Research-Environment/amira-mcp-server in README.md.
Read-only Model Context Protocol server for the Africa Multiple Interactive Research Atlas (AMIRA), the research-data platform of the Africa Multiple Cluster of Excellence at the University of Bayreuth. AMIRA is published on the cluster's public Omeka S site at data.africamultiple.uni-bayreuth.de.
AMIRA is built and maintained by the Cluster's Digital Research Environment (DRE), its digital infrastructure unit. The DRE designs, builds, and maintains the data systems that connect researchers across the Africa Multiple Research Centres (AMRCs) and partner institutions worldwide. Curation and description are joint efforts with partners at the AMRCs in Université Joseph Ki-Zerbo, Rhodes University, the University of Lagos, and Moi University. Bayreuth hosts the coordinating DRE infrastructure and metadata layer. Federal University of Bahia is a privileged partner, not an AMRC.
Storage stays distributed by default: data remains in its local repository, while Bayreuth holds the metadata layer that points to it. Research data becomes findable without being relocated. High-quality metadata makes researchers' data more discoverable and their work more visible to a wider community of peers.
This server exposes AMIRA's projects, thematic research sections, ~4,000 digitised research items, people, institutions, groups, collections, the cluster bibliography with searchable full text (extracted from the open-access PDFs) and its journals, podcast episodes with transcripts, and the cluster's YouTube videos with searchable transcripts — as 33 core tools an LLM can query, five research-workflow prompts, and resources for citable records and bulk exports. From one MCP interface, clients can move across records and the places, languages, and subjects that connect them.
Every record carries an amira_url — its public page on the Omeka S site
(…/s/amira/item/<id>) — so findings can be cited as links back to the source.
AMIRA focuses on the Cluster's research data. For news, events, and general
information about the Africa Multiple Cluster of Excellence, visit
africamultiple.uni-bayreuth.de.
Coverage checked 6 October 2026: 3,975 research items (1,419 with digitised
media), 93 projects, 555 publications (60 with extracted full text), 87 journals,
43 podcast episodes, 140 videos and 3,085 subject authorities (646 Library of
Congress headings). Counts are a dated snapshot; use get_collection_overview for
what your running server actually holds. See the publication guide
and the roadmap for further improvements.
Two October 2026 reviews document the current design: the 5 October review and its implementation in 1.18 (research tools, apps, durable snapshots), and the 6 October review and its implementation in 1.19 (matching fixes, prompts, resources, authority identifiers, media and hardening).
How it gets its data — and why nothing else is needed
The server is self-contained and offline-first. It reads from two sources:
- A complete data snapshot bundled inside the
.mcpb— crawled from the public Omeka S REST API at build time and transformed into compact typed records, so the server works fully offline with zero setup. - (optional, on by default) the Omeka S API over HTTPS. At startup and daily, probes compare item and item-set modification signatures and totals. A changed signal triggers a complete crawl; a weekly forced crawl also catches vocabulary-only edits. A writer lock, immutable generations and an atomic pointer preserve the previous usable snapshot if refresh fails. The API origin and installation path partition the cache and must match the manifest before its records can be served under that site's citation URLs.
End users need no API key, no credentials, and no VPN — it's the same openly published data that powers the public site.
Tools
Call get_collection_overview first to scope the data, then drill in.
| Tool | Purpose |
|---|---|
get_collection_overview | Counts and breakdowns across the whole collection + snapshot freshness |
search_research_items | Find items by keyword (every word must match; quote a phrase), subject, location (hierarchy-aware), country (exact, with aliases such as Côte d'Ivoire = Ivory Coast), contributor, project, section, university, resource type, format/genre, language, year, has_media, added_since / modified_since; `export=csv |
get_research_item | Full metadata for one item (by Omeka id or typed id): typed dates, roles with the affiliation at the time, places with their region/country chain, provenance as linked institutions, typed identifiers, collections, related items, media files and the IIIF manifest — plus a generated citation and a BibTeX/RIS/CSL-JSON export (citation_format) |
search_projects / get_project | Projects by keyword or acronym, university, section, PI, member, funder — detail with item breakdown, top subjects and sample items |
list_research_sections / get_research_section | Thematic sections with funding phases (AM 1.0 / AM 2.0), PIs, counts, projects; detail by name or id |
search_persons / get_person | People (either name order works) — profile with GND and other authority identifiers, projects, items, publications and top collaborators |
list_institutions / get_institution / list_cluster_partners / list_groups | Organisations (institutions, Africa Multiple partner categories, and research groups) by name, acronym or id, their projects, people and items |
list_subjects | Subject headings ranked by item frequency, each marked as a Library of Congress heading (with its id.loc.gov URI) or a free tag (`vocabulary=lcsh |
list_locations | Every place — countries and cities in one flat list (hierarchy rolled up) — ranked by item count, with coordinates and Wikidata ids; filter by country, name, near (radius) or bbox |
list_collections | Collections (item sets) ranked by research-item count, with IIIF collection links — pair with the collection filter |
list_categories | Facet values: formats/genres, languages, resource types |
list_years | Date histogram of research items by year or decade — coverage over time, most-covered year/decade |
search_publications / get_publication | Filter the bibliography by author, subject, language, type, venue, year, date added and full-text availability; search abstracts in every language and extracted PDF text. Detail adds every abstract with its language, linked authorities, conference/extent/access/thesis metadata. citation_format selects BibTeX/RIS/CSL-JSON per page; export links the whole result as one file |
list_publication_facets | Counts by type, year, language, subject, author/editor or venue across the complete filtered bibliography; ranked and paginated |
list_journals | The journals the cluster publishes in, ranked by publication count, with ISSN and country — pair with the venue filter |
find_related | Cross-entity discovery: pivot from a subject/place/person/project to co-occurring entities (incl. publications) |
search_podcasts / get_podcast | Cluster podcast episodes with searchable transcripts, filterable by language; detail gives the duration, audio file and the model that generated the transcript; transcript text is opt-in |
search_videos / get_video | The cluster's YouTube videos — full-text search over transcripts (match snippets; thumbnail; transcript opt-in on detail) |
resolve_entity | Resolve names or typed IDs (either vocabulary: item:7392 = research_item:7392); return separate candidates for homonyms and unreconciled literals |
get_entity_graph | Bounded one-hop graph with distinct explicit links/co-occurrences and paginated cited evidence |
get_text_passages | Find passages in selected publication/video/podcast records, with exact original-text offsets |
compare_collections | Compare 2–4 projects or collections using common filters, denominators and missingness |
get_data_quality | Snapshot coverage, missing metadata and unresolved-reference counts |
get_snapshot_changes | Page added, updated or deleted records across retained snapshots from the same instance |
Prompts and resources
Five prompts package multi-step research workflows. Hosts show them as slash
commands; their arguments autocomplete from the snapshot's own authority names, and
each restates the citation rules. They are listed through prompts/list, so they
add nothing to the per-turn tool payload.
| Prompt | Arguments | Workflow |
|---|---|---|
literature_review | topic, year_from?, year_to? | Facets, publications, full-text passages and related research data on a topic |
project_dossier | project | Team, holdings, periods, places and connections of one project |
person_profile | name | Disambiguation, profile, graph and publication list of one person |
place_report | place | Items, projects, subjects and periods for a country, city or region |
transcript_evidence | query | Quoted passages from transcripts and open-access full texts |
Resources expose data outside the tool calls:
| URI | Content |
|---|---|
amira://record/{kind}/{id} | Any record as citable JSON — the same projection as its get_* tool; kinds follow resolve_entity ids (research_item, publication, person, location, subject, …) |
amira://export/{corpus}/{format}/{query} | A filtered export (CSV/JSONL for research items; CSV/JSONL/BibTeX/RIS/CSL-JSON for publications). The search tools return these links as resource_links when called with export |
amira://dataset | A schema.org Dataset description of the served snapshot: counts, coverage, licence mix, funder and how to cite it |
Profiles and research workflows
AMIRA_TOOL_PROFILE=full is the default: 33 core tools, or 35 over HTTP.
research, discovery and visualization expose smaller documented subsets
(22/12/14 core tools respectively; HTTP adds search and fetch). Profiles are
fixed at startup; they reduce discovery cost, not access permissions. Tools
outside a profile are removed, and so are the apps that render them. See
src/toolProfiles.ts for the exact lists.
Typed ids work everywhere in either vocabulary (item: / research_item:,
pub: / publication:, section:, person:, organisation:, project:,
video:, podcast:), including the id of every get_* tool. Paged results
stop early, with response_limited: true and a correct next_offset, when a page
would exceed 40,000 characters — Claude Code moves larger tool results out of
the conversation. Error results carry the error as JSON text only.
Use resolve_entity before graph traversal; pass its typed id unchanged to
get_entity_graph. Graph edges count distinct source records, separate each
corpus, and distinguish catalogue links from co-occurrence. Neither co-occurrence
nor a shared subject establishes collaboration. Follow an edge with edge_id
and the returned snapshot_id for stable evidence paging. Graphs cap at 100 nodes,
200 edges and 42,000 JSON bytes, so check truncated.
get_text_passages takes 1–10 publication:ID (or pub:ID), video:ID or
podcast:ID identifiers. It returns at most 20 passages per page, original UTF-16
offsets and source citations; every match is reachable through next_offset
(scanned_matches_capped reports a document with more than 10,000 matches). These are extracted-text offsets, not
PDF page numbers or audio timestamps. get_snapshot_changes requires two local
same-source generations; it reports available history instead of inventing a diff.
Example questions it can answer
| You ask… | The model uses… |
|---|---|
| "What does the collection hold on Islam?" | list_subjects keyword=Islam → search_research_items subject=Islam |
| "Show me all the French-language items" | search_research_items language=French (codes fr/fra/legacy fre work too) |
| "What has Ulli Beier contributed?" | get_person name="Ulli Beier" (resolves to 'Beier, Ulli') |
| "Which items come from Nigeria?" | search_research_items location=Nigeria (Lagos items count too — location walks the place hierarchy) or country=Nigeria (exact: Niger stays out) |
| "What does AMIRA hold from Côte d'Ivoire?" | search_research_items country="Côte d'Ivoire" (the authority stores "Ivory Coast"; common aliases resolve) |
| "What has been added since September?" | search_research_items added_since=2026-09-01 (or search_publications added_since=…) |
| "Which Lagos items have digitised images?" | search_research_items location=Lagos has_media=true → get_research_item for the files and IIIF manifest |
| "Give me all Islam-related items as a spreadsheet" | search_research_items subject=Islam export=csv → read the returned amira://export/… link |
| "What audio recordings are in the ILAM collection?" | search_research_items project_id=37700 resource_type=Audio |
| "In which talks does anyone discuss decoloniality?" | search_videos keyword=decolonial (matches inside transcripts, flagged matched_in) |
| "Which cluster publications discuss migration control — and what do they actually say?" | search_publications keyword="migration control" (matches inside the extracted full text, flagged matched_in) → get_publication include_fulltext=true |
| "How many French-language publications are there, and of which types?" | list_publication_facets facet=type language=fr → search_publications language=fr |
| "Export the French-language bibliography for my reference manager" | search_publications language=fr citation_format=ris → follow next_offset; combine the complete ris records |
| "Which journals does the cluster publish in?" | list_journals (ranked by publication count) |
| "What themes travel with Architecture across projects?" | find_related entity_type=subject value=Architecture |
| "When was this photograph taken?" | get_research_item → typed dates (created/collected/issued/…) |
| "Give me a citation for this item — and the BibTeX" | get_research_item → generated_citation + bibtex (citation_format=ris/csl-json for a reference manager) |
| "Which decade does the collection cover most?" | list_years bucket=decade sort=count |
Companion skill
A research-workflow skill ships in .claude/skills/amira-mcp/:
cluster context, a query workflow, a tool-by-task map, the citation discipline, and coverage caveats.
Install it either way:
- Download
amira-mcp-skill.zipfrom the releases page and unzip it into your Claude skills directory (e.g.~/.claude/skills/) — it expands to~/.claude/skills/amira-mcp/. - Or copy the
.claude/skills/amira-mcp/folder from this repo there.
…or let the server hand it over (Skills over MCP)
Official extension; host support varies. SEP-2640 is now Final. The server follows the published Skills extension, including SHA-256 digests and byte sizes for every file. The zip above remains available for clients without extension support. The research tools and citation contract work either way.
The server also serves that same skill over the connection, following the
SEP-2640 extension
(io.modelcontextprotocol/skills). Nothing to download and nothing to keep in sync: a host that
implements the extension discovers the skill on connect, and the copy it gets is the one built into
the running server rather than a zip that may predate the tool surface it describes. It is the only
route that reaches the remote HTTP surface — ChatGPT, Claude.ai connectors and the APIs have no
local skills directory.
| Method | What it returns |
|---|---|
skills/list | The catalog: skill://amira-mcp/SKILL.md, its frontmatter, and a SHA-256 digest and byte size for every file |
skills/get | One skill by URI, to refresh digests without re-listing |
resources/read | Any single file — SKILL.md or a reference — read on demand |
resources/directory/read | One directory level at a time (skill://amira-mcp, skill://amira-mcp/references) |
Cost when unused: zero. Skill text is never injected into the server instructions or into any
tool description, so the tools/list payload — the part re-sent every turn — is byte-identical
whether or not a host supports the extension. Disclosure timing stays a host decision; the three
reference files load only when something actually reads them.
Set AMIRA_SKILLS=0 to drop the capability and the three methods entirely.
Install (end users)
Download amira-mcp-server.mcpb from the
releases page
and double-click it. Claude Desktop shows an install dialog; click Install.
No further configuration is required. The latest tagged release carries the
current data; the rolling data-latest pre-release always tracks the freshest
site snapshot.
Use it from ChatGPT, the API, or any remote client
The .mcpb is the local, offline option for Claude Desktop. The same server can
also run as a remote Streamable HTTP endpoint — one HTTPS URL that ChatGPT,
Claude (web + desktop remote connectors), the OpenAI and Anthropic APIs, Cursor,
VS Code and other clients connect to by pasting a URL (no download, always-fresh
data). The remote surface serves the same 33 tools plus the
OpenAI-compatible search / fetch tools for ChatGPT research integrations (35
total). Access is unauthenticated — the data is public and read-only.
search takes plain keywords (matched term-by-term, not as an exact phrase),
with optional limit and types to keep the result set tight, and returns
ranked hits across research items, the bibliography (reaching into publication
full text), podcasts, videos, projects and research sections. The url on
search results and fetch documents is
always the AMIRA/Omeka public record page; DOI, YouTube/watch, and podcast/listen
URLs are kept as secondary metadata/text links. fetch returns one record's full
text by the id search hands back — for videos and podcasts the transcript is omitted by default
(metadata + description only, since a full one can run to tens of thousands of
characters), and include_transcript=true pulls it in — paged with
transcript_offset / transcript_max_chars (the same names get_video /
get_podcast use). Publication full text works the same way
(include_fulltext=true, paged with fulltext_offset / fulltext_max_chars,
matching get_publication), and max_chars caps the whole text body. The
appended window is sized against what max_chars leaves after the metadata
header, so *_returned_chars is exactly what landed in text and the next page
starts at offset + returned_chars with no gap.
npm run build && npm run start:http # → http://localhost:8787/mcp
# or, self-contained, via Docker:
docker build -t amira-mcp . && docker run -p 8787:8787 amira-mcp
Endpoints: POST /mcp (the MCP endpoint) and GET /healthz. Bind with PORT /
HOST.
-
ChatGPT → enable Developer mode under Settings → Security and login, then add your server URL from ChatGPT Plugins. Research integrations use
search+fetch; see the current OpenAI setup guide. -
Claude (web or desktop) → Settings → Connectors → Add custom connector →
https://<your-host>/mcp. -
OpenAI API (Responses) — point the
mcptool at the endpoint:{ "type": "mcp", "server_label": "amira", "server_url": "https://<your-host>/mcp", "allowed_tools": ["search", "fetch"], "require_approval": "never" }
Deploy (self-hosted, e.g. alongside the amira site)
The multi-stage Dockerfile crawls a fresh snapshot at build time
and runs the self-contained bundle (no node_modules at runtime). On a Linux
host you can equally run it under systemd behind a reverse proxy:
# /etc/systemd/system/amira-mcp.service
[Service]
ExecStart=/usr/bin/node /opt/amira-mcp/server/http.js
Environment=PORT=8787 HOST=127.0.0.1 AMIRA_LIVE_REFRESH=true
Restart=always
User=www-data
location /mcp { proxy_pass http://127.0.0.1:8787/mcp; proxy_buffering off; client_max_body_size 64k; }
proxy_buffering off keeps the Streamable-HTTP/SSE responses flowing. The refresh uses AMIRA_SITE_BASE (the public HTTPS API by default), even when
co-located with Omeka. It performs a full crawl when changes are detected;
co-location does not eliminate API load. The server caps request bodies at 64 KiB
and rate-limit state at 10,000 clients. Configure connection, request-rate and
concurrency limits at the reverse proxy for public deployments; the in-process
limiter is a courtesy control. Shutdown cancels refresh work and gives HTTP
connections a ten-second drain window.
For a build that consumes a previously reviewed data/ snapshot without calling
Omeka during the build:
docker build --build-arg SNAPSHOT_STAGE=bundled -t amira-mcp .
The default SNAPSHOT_STAGE=fetch still crawls fresh data. Both stages validate the
snapshot. The image pins its base digest; npm ci uses the committed lockfile.
Develop / rebuild
npm ci
npm run fetch-data # crawl the public Omeka API -> ./data snapshot (~1 min)
npm run typecheck # tsc --noEmit
npm run build # esbuild -> server/{index,http,fetchCli,lib}.js
npm test # unit tests: transform fixtures, folding, snapshot + store lifecycle,
# and the full tool layer against a fixture snapshot via
# InMemoryTransport (offline)
npm run test:live # integration tests against the live API (network)
npm run smoke # test/smoke/stdio.mjs: spawn the stdio server, exercise every tool offline
npm run smoke:http # test/smoke/http.mjs: spawn the HTTP server on a free port: search/fetch +
# parity, CORS preflight for both protocol revisions, the rate limiter
npm run weigh # token budget report (needs ./data — run fetch-data first)
npm run benchmark # offline latency baseline against ./data (30 warm samples per query)
Token budgets
Shortened here. Read the whole README on GitHub.
Advanced
- Delivery
- amira-mcp-server MCP server → your ahel connector (mcp.ahel.ai) → your AI.
- Item type
- mcp-server
- Key
io-github-am-digital-research-environment-amira-mcp-server- Source
- github.com/AM-Digital-Research-Environment/amira-mcp-server
- Hosted endpoint
https://data.africamultiple.uni-bayreuth.de/mcp