Semantic Scholar MCP Server

MCP serverSearch

MCP server for the Semantic Scholar API: search 200M+ papers, citations, authors, recommendations.

Unavailable. This server has no hosted endpoint yet, so ahel can't serve it.

Connect ahel once, and every AI you use reads what you have installed.

From the project's README

As published by smaniches/semantic-scholar-mcp in README.md.

A 14-tool Semantic Scholar MCP server for academic research workflows. Direct access to 200M+ papers from Semantic Scholar — paper search, citation graph traversal, author profiles, and recommendations — from any Model Context Protocol client (e.g., Claude Desktop, Claude Code, Cursor, Cline, Continue, and others).

Every release ships verifiable supply-chain provenance: Sigstore-signed SLSA build-provenance attestations on the wheel, sdist, and container image; PEP 740 attestations on the PyPI upload; and a CycloneDX SBOM — so you can prove the artifact you installed was built from this repo. See Provenance & supply chain.

Author: Santiago Maniches · ORCID 0009-0005-6480-1987 · TOPOLOGICA LLC


Quick start

uvx s2-mcp-server                                      # run instantly, no install
claude mcp add semantic-scholar -- uvx s2-mcp-server   # or register it in Claude Code

No API key is needed to start (public rate limit: 1 req/sec); set SEMANTIC_SCHOLAR_API_KEY for 10 req/sec. Claude Desktop, Docker, pip, and remote (Streamable HTTP) setups are in Installation.


Provenance & supply chain

A research tool is only as trustworthy as the chain from its source to the binary you run. Every release of this server ships cryptographically verifiable supply-chain evidence, all generated in CI from the tagged commit:

GuaranteeWhat it provesWhere it is produced
SLSA build provenance (wheel + sdist)the published distributions were built by this repo's publish.yml from the released tag, not hand-uploadedpublish.ymlactions/attest-build-provenance (build job)
SLSA build provenance (container image)the ghcr.io image digest was built by this repo's docker.ymldocker.ymlactions/attest-build-provenance, push-to-registry (lines 141–147)
PEP 740 attestationsthe PyPI upload itself carries Sigstore-backed attestations under Trusted Publishingpublish.ymlattestations: true (publish-pypi job)
CycloneDX SBOMa machine-readable bill of materials, generated in an unprivileged job from the exact wheel's statically resolved dependency metadata (wheels only, none of it executed), SHA-256-bound to that wheel, then attested against the wheel alonepublish.ymlcyclonedx-py + scripts/release_sbom.py (sbom job) + actions/attest-sbom (attest-sbom job)
SHA-pinned Actionsevery CI action is pinned to a commit SHA, so the release pipeline itself cannot silently changeall jobs in .github/workflows/ (e.g. publish.yml, docker.yml)

Verify the wheel and the container image against their attestations with the GitHub CLI:

# Wheel / sdist (download from the PyPI project or the release assets first)
gh attestation verify s2_mcp_server-*.whl --repo smaniches/semantic-scholar-mcp

# Container image
gh attestation verify oci://ghcr.io/smaniches/semantic-scholar-mcp:latest \
  --repo smaniches/semantic-scholar-mcp

The full supply-chain posture, including the known-limitations list, is in SECURITY.md. This is release-time provenance (proving how the artifact was built); the server does not currently attach a per-response receipt to individual API results.


How it compares

There is no public Semantic Scholar MCP standard, so the most useful comparison is against the obvious alternative: calling the Semantic Scholar REST API yourself from an agent. Everything in the right-hand column is plumbing this server already owns and the caller would otherwise reimplement.

This serverRaw S2 REST API from an agent
Tool surface14 typed MCP tools (search, retrieval, recommendations, status)caller composes raw HTTP requests
Citation graphboth directions (citations and references) in get_papermanual paging over two endpoints
Bulk operationspapers (≤500) and authors (≤1000) in one callcaller batches and paginates
Full-text snippet searchsnippet_search with surrounding contextseparate endpoint, caller-assembled
Paper-ID resolutionseven formats — Semantic Scholar ID, DOI, ArXiv, PubMed, Corpus ID, ACL, URL — validated pre-flight (validators.py)caller normalizes and validates IDs
Rate limitingclient-side per-tier limiter, never exceeds the interval (client.py)caller throttles by hand
Retry / backoffbounded, jittered retry on 429/502/503/timeout, honors Retry-After (client.py)caller implements retry
Errorstyped exception hierarchy, branchable by caller (errors.py)parse HTTP status strings
Outputchat-tuned Markdown or JSON per call (formatters.py)raw JSON
Supply-chain provenanceSLSA + PEP 740 + CycloneDX SBOM per release (see above)n/a
Citabilityminted Zenodo DOI, MIT licensedn/a

Installation

Option 1: One-Line Install (Recommended)

# No cloning needed — runs directly from PyPI
uvx s2-mcp-server

Option 2: Claude Code

claude mcp add semantic-scholar -- uvx s2-mcp-server

Option 3: Claude Desktop (Windows)

Add to %APPDATA%\Claude\claude_desktop_config.json:

{
  "mcpServers": {
    "semantic-scholar": {
      "command": "uvx",
      "args": ["s2-mcp-server"],
      "env": {
        "SEMANTIC_SCHOLAR_API_KEY": "your-key-here"
      }
    }
  }
}

Option 4: Claude Desktop (macOS)

Add to ~/Library/Application Support/Claude/claude_desktop_config.json:

{
  "mcpServers": {
    "semantic-scholar": {
      "command": "uvx",
      "args": ["s2-mcp-server"],
      "env": {
        "SEMANTIC_SCHOLAR_API_KEY": "your-key-here"
      }
    }
  }
}

Option 5: pip / From Source

pip install s2-mcp-server
# or
git clone https://github.com/smaniches/semantic-scholar-mcp.git
cd semantic-scholar-mcp && pip install -e .

Option 6: Docker

docker pull ghcr.io/smaniches/semantic-scholar-mcp:latest
docker run -e SEMANTIC_SCHOLAR_API_KEY=your-key ghcr.io/smaniches/semantic-scholar-mcp

Option 7: Remote server (Streamable HTTP) — requires ≥ 1.5.0

# Serve MCP over HTTP at http://127.0.0.1:8000/mcp instead of stdio
# (--from pins the floor: uvx may otherwise reuse a cached older version)
uvx --from "s2-mcp-server>=1.5.0" s2-mcp-server --transport http

See Remote access (Streamable HTTP) for client configuration, per-request API keys, and deployment guidance.

Note: Get a free API key at semanticscholar.org/product/api. Without a key, you get rate-limited public access (1 req/sec).


Architecture

flowchart LR
  Client["MCP client<br/>(Claude Desktop, Claude Code,<br/>Cursor, Cline, Continue, …)"]
  subgraph Server ["s2-mcp-server (this package)"]
    direction TB
    FastMCP["FastMCP runtime<br/>(stdio / Streamable HTTP, lifespan)"]
    Tools["14 @mcp.tool functions<br/>(server.py)"]
    Models["Pydantic input models<br/>+ field sets (models.py)"]
    Validators["Paper-ID validator<br/>(validators.py)"]
    Cache["TTL cache<br/>(cache.py)"]
    Fmt["Markdown formatters<br/>(formatters.py)"]
    HTTP["httpx client<br/>+ rate limit + retry/backoff<br/>(client.py)"]
    Errors["Typed exceptions<br/>(errors.py)"]
    Log["Structured JSON logger<br/>(logging_config.py)"]
  end
  S2Graph["Semantic Scholar<br/>Graph API"]
  S2Recs["Semantic Scholar<br/>Recommendations API"]

  Client <-- "stdio or Streamable HTTP<br/>(JSON-RPC)" --> FastMCP
  FastMCP --> Tools
  Tools --> Models
  Tools --> Validators
  Tools --> Cache
  Tools --> HTTP
  Tools --> Fmt
  HTTP --> Errors
  HTTP --> Log
  HTTP -- "GET / POST<br/>x-api-key" --> S2Graph
  HTTP -- "GET / POST<br/>x-api-key" --> S2Recs

Module responsibilities (src/semantic_scholar_mcp/):

ModuleResponsibility
server.pyFastMCP instance, 14 @mcp.tool registrations, lifespan, main() entry. Re-exports the helper surface for back-compat.
transport.pyStreamable HTTP transport: CLI/env parsing (--transport http), uvicorn wiring, and per-request API-key extraction (header / query param / Smithery config) into a request-scoped contextvar.
client.pyShared httpx.AsyncClient singleton, per-tier rate limiter (1 req/s public, 10 req/s keyed), retry loop with exponential backoff + jitter on 429/502/503/timeout, HTTP→typed-exception mapping.
models.pyPydantic input models per tool, ResponseFormat enum, the four tiered field-set constants (PAPER_SEARCH_FIELDS, …_LITE, PAPER_BULK_SEARCH_FIELDS, PAPER_DETAIL_FIELDS, AUTHOR_FIELDS).
validators.pyPre-flight paper-ID validation. Rejects NUL bytes, ?, #, path traversal; accepts the seven canonical ID formats.
cache.pyIn-memory TTL cache (5 min, 200 entries, oldest-first eviction) for paper/author lookups within a session.
formatters.pyMarkdown renderers for paper and author dicts, tuned for chat-surface readability.
errors.pySemanticScholarError hierarchy: AuthenticationError, RateLimitError, NotFoundError, ValidationError, ServerError.
logging_config.pyOne-JSON-per-line StructuredFormatter on stderr; safe to ship through any log aggregator.

Design choices worth knowing

  • Single httpx.AsyncClient per process. Created lazily, closed in the FastMCP lifespan teardown. Amortizes connection setup; respects keep-alive limits. The lifespan is reference-counted: under the Streamable HTTP transport the SDK enters it per request, so teardown only runs when the last holder exits.
  • Rate limit is enforced at the client, not the API. A semaphore + last-request timestamp ensures we never exceed the per-tier interval even when the MCP host issues tool calls in parallel.
  • Retry is bounded and jittered. Up to MAX_RETRIES = 3, base 1 s, capped at 30 s. Honors Retry-After when present.
  • Errors are typed. Status codes map onto a small exception hierarchy so callers can branch on AuthenticationError vs RateLimitError vs NotFoundError instead of parsing strings.
  • Input validation is pre-flight. Paper IDs are checked before any outbound request; bad IDs never hit the wire.
  • Version is single-source. __version__ is derived from importlib.metadata.version("s2-mcp-server"), so bumping pyproject.toml is sufficient; release-please bumps the manifest, server.json (×2 paths), CITATION.cff, and .zenodo.json in lockstep on every release.

Configuration

API Key Options

You can provide your API key in three ways:

  1. Environment Variable (recommended for persistent use):

    export SEMANTIC_SCHOLAR_API_KEY="your-api-key-here"
    
  2. Per-request HTTP header (Streamable HTTP transport only): send x-api-key: your-key with each request — see Remote access (Streamable HTTP).

  3. Per-Request Parameter (overrides env var):

    {
      "api_key": "your-api-key-here"
    }
    

    Deprecated: per-request api_key is deprecated and will be removed in v2.0.0. Tool-call arguments may be visible in MCP transcripts, client logs, and the LLM's tool-call history. Use the SEMANTIC_SCHOLAR_API_KEY environment variable instead. See SECURITY.md for details.

Get a free API key at: https://www.semanticscholar.org/product/api

Claude Desktop Setup

Add to your Claude Desktop config file:

Windows: %APPDATA%\Claude\claude_desktop_config.json macOS: ~/Library/Application Support/Claude/claude_desktop_config.json Linux: ~/.config/Claude/claude_desktop_config.json

{
  "mcpServers": {
    "semantic-scholar": {
      "command": "python",
      "args": ["-m", "semantic_scholar_mcp"],
      "env": {
        "SEMANTIC_SCHOLAR_API_KEY": "your-api-key-here"
      }
    }
  }
}

Then restart Claude Desktop.


Remote access (Streamable HTTP)

stdio remains the default transport. --transport http serves the same 14 tools over the MCP Streamable HTTP transport, which is what remote clients — claude.ai custom connectors, Smithery listings, mcp-remote bridges — connect to.

Requires s2-mcp-server ≥ 1.5.0. Earlier releases (≤ 1.4.0) do not parse CLI flags: they silently ignore --transport http and start a stdio server instead, never opening the port.

# Local HTTP endpoint at http://127.0.0.1:8000/mcp
# (--from pins the floor: uvx may otherwise reuse a cached older version)
uvx --from "s2-mcp-server>=1.5.0" s2-mcp-server --transport http

# Bind a public interface and custom port (only behind a TLS proxy — see Security)
uvx --from "s2-mcp-server>=1.5.0" s2-mcp-server --transport http --host 0.0.0.0 --port 8080

# Docker
docker run -p 8000:8000 ghcr.io/smaniches/semantic-scholar-mcp --transport http

Flags and environment variables

FlagEnv varDefaultMeaning
--transportMCP_TRANSPORTstdiostdio, http (alias: streamable-http)
--hostMCP_HOST127.0.0.1Bind address (0.0.0.0 in the Docker image)
--portMCP_PORT, then PORT8000Bind port (PORT is honored for hosting platforms)
--pathMCP_PATH/mcpURL path of the MCP endpoint
MCP_STATELESS_HTTPtrueOne independent server interaction per request (recommended)
MCP_JSON_RESPONSEtruePlain JSON responses instead of SSE streams

CLI flags beat environment variables. The server is stateless and returns JSON by default — the configuration recommended for production Streamable HTTP deployments — and no tool relies on sessions, streaming, or server-initiated messages, so there is no functional trade-off.

Per-request API keys (bring your own key)

When served over HTTP, each request may carry its own Semantic Scholar API key; concurrent users never share or observe each other's keys. Sources, in precedence order:

  1. x-api-key HTTP header (recommended)
  2. SEMANTIC_SCHOLAR_API_KEY query parameter (Smithery session config)
  3. api_key query parameter
  4. Legacy base64 ?config= parameter (older Smithery deployments)

A request without a key falls back to the server's SEMANTIC_SCHOLAR_API_KEY environment variable, or to keyless public-tier access.

Client configuration

Claude Code

claude mcp add --transport http semantic-scholar http://127.0.0.1:8000/mcp \
  --header "x-api-key: your-key-here"

JSON config (clients that accept a url)

{
  "mcpServers": {
    "semantic-scholar": {
      "type": "http",
      "url": "http://127.0.0.1:8000/mcp",
      "headers": { "x-api-key": "your-key-here" }
    }
  }
}

claude.ai custom connectors require a public HTTPS URL and accept either authless servers or OAuth — API keys in the connector URL are not supported by claude.ai. Host the server with the key supplied server-side (SEMANTIC_SCHOLAR_API_KEY env var) and register the public /mcp URL as the connector.

Smithery lists remote servers by URL (smithery mcp publish <url>); the per-request key extraction above is compatible with Smithery session config out of the box.

Security notes

  • The HTTP transport performs no authentication of inbound callers. The default bind is loopback (127.0.0.1). Expose it publicly only behind a TLS-terminating reverse proxy, and prefer the x-api-key header over query parameters (URLs end up in access logs).
  • API keys are request-scoped, and the server itself never logs them. (A key placed in a URL query parameter can still appear in access logs, as noted above — prefer the x-api-key header.)
  • See SECURITY.md for the project's broader threat model.

Supported ID Formats

The server accepts the following paper identifier formats:

FormatPatternExample
Semantic Scholar ID40-character hex649def34f8be52c8b66281af98ae884c09aef38b
DOIDOI:xxxDOI:10.1038/s41586-021-03819-2
ArXivARXIV:xxxARXIV:2106.15928 or ARXIV:2106.15928v2
PubMedPMID:xxxPMID:32908142
Corpus IDCorpusId:xxxCorpusId:215416146
ACLACL:xxxACL:P19-1285
URLURL:xxxURL:https://arxiv.org/abs/2106.15928

Tools Reference

1. semantic_scholar_search_papers

Search for academic papers with advanced filters.

Parameters:

ParameterTypeRequiredDescription
querystringYesSearch query (supports AND, OR, NOT operators and "phrase search")
yearstringNoYear filter: "2024", "2020-2024", or "2020-"
fields_of_studystring[]NoFilter by fields: ["Computer Science", "Biology"]
publication_typesstring[]NoFilter by type: ["Review", "JournalArticle"]
open_access_onlybooleanNoOnly return open access papers (default: false)
min_citation_countintegerNoMinimum citation count
limitintegerNoMax results 1-100 (default: 10)
offsetintegerNoPagination offset (default: 0)
response_formatstringNo"markdown" or "json" (default: markdown)
api_keystringNoOverride environment API key

Example:

Search for "transformer attention mechanism" papers from 2023 with at least 100 citations

JSON Example:

{
  "query": "transformer attention mechanism",
  "year": "2023",
  "min_citation_count": 100,
  "fields_of_study": ["Computer Science"],
  "limit": 20
}

2. semantic_scholar_get_paper

Get detailed information about a specific paper.

Parameters:

ParameterTypeRequiredDescription
paper_idstringYesPaper ID in any supported format
include_citationsbooleanNoInclude citing papers (default: false)
include_referencesbooleanNoInclude referenced papers (default: false)
citations_limitintegerNoMax citations to return 1-100 (default: 10)
references_limitintegerNoMax references to return 1-100 (default: 10)
response_formatstringNo"markdown" or "json" (default: markdown)
api_keystringNoOverride environment API key

Example:

Get details for DOI:10.1038/s41586-021-03819-2 including its top 20 citations

JSON Example:

{
  "paper_id": "DOI:10.1038/s41586-021-03819-2",
  "include_citations": true,
  "citations_limit": 20
}

3. semantic_scholar_search_authors

Search for academic authors by name.

Parameters:

ParameterTypeRequiredDescription
querystringYesAuthor name to search
limitintegerNoMax results 1-100 (default: 10)
offsetintegerNoPagination offset (default: 0)
response_formatstringNo"markdown" or "json" (default: markdown)
api_keystringNoOverride environment API key

Example:

Find author "Yoshua Bengio"

JSON Example:

{
  "query": "Yoshua Bengio",
  "limit": 5
}

4. semantic_scholar_get_author

Get author profile with publications.

Parameters:

ParameterTypeRequiredDescription
author_idstringYesSemantic Scholar author ID
include_papersbooleanNoInclude publications (default: true)
papers_limitintegerNoMax papers to return 1-100 (default: 20)
response_formatstringNo"markdown" or "json" (default: markdown)
api_keystringNoOverride environment API key

Example:

Get author profile for author ID 1741101 with their top 50 publications

JSON Example:

{
  "author_id": "1741101",
  "include_papers": true,
  "papers_limit": 50
}

5. semantic_scholar_recommendations

Get AI-powered paper recommendations based on a seed paper.

Parameters:

ParameterTypeRequiredDescription
paper_idstringYesSeed paper ID in any supported format
from_poolstringNoRecommendation pool: "recent" (default) or "all-cs"
limitintegerNoMax recommendations 1-100 (default: 10)
response_formatstringNo"markdown" or "json" (default: markdown)
api_keystringNoOverride environment API key

Example:

Get recommendations based on paper 649def34f8be52c8b66281af98ae884c09aef38b

JSON Example:

{
  "paper_id": "ARXIV:1706.03762",
  "limit": 15
}

6. semantic_scholar_bulk_papers

Retrieve multiple papers in a single request (max 500).

Parameters:

ParameterTypeRequiredDescription
paper_idsstring[]YesList of paper IDs (max 500)
response_formatstringNo"markdown" or "json" (default: json)
api_keystringNoOverride environment API key

Example:

Retrieve these papers: DOI:10.1038/nature12373, ARXIV:2106.15928, PMID:32908142

JSON Example:

{
  "paper_ids": [
    "DOI:10.1038/nature12373",
    "ARXIV:2106.15928",
    "PMID:32908142"
  ]
}

7. semantic_scholar_bulk_search

Shortened here. Read the whole README on GitHub.

Signals

GitHub stars
21
Forks
2
Last commit
Sep 2026
Advanced
Delivery
semantic-scholar-mcp MCP server → your ahel gateway (mcp.ahel.ai) → every connected AI client.
Catalog kind
mcp-server
Gateway key
io-github-smaniches-semantic-scholar-mcp
Source
github.com/smaniches/semantic-scholar-mcp