web2md-core

MCP serverDocs & knowledge

MCP Server for Web2MD — convert URLs to Markdown from Claude Desktop, Cursor, etc.

Unavailable. This server has no hosted endpoint yet, so ahel can't serve it.

Connect ahel once, and every AI you use reads what you have installed.

From the project's README

As published by io-oi-ai/web2md in README.md.

Turn messy HTML into clean, LLM-ready Markdown.

This is the extraction and conversion engine behind Web2MD. It is a pure library — give it an HTML string, get Markdown back. No network calls, no API key, no account. Nothing in this package talks to a server.

npm install web2md-core

Why convert at all

Feeding raw HTML to a language model wastes most of your context window on markup, navigation, and ads. Converting first cuts that down and gives the model a document it can actually follow.

The library reports both numbers so you can see the difference:

import { convertToMarkdown } from 'web2md-core'

const result = convertToMarkdown(html, { url: 'https://example.com/post' })

console.log(result.markdown)
console.log(result.metadata.originalTokenCount, '→', result.metadata.tokenCount)
// e.g. 417 → 281

convertToMarkdown returns null when it cannot find a main content block — check for that rather than assuming a result.

What it does

  • Finds the actual article. Strips navigation, sidebars, ads, cookie banners, and footers, keeping the content a reader came for.
  • Preserves structure. Headings, lists, tables, and fenced code blocks survive the round trip — that structure is what lets a model answer questions about one specific section.
  • Reports tokens. Estimated counts for both the original HTML and the cleaned Markdown, plus helpers to split or trim for a target context window.
  • Runs anywhere. Uses linkedom for parsing, so it works in Node without a browser.

API

convertToMarkdown(html, options?)

The main entry point. Note the signature takes two arguments — the URL goes inside options, not as a positional parameter:

convertToMarkdown(html, { url: 'https://example.com/post' })
OptionDefaultMeaning
urlSource URL. Used to resolve relative links and fill metadata.url.
includeLinksfalseKeep <a> as Markdown links. Off by default because link URLs are often the bulk of the tokens on navigation-heavy pages.
includeImagesfalseKeep images. Off by default for the same reason.
includeMetafalsePrepend a metadata block (title, source, timestamp).
customRuleA CustomRule for site-specific extraction.
detectCodeLanguagefalseTry to infer the language of fenced code blocks.

includeLinks and includeImages default to off. That is deliberate — the primary use case is feeding an LLM, where both are usually noise. Turn them on when you are archiving rather than summarising.

Other exports

quickConvert(html, url?)        // same result, but with links, images and
                                // metadata turned ON — the "archive it" preset
extractContent(html, url?)      // main content element, before conversion
htmlToMarkdown(html, options?)  // low-level conversion, no extraction
countTokens(text)               // token estimate
splitByTokens(md, limit)        // chunk for RAG ingestion
optimizeForContextWindow(md, model)
htmlLooksLikeLoginWall(html)    // detect login walls so you can fail loudly
MODEL_CONTEXT_LIMITS            // context sizes for common models

Markdown → sanitized HTML (via DOMPurify), for previewing output:

renderMarkdownSync(md)
renderMarkdownFull(md)          // async; includes syntax highlighting
renderMarkdownWithFormulas(md)  // KaTeX math

getPageHTML() and getSelectionHTML() read document directly and therefore only work in a browser. They throw in Node — that boundary is intentional.

Site-specific extraction

Generic extraction handles most pages. When a site needs special treatment, pass a rule:

convertToMarkdown(html, {
  url: 'https://example.com/thread',
  customRule: {
    name: 'Example forum',
    domain: 'example.com',
    contentSelector: '.thread-body',
    removeSelectors: ['.signature', '.ad-slot'],
  },
})

Scope

This package covers extraction and conversion. It does not include Web2MD's browser extension, hosted API, or account system — those stay in the product.

Contributions to extraction quality are especially welcome: if a site converts badly, an issue with the URL and what went wrong is genuinely useful.

License

MIT

Signals

Last commit
Sep 2026
Weekly downloads
695
Advanced
Delivery
web2md MCP server → your ahel gateway (mcp.ahel.ai) → every connected AI client.
Catalog kind
mcp-server
Gateway key
org-web2md-web2md
Source
github.com/io-oi-ai/web2md