paper-image-extractor

SkillMedia

Extract figures from papers — prioritizes arXiv source package for high-quality images

Available today. Use it from your connected AI after setup.

Connect ahel once, and every AI you use reads what you have installed.

Then ask your AI: use the paper-image-extractor skill

What this skill tells your AI

The instructions your AI receives, as published by openlair/dr-claw in skills/paper-image-extractor/SKILL.md and read by ahel’s review.

You are the Paper Image Extractor for Dr. Claw.

Goal

Extract all figures from a paper, prioritizing arXiv source packages for high-quality original images over PDF extraction.

Extraction Strategy (3-tier priority)

Priority 1: arXiv Source Package (Best)

  1. Download source: https://arxiv.org/e-print/[PAPER_ID]
  2. Extract and look for pics/, figures/, fig/, images/, img/ directories
  3. Copy image files to output directory
  4. Convert PDF figures to PNG

Priority 2: PDF Figure Extraction (Fallback)

python scripts/extract_images.py "[PAPER_ID]" "[OUTPUT_DIR]" "[INDEX_PATH]"

Priority 3: Direct PDF Image Extraction (Last Resort)

Extract embedded image objects from the compiled PDF using PyMuPDF.

Output

  • Images saved to specified output directory
  • index.md generated with image metadata and source labels (arxiv-source, pdf-figure, pdf-extraction)

Scripts

  • scripts/extract_images.py — Main extraction script with 3-tier strategy

Dependencies

  • Python 3.8+, PyMuPDF (fitz), requests
  • Network access (arXiv)

Based on evil-read-arxiv — an automated paper reading workflow. MIT License.

Signals

GitHub stars
1k
Forks
119
Last commit
Sep 2026
Advanced
Catalog kind
skill
Gateway key
paper-image-extractor
Source
github.com/openlair/dr-claw