PDF Extract

SkillFiles & storage

Extract text and tables from PDF files using command-line tools in the shared volume at {{SHARED_VOLUME}}.

Instructions available. Your AI can read the instructions. Execution depends on the setup they require.

Add ahel to your AI once: Claude, ChatGPT, Cursor, Claude Code or Codex. Then ask it to use this.

Then ask your AI: use the PDF Extract skill

What this skill tells your AI

The instructions your AI receives, as published by bidewio/better-openclaw in skills/pdf-extract/SKILL.md and read by ahel’s review.

Process PDF files from the shared volume at {{SHARED_VOLUME}}.

Extract Text

# Extract all text
pdftotext {{SHARED_VOLUME}}/input/document.pdf {{SHARED_VOLUME}}/output/document.txt

# Extract text from specific pages
pdftotext -f 1 -l 5 {{SHARED_VOLUME}}/input/document.pdf {{SHARED_VOLUME}}/output/pages.txt

Get PDF Info

pdfinfo {{SHARED_VOLUME}}/input/document.pdf

Tips for AI Agents

  • Use pdftotext (poppler-utils) for reliable text extraction.
  • Use tabula-java or camelot for structured table extraction.
  • OCR-scanned PDFs require tesseract for text recognition.

Signals

GitHub stars
58
Forks
6
Last commit
Aug 2026
Advanced
Item type
skill
Key
pdf-extract
Source
github.com/bidewio/better-openclaw