Document processing
SkillDocs & knowledgeUse when the deliverable is a document's bytes or its literal content — text/tables out of PDFs, AcroForm fill and flatten, page merge/split, PDF/DOCX from templates, OCR of image-only scans. NOT schema-typed fields pulled from text (that is structured-extraction), NOT signature routing (e-signature) or spreadsheet cells/formulas (spreadsheet-ops).
Available today. Use it from your connected AI after setup.
No other account needed.
Connect ahel once, and every AI you use reads what you have installed.
Then ask your AI: use the Document processing skill
What this skill tells your AI
The instructions your AI receives, as published by ericrisco/rsc-harness in skills/document-processing/SKILL.md and read by ahel’s review.
File in, content out — or data in, file out. You open a byte stream (PDF, DOCX, scan) and either pull the content out, or you build a new document from a template and a data dict. That is the whole job: the deliverable is bytes of a document or the literal content of one.
The boundary test, apply it first:
- Deliverable is raw text / Markdown / table cells / a generated file → you are in the right place.
- Deliverable is a typed object matching a schema (
{parties: [...], total: 1234.50}) → that isstructured-extraction. This skill stops at "clean Markdown out of the file"; the schema-constrained extraction runs on that Markdown.
Everything else routes too: signing with an audit trail → e-signature, spreadsheet grids/formulas/XLSX-as-data → spreadsheet-ops, indexing for cross-document Q&A → rag (this skill produces the text rag ingests, it does not index it), downloading the files off a site → data-scraper.
Step 0 — does the PDF have a text layer?
The most expensive mistake in this skill is OCR'ing a PDF that already has a text layer. A digital PDF (exported from Word, a browser, a report tool) carries selectable text — extracting it is free, instant, and lossless. OCR is slow, costs money or GPU, and introduces errors. Never OCR a PDF you can extract.
Check before you pick an engine:
import pdfplumber
with pdfplumber.open("doc.pdf") as pdf:
txt = pdf.pages[0].extract_text() or ""
if len(txt.strip()) > 20:
print("text layer present -> extract directly (pdfplumber / pypdf)")
else:
print("image-only or empty -> this is an OCR job")
If extract_text() returns empty (or near-empty) across the first few pages, it is a scan or image-only PDF and you go to the OCR branch. Symptom from the user's side: "the text copies out as garbage / random symbols" usually means a broken/embedded font, not a missing text layer — try pypdf extraction too before assuming OCR.
Engine selection
| Goal | Use | Why |
|---|---|---|
| Extract text + tables with layout | pdfplumber | Layout-aware; extract_tables() returns rows/cols as Python lists → pandas/CSV. |
| Raw text, merge, split, rotate, page ops | pypdf (6.12.2) | Pure-Python, no C deps, runs in Lambda/containers; the maintained core — import pypdf, never the dead PyPDF2, which was merged back into it. |
| Fill an interactive PDF form | pypdf | update_page_form_field_values writes AcroForm fields; can flatten. |
| Generate a Word/DOCX from a template | docxtpl (0.20.x) | A real .docx becomes a Jinja2 template; author in Word, tag, render. |
| Generate a PDF from scratch | ReportLab | Canvas / Platypus flowables for laid-out PDFs. |
| OCR a scan, local / no API budget | Docling (or Marker) | Layout + reading order + table structure, fully local, wraps Tesseract/RapidOCR. |
| OCR messy scans / handwriting / hard tables, API ok | Mistral OCR | mistral-ocr-2512 (OCR 3), ~$2 / 1,000 pages, tuned for forms + handwriting. |
| Fastest extract / easiest page→PNG raster | PyMuPDF ⚠️ AGPL | Fast, but AGPL: shipping it imposes an open-source obligation or needs a paid license. Flag this before recommending. |
Extraction recipes
Text + tables with pdfplumber, straight to CSV:
import csv
import pdfplumber
rows = []
with pdfplumber.open("invoice.pdf") as pdf:
for page in pdf.pages:
for table in page.extract_tables():
rows.extend(table)
with open("out.csv", "w", newline="") as f:
csv.writer(f).writerows(rows)
Raw text, merge, split, rotate with pypdf:
from pypdf import PdfReader, PdfWriter
# raw text
text = "\n".join(p.extract_text() or "" for p in PdfReader("doc.pdf").pages)
# merge two files
w = PdfWriter()
for src in ("a.pdf", "b.pdf"):
w.append(src)
with open("merged.pdf", "wb") as f:
w.write(f)
# split first 3 pages + rotate one
w2 = PdfWriter()
reader = PdfReader("doc.pdf")
for page in reader.pages[:3]:
w2.add_page(page)
w2.pages[0].rotate(90)
with open("first3.pdf", "wb") as f:
w2.write(f)
Form filling (AcroForm)
Dump the field names first — guessing them is the #1 reason a fill silently does nothing:
from pypdf import PdfReader
fields = PdfReader("form.pdf").get_fields() or {}
for name, f in fields.items():
print(name, "->", f.get("/FT")) # /Tx text, /Btn checkbox/radio, /Ch choice
Then write the values. Set auto_regenerate=False and bake with flatten=True if it must not be editable:
from pypdf import PdfReader, PdfWriter
reader = PdfReader("form.pdf")
writer = PdfWriter()
writer.append(reader)
for page in writer.pages:
writer.update_page_form_field_values(
page,
{"applicant_name": "Eric Risco", "agree": "/Yes"}, # checkbox = its on-state
auto_regenerate=False, # else a spurious "save changes?" prompt fires on open
)
# flatten=True bakes the values and drops the editable widgets
with open("filled.pdf", "wb") as f:
writer.write(f)
auto_regenerate defaults to True for legacy reasons, and you almost never want it. Checkbox/radio values are the field's /V on-state (often /Yes), not True — read the field to find it.
Generation
DOCX from a Word template you authored and tagged with Jinja2 ({{ client }}, {% tr for row in items %} on a table row, InlineImage for pictures):
from docxtpl import DocxTemplate
doc = DocxTemplate("contract_template.docx")
doc.render({
"client": "Acme SL",
"date": "2026-06-02",
"items": [{"desc": "Audit", "amount": "1.200,00 €"}],
})
doc.save("contract_2026-06-02.docx")
PDF from scratch with ReportLab Platypus:
from reportlab.lib.pagesizes import A4
from reportlab.platypus import SimpleDocTemplate, Paragraph, Spacer
from reportlab.lib.styles import getSampleStyleSheet
styles = getSampleStyleSheet()
doc = SimpleDocTemplate("report.pdf", pagesize=A4)
doc.build([
Paragraph("Quarterly Report", styles["Title"]),
Spacer(1, 12),
Paragraph("Generated automatically from the data dict.", styles["BodyText"]),
])
OCR
Branch on cost and privacy. Local, no API budget, or data must not leave the machine → Docling/Marker. Messy scans, handwriting, brutal tables, and an API budget is fine → Mistral OCR.
Local with Docling (wraps Tesseract / RapidOCR, exports Markdown preserving tables):
from docling.document_converter import DocumentConverter
result = DocumentConverter().convert("scan.pdf")
markdown = result.document.export_to_markdown()
open("scan.md", "w").write(markdown)
Hosted with Mistral OCR (~$2 / 1,000 pages, 50% off via Batch API; outputs interleaved text+images as Markdown):
from mistralai import Mistral
client = Mistral(api_key=os.environ["MISTRAL_API_KEY"])
resp = client.ocr.process(
model="mistral-ocr-2512",
document={"type": "document_url", "document_url": signed_url},
)
markdown = "\n\n".join(p.markdown for p in resp.pages)
Never trust OCR output blind. OCR confuses 0/O, 1/l/I, and drops or shifts decimal points — a 1.234,50 can come back as 1234,50 or 1,234.50. Always spot-check totals, dates, and ID numbers against the rendered page before you hand the text downstream. For clean scans with no budget, plain pytesseract is the zero-cost baseline, but it is weak on layout/tables versus the pipelines above.
Scale
Batch jobs: parallelize per-file, cap concurrency on the hosted API (rate limits + cost), and use Mistral's Batch API for the 50% discount on large runs. Engine install matrix, exact version pins, the full licensing table, the Docling-vs-Marker-vs-Mistral feature/cost comparison, and troubleshooting (encrypted PDFs, mangled AcroForm field names, multi-column reading order, CJK/handwriting) live in references/engines.md — read it before a non-trivial install.
Anti-patterns
| Anti-pattern | Why it is wrong | Do instead |
|---|---|---|
| Pipe every PDF straight to OCR | OCR'ing a digital PDF is slow, costs money, and adds errors to text you could extract losslessly | Step 0: check the text layer first; OCR only image-only PDFs |
import PyPDF2 | Unmaintained; merged into pypdf years ago — a stale-code smell | from pypdf import PdfReader, PdfWriter |
| Recommend PyMuPDF without a word about its license | PyMuPDF is AGPL; shipping it silently creates an open-source obligation | Flag AGPL; prefer pdfplumber/pypdf, or get a commercial license knowingly |
Leave auto_regenerate=True on a form fill | Marks the AcroForm dirty → a spurious "save changes?" prompt for every user | Pass auto_regenerate=False |
| Trust OCR'd totals/numbers as-is | 0/O, 1/l, shifted decimals silently corrupt amounts | Spot-check totals/dates/IDs against the page image |
| Hand-roll a regex to pull typed fields from the Markdown | Brittle, re-implements a sibling, breaks on layout drift | Output clean Markdown, hand it to structured-extraction |
| Use Mistral OCR when the user said "no cloud / local only" | Sends documents off-machine, violating the privacy constraint | Use Docling/Marker + Tesseract/RapidOCR locally |
Signals
- GitHub stars
- 82
- Forks
- 3
- Last commit
- Sep 2026
Advanced
- Catalog kind
- skill
- Gateway key
document-processing- Source
- github.com/ericrisco/rsc-harness