OCR Extraction
SkillDocs & knowledgeOCR text extraction: perform OCR on scanned documents, images, and PDFs without a text layer to extract editable text
Available today. Use it from your connected AI after setup.
No other account needed.
Connect ahel once, and every AI you use reads what you have installed.
Then ask your AI: use the OCR Extraction skill
What this skill tells your AI
The instructions your AI receives, as published by sunyifeisb-art/legalwork in skills/ocr_extraction/SKILL.md and read by ahel’s review.
何时使用
当用户上传或引用的文档符合以下任一情况时,必须执行 OCR,不能直接判定“无法提取”:
- 扫描件 PDF(页面本质是图片,没有可选中文字)
- 照片、截图、扫描图片(PNG / JPG / TIFF / BMP / WebP)
read返回乱码、空白或提示是二进制的 PDF- 用户提到“扫描件”“图片文字”“无法提取”“识别文字”等
提取流程
- 先判断文件位置:使用
ls或bash确认文件在用户工作区中的绝对路径。 - 优先调用
ocr_agent.py:如果工作区根目录存在ocr_agent.py,直接运行:python3 ocr_agent.py auto <file_path>auto模式会自动判断是否需要 OCR,并返回结构化 JSON(含text、markdown、confidence、pages_count等)。- 若需更高精度,可换用
python3 ocr_agent.py scan <file_path> --profile complex_document。
- 没有
ocr_agent.py时的兜底:使用bash内联 Python 脚本:python3 << 'PYEOF' import fitz, pytesseract, io from PIL import Image path = "<file_path>" doc = fitz.open(path) texts = [] for page in doc: pix = page.get_pixmap(dpi=300) img = Image.frombytes("RGB", [pix.width, pix.height], pix.samples) text = pytesseract.image_to_string(img, lang="chi_sim+eng") if text.strip(): texts.append(text) print("\n---PAGE_BREAK---\n".join(texts)) PYEOF - 验证结果:如果 OCR 结果为空或明显残缺,尝试:
- 提高分辨率(
dpi=300或dpi=400) - 指定中文语言包(
lang="chi_sim") - 对图片做预处理(灰度、二值化)
- 提高分辨率(
- 继续后续任务:提取到文本后,立即用于用户要求的分析、摘要、起草、审查等,不要把 OCR 结果甩给用户就结束。
输出格式
- 直接返回提取到的文本内容,并简要说明识别到的页数/置信度。
- 如果文本较长,先给出关键片段,再告诉用户完整内容已保存到哪个文件(可用
write保存到工作区)。
注意事项
- 不要在没有尝试 OCR 的情况下告诉用户“这是扫描件,无法处理”。
- 如果 OCR 依赖的 Python 包未安装,先尝试
pip install pymupdf pytesseract pillow。 - 如果系统中没有 Tesseract,使用
brew install tesseract tesseract-lang(macOS)或apt install tesseract-ocr tesseract-ocr-chi-sim(Linux)。
Signals
- GitHub stars
- 57
- Forks
- 9
- Last commit
- Sep 2026
ahel review
K1binfo
installs-packages
Automated review, not a security audit. Ruleset v1+k2.
Advanced
- Catalog kind
- skill
- Gateway key
ocr-sunyifeisb-art- Source
- github.com/sunyifeisb-art/legalwork