openakita/skills@baidu-paddleocr-doc

SkillDocs & knowledge

Lets your agent extract text, tables, and formulas from documents and images using PaddleOCR.

Instructions available. Your AI can read the instructions. Execution depends on the setup they require.

Add ahel to your AI once: Claude, ChatGPT, Cursor, Claude Code or Codex. Then ask it to use this.

Then ask your AI: use the openakita/skills@baidu-paddleocr-doc skill

About this skill

PaddleOCR document parsing skill based on PaddleOCR-VL-1.5. Provides SOTA-level document understanding with ultra-high precision recognition and parsing. Use when user needs to parse, extract, or understand document content.

What this skill tells your AI

The instructions your AI receives, as published by openakita/openakita in skills/baidu-paddleocr-doc/SKILL.md and read by ahel’s review.

基于 SOTA 文档解析模型 PaddleOCR-VL-1.5 构建,为 Agent 加上"眼睛",对文档进行超高精度识别、解析。

配置

export BAIDU_API_KEY="your_key"

功能

  • 文档结构识别
  • 表格提取与还原
  • 公式识别
  • 图文混排解析
  • 多语言文档支持

预置脚本

scripts/baidu_ocr_doc.py

百度文档/表格 OCR 识别,需设置 BAIDU_OCR_AK 和 BAIDU_OCR_SK。

python3 scripts/baidu_ocr_doc.py doc /path/to/document.jpg
python3 scripts/baidu_ocr_doc.py table /path/to/table.png

Signals

GitHub stars
2k
Forks
280
Last commit
Sep 2026
Advanced
Item type
skill
Key
openakita-skills-baidu-paddleocr-doc
Source
github.com/openakita/openakita