PPT Master Skill

SkillDocs & knowledge

AI-driven multi-format SVG content generation system. Converts source documents (PDF/DOCX/URL/Markdown) into high-quality SVG pages and exports to PPTX through multi-role collaboration. Use when user asks to "create PPT", "make presentation", "生成PPT", "做PPT", "制作演示文稿", or mentions "ppt-master".

Instructions available. Your AI can read the instructions. Execution depends on the setup they require.

Add ahel to your AI once: Claude, ChatGPT, Cursor, Claude Code or Codex. Then ask it to use this.

Then ask your AI: use the PPT Master Skill skill

What this skill tells your AI

The instructions your AI receives, as published by hepai-lab/drsai in examples/agent_groupchat/docmaster/skills/ppt-master/skills/ppt-master/SKILL.md and read by ahel’s review.

AI-driven multi-format SVG content generation system. Converts source documents into high-quality SVG pages through multi-role collaboration and exports to PPTX.

Core Pipeline: Source Document → Create Project → [Template] → Strategist → [Image_Generator] → Executor Live Preview → Quality Check → Post-processing → Export

[!CAUTION]

🚨 Global Execution Discipline (MANDATORY)

This workflow is a strict serial pipeline. The following rules have the highest priority — violating any one of them constitutes execution failure:

  1. SERIAL EXECUTION — Steps MUST be executed in order; the output of each step is the input for the next. Non-BLOCKING adjacent steps may proceed continuously once prerequisites are met, without waiting for the user to say "continue"
  2. BLOCKING = HARD STOP — Steps marked ⛔ BLOCKING require a full stop; the AI MUST wait for an explicit user response before proceeding and MUST NOT make any decisions on behalf of the user
  3. NO CROSS-PHASE BUNDLING — Cross-phase bundling is FORBIDDEN. (Note: the Eight Confirmations in Step 4 are ⛔ BLOCKING — the AI MUST present recommendations and wait for explicit user confirmation before proceeding. Once the user confirms, all subsequent non-BLOCKING steps — design spec output, SVG generation, speaker notes, and post-processing — may proceed automatically without further user confirmation)
  4. GATE BEFORE ENTRY — Each Step has prerequisites (🚧 GATE) listed at the top; these MUST be verified before starting that Step
  5. NO SPECULATIVE EXECUTION — "Pre-preparing" content for subsequent Steps is FORBIDDEN (e.g., writing SVG code during the Strategist phase)
  6. NO SUB-AGENT SVG GENERATION — Executor Step 6 SVG generation is context-dependent and MUST be completed by the current main agent end-to-end. Delegating page SVG generation to sub-agents is FORBIDDEN
  7. SEQUENTIAL PAGE GENERATION ONLY — In Executor Step 6, after the global design context is confirmed, SVG pages MUST be generated sequentially page by page in one continuous pass. Grouped page batches (for example, 5 pages at a time) are FORBIDDEN
  8. SPEC_LOCK RE-READ PER PAGE — Before generating each SVG page, Executor MUST read_file <project_path>/spec_lock.md. All colors / fonts / icons / images MUST come from this file — no values from memory or invented on the fly. Executor MUST also look up the current page's page_rhythm (anchor / dense / breathing), page_layouts (which template SVG to inherit, if any), and page_charts (which chart template to adapt, if any). Empty / absent entries are intentional Strategist signals — see executor-base.md §2.1. This rule exists to resist context-compression drift on long decks and to break the uniform "every page is a card grid" default
  9. SVG MUST BE HAND-WRITTEN, NOT SCRIPT-GENERATED — Every SVG page is written by the main agent directly, one page at a time (see rules 6 and 7). Writing or running a Python / Node / shell script that produces the SVG files in batch — looping over pages, templating from data, or emitting them via a generator — is FORBIDDEN, including under "save tokens", "quick draft", or "user is in a hurry" pretexts. The script-generation path was tried on a feature branch and abandoned: cross-page visual consistency depends on per-page authoring with full upstream context, which a generator script cannot reproduce

[!IMPORTANT]

🌐 Language & Communication Rule

  • Response language: match the user's input and source materials. Explicit user override (e.g., "请用英文回答") takes precedence.
  • Template format: design_spec.md MUST follow its original English template structure (section headings, field names) regardless of conversation language. Content values may be in the user's language.

[!IMPORTANT]

🔌 Compatibility With Generic Coding Skills

  • ppt-master is a repository-specific workflow, not a general application scaffold
  • Do NOT create .worktrees/, tests/, branch workflows, or generic engineering structure by default
  • On conflict with a generic coding skill, follow this skill unless the user explicitly says otherwise

Main Pipeline Scripts

ScriptPurpose
${SKILL_DIR}/scripts/source_to_md/pdf_to_md.pyPDF to Markdown
${SKILL_DIR}/scripts/source_to_md/doc_to_md.pyDocuments to Markdown — native Python for DOCX/HTML/EPUB/IPYNB, pandoc fallback for legacy formats (.doc/.odt/.rtf/.tex/.rst/.org/.typ)
${SKILL_DIR}/scripts/source_to_md/excel_to_md.pyExcel workbooks to Markdown — supports .xlsx/.xlsm; legacy .xls should be resaved as .xlsx
${SKILL_DIR}/scripts/source_to_md/ppt_to_md.pyPowerPoint to Markdown
${SKILL_DIR}/scripts/source_to_md/web_to_md.pyWeb page to Markdown (supports WeChat via curl_cffi)
${SKILL_DIR}/scripts/project_manager.pyProject init / validate / manage
${SKILL_DIR}/scripts/analyze_images.pyImage analysis
${SKILL_DIR}/scripts/latex_render.pyLaTeX formula rendering (manifest-driven PNG assets)
${SKILL_DIR}/scripts/image_gen.pyAI image generation (multi-provider)
${SKILL_DIR}/scripts/svg_quality_checker.pySVG quality check
${SKILL_DIR}/scripts/total_md_split.pySpeaker notes splitting
${SKILL_DIR}/scripts/finalize_svg.pySVG post-processing (unified entry)
${SKILL_DIR}/scripts/svg_to_pptx.pyExport to PPTX
${SKILL_DIR}/scripts/update_spec.pyPropagate a spec_lock.md color / font_family change across all generated SVGs

For complete tool documentation, see ${SKILL_DIR}/scripts/README.md.

Windows note: if a python3 ... command fails (common on python.org installs, which provide python.exe but not python3.exe), rerun the same command with python instead.

Template Index

IndexPathPurpose
Layout templates${SKILL_DIR}/templates/layouts/layouts_index.jsonQuery available page layout templates
Brand presets${SKILL_DIR}/templates/brands/brands_index.jsonQuery available brand identity presets (color / typography / logo / voice)
Visualization templates${SKILL_DIR}/templates/charts/charts_index.jsonQuery available visualization SVG templates (charts, infographics, diagrams, frameworks)
Icon library${SKILL_DIR}/templates/icons/See ${SKILL_DIR}/templates/icons/README.md; search icons on demand with ls templates/icons/<library>/ | grep <keyword>

Standalone Workflows

WorkflowPathPurpose
topic-researchworkflows/topic-research.mdPre-pipeline — gather web sources when the user supplies only a topic with no source files
template-fillworkflows/template-fill-pptx.mdGive a native PPTX template deck plus source material; select fitting pages (a page may be reused for several output slides) and fill text back without SVG conversion
create-templateworkflows/create-template.mdStandalone layout template creation workflow
create-brandworkflows/create-brand.mdStandalone brand-only template creation (identity preset; no SVG page roster)
resume-executeworkflows/resume-execute.mdPhase B entry — resume execution in a fresh chat after Phase A (Step 1–5) completed in another session (split mode)
verify-chartsworkflows/verify-charts.mdChart coordinate calibration — run after SVG generation if the deck contains data charts
customize-animationsworkflows/customize-animations.mdObject-level PPTX animation customization — run only when the user explicitly asks to tune animation order/effects/timing
live-previewworkflows/live-preview.mdBrowser-based live preview — auto-started during generation and re-enterable any time the user mentions "live preview", "preview", "看效果", or wants to click/select a slide element
visual-reviewworkflows/visual-review.mdPer-page rubric-based visual self-check — run only when the user explicitly asks for a visual re-pass on the generated SVGs (between Executor and post-processing). Opt-in only; never invoked by the main pipeline.

Workflow

Step 1: Source Content Processing

🚧 GATE: User has provided source material (PDF / DOCX / EPUB / URL / Markdown file / text description / conversation content — any form is acceptable).

No source content? When the user supplies only a topic name or requirements without any file or substantive description, run the topic-research workflow first, then return here with its products as input.

When the user provides non-Markdown content, convert immediately:

User ProvidesCommand
PDF filepython3 ${SKILL_DIR}/scripts/source_to_md/pdf_to_md.py <file>
DOCX / Word / Office documentpython3 ${SKILL_DIR}/scripts/source_to_md/doc_to_md.py <file>
XLSX / XLSM / Excel workbookpython3 ${SKILL_DIR}/scripts/source_to_md/excel_to_md.py <file>
CSV / TSVRead directly as plain-text table source
PPTX / PowerPoint deckpython3 ${SKILL_DIR}/scripts/source_to_md/ppt_to_md.py <file>
EPUB / HTML / LaTeX / RST / otherpython3 ${SKILL_DIR}/scripts/source_to_md/doc_to_md.py <file>
Web linkpython3 ${SKILL_DIR}/scripts/source_to_md/web_to_md.py <URL>
WeChat / high-security sitepython3 ${SKILL_DIR}/scripts/source_to_md/web_to_md.py <URL> (requires curl_cffi, included in requirements.txt)
MarkdownRead directly

Office vector assets (EMF/WMF) from DOCX/PPTX sources: doc_to_md.py / ppt_to_md.py extract embedded Office vector images (.emf/.wmf) alongside bitmap images. After import-sources, these land in images/ together with image_manifest.json and are first-class assets in §VIII Image Resource List.

Do NOT convert EMF/WMF to PNG. The PPT Master pipeline preserves them as external references (finalize_svg.py skips them) and svg_to_pptx.py embeds them as PPTX-native media via image/x-emf / image/x-wmf MIME — PowerPoint renders them at full vector fidelity. Converting via LibreOffice/Inkscape introduces CJK font substitution drift and rasterization loss; the original EMF/WMF is always higher fidelity than the converted PNG.

Browser-based live preview cannot render EMF (will show blank) — this is expected; the PPTX output is the source of truth.

✅ Checkpoint — Confirm source content is ready, proceed to Step 2.


Step 2: Project Initialization

🚧 GATE: Step 1 complete; source content is ready (Markdown file, user-provided text, or requirements described in conversation are all valid).

python3 ${SKILL_DIR}/scripts/project_manager.py init <project_name> --format <format>

Format options: ppt169 (default), ppt43, xhs, story, etc. For the full format list, see references/canvas-formats.md.

Import source content (choose based on the situation):

SituationAction
Has source files (PDF/MD/etc.)python3 ${SKILL_DIR}/scripts/project_manager.py import-sources <project_path> <source_files...> --move
User provided text directly in conversationNo import needed — content is already in conversation context; subsequent steps can reference it directly

⚠️ MUST use --move (not copy): all source files — Step 1's generated Markdown, original PDFs / MDs / images — go into sources/ via import-sources --move. After execution they no longer exist at the original location. Intermediate artifacts (e.g., _files/) are handled automatically.

✅ Checkpoint — Confirm project structure created successfully, sources/ contains all source files, converted materials are ready. Proceed to Step 3.


Step 3: Template Option

🚧 GATE: Step 2 complete; project directory structure is ready.

Default — free design. Proceed directly to Step 4. Do NOT query any *_index.json unless triggered. Do NOT ask the user. Do NOT proactively suggest, hint at, or fuzzy-match any template based on content, slug-like words, or vague style descriptions.

Template flow triggers ONLY on explicit directory paths supplied by the user in their initial message. The trigger rule is mechanical, not interpretive:

User input containsStep 3 action
One or more explicit template directory paths (each resolves to a directory containing design_spec.md with kind: brand / kind: layout / kind: deck in its YAML frontmatter)Read each spec's kind, dispatch per the kind matrix below, fuse if multiple
Anything else — bare template names ("用 academic_defense"), style descriptions ("麦肯锡风格"), brand mentions ("招商银行风格"), vague intent ("想用个模板"), or silenceSkip Step 3, free design

There is no slug matching, no name lookup, no fuzzy resolution. A name without a path does not trigger — the user must give a path the AI can cd into.

Style descriptions ("麦肯锡风格" / "Keynote 风" / "极简风" / etc.) never trigger Step 3. They flow into Strategist's Eight Confirmations as a style brief (color / typography / tone in confirmations e–g).

Bare names ("academic_defense", "招商银行", "anthropic") do NOT trigger Step 3 even if a matching directory exists in the library. The user must give a path. AI must not "helpfully" resolve a name to a path.

"What templates exist?" is out-of-band Q&A — answer by listing entries from brands_index.json / layouts_index.json / decks_index.json together with their paths. Listing alone does not advance the pipeline; the user must send a path back to trigger Step 3.

To create a new layout or deck, read workflows/create-template.md. To create a new brand, read workflows/create-brand.md.

Three template kinds

The architecture has three independent reference bundles. Full schema in docs/zh/templates-architecture.md. Summary:

KindPhysical dirContainsFrontmatter
brandtemplates/brands/<id>/identity-only segment: color / typography / logo / voice / icon stylekind: brand
layouttemplates/layouts/<id>/structure-only segment: canvas / page structure / page types / SVG rosterkind: layout
decktemplates/decks/<id>/full replica: identity + structure + middle (template overview) segmentskind: deck

Segment ownership (governs fusion override priority):

SegmentSectionsOwner kind on fusion
IdentityColor Scheme / Typography / Logo / Voice & Tone / Icon Stylebrand
StructureCanvas / Page Structure / Page Types / SVG Rosterlayout
MiddleTemplate Overview (use cases / design intent)deck (no other kind writes this)
Single-path dispatch
User path's kindStep 3 action
kind: brandCopy design_spec.md + logo files + asset subdirs (images/ / illustrations/ / icons/) into <project>/templates/. Strategist locks identity segment as truth; structure stays free.
kind: layoutCopy design_spec.md + SVG roster + asset files into <project>/templates/. Strategist locks structure; identity decided in Eight Confirmations e–g.
kind: deckCopy everything (design_spec.md + SVGs + logos + assets) into <project>/templates/. Strategist locks all segments; Eight Confirmations narrows to deck-content fields (audience / page count / outline / tone tweaks).
TEMPLATE_DIR=<user-supplied path>
cp -r ${TEMPLATE_DIR}/* <project_path>/templates/

The single-line copy suffices for all three kinds — the spec's kind field tells Strategist how to read it; downstream code doesn't distinguish.

Multi-path fusion

When the user gives two or more paths of different kinds, Step 3 fuses them into a single <project>/templates/design_spec.md. Default granularity is segment-level integer replacement — entire identity / structure / middle segments are taken from the highest-priority source for that segment, no implicit field-level mixing.

Override priority by segment:

CombinationIdentity fromStructure fromMiddle from
brand onlybrand(free design)(none)
layout only(free design)layout(none)
deck onlydeckdeckdeck
brand + layoutbrandlayout(none)
brand + deckbrand (overrides deck)deckdeck
layout + deckdecklayout (overrides deck)deck
brand + layout + deckbrandlayoutdeck

Field-level micro-adjustment (e.g. "use anthropic brand but primary changed to #FF0000") is not part of Step 3 fusion — it flows into Strategist Eight Confirmations e–g as a normal user request.

Same-kind multiple paths — conflict resolution

When the user gives two paths of the same kind (e.g. brands/anthropic + brands/google), Step 3 surfaces a conflict prompt before fusing — like resolving a git merge conflict:

AI: 你给了两个 brand,检测到段级冲突:
    - Color Scheme(Anthropic 橙红 vs Google 多色)
    - Typography(Styrene/AnthropicSans vs GoogleSans/Roboto)
    - Logo(Anthropic 标 vs Google 标)
    - Voice & Tone(restrained vs friendly)
    - Icon Style(stroke vs filled)

    要 (a) 全部按 Anthropic / (b) 全部按 Google / (c) 逐段挑?

Rules:

  • Default: no implicit ordering — every cross-source segment difference is reported as a conflict
  • Only when the user picks (c) does AI walk through each segment one by one
  • Field-level conflicts are out of scope — segment-level only
  • Three or more same-kind paths are not supported — ask the user to converge to at most two
Fused spec provenance

When fusion happens (any multi-path case), the resulting <project>/templates/design_spec.md carries a provenance block immediately under its H1:

> **Fused from:**
> - deck: `templates/decks/招商银行/` (base)
> - brand: `templates/brands/anthropic/` (identity override)
> - layout: `templates/layouts/academic_defense/` (structure override)
> - conflicts resolved: Color Scheme from anthropic(user picked a)

Single-path Step 3 does not add provenance (the source is self-evident from the copied files).

✅ Checkpoint — Default path proceeds to Step 4 without user interaction. If the user supplied one or more explicit template paths, those have been dispatched (or fused) into <project_path>/templates/ before advancing.


Step 4: Strategist Phase (MANDATORY — cannot be skipped)

🚧 GATE: Step 3 complete; default free-design path taken, or (if triggered) template files copied into the project.

First, read the role definition:

Read references/strategist.md

⚠️ Mandatory gate: before writing design_spec.md, Strategist MUST read_file templates/design_spec_reference.md and follow its full I–XI section structure. See strategist.md Section 1.

Eight Confirmations (full template: templates/design_spec_reference.md):

⛔ BLOCKING: present the Eight Confirmations as a single bundled recommendation set and wait for explicit user confirmation or modification before outputting Design Specification & Content Outline. This is the single core confirmation point — once confirmed, all subsequent steps proceed automatically.

  1. Canvas format
  2. Page count range
  3. Target audience
  4. Style objective
  5. Color scheme
  6. Icon usage approach
  7. Typography plan, including formula rendering policy
  8. Image usage approach

Mandatory — split-mode note (not a ninth confirmation): after listing the eight confirmation details, you MUST append exactly one short line (rendered in the user's language, prefixed with 💡) about generation mode. Pick the variant by qualitative read of Phase A signals — recommended page count, source-material bulk, whether topic-research ran with substantial web-fetch accumulation:

Signal readLine content
Heavy (long page count / bulky sources / heavy web-fetch accumulation)State estimated page count and large source size; recommend switching to split mode after Step 5 — stop this chat, open a fresh window and input 继续生成 projects/<project_name> to enter Phase B (SVG generation + export); no response or "continue" = default continuous mode.
Normal (default)State scale is moderate, default continuous mode generates in one go; if mid-way window switch is desired, input 继续生成 projects/<project_name> after Step 5 to switch to split mode.

This line is required output every run — the user must always see the mode choice exists. Whether to act on it is the user's call.

Formula rendering policy lives inside item 7 (Typography plan):

PolicyBehavior
mixed (default)Strategist renders complex formula-worthy expressions as PNG assets; simple inline expressions remain editable text / Unicode
render-allStrategist renders every formula-worthy expression as PNG assets
text-onlyNo formula rendering; formulas remain editable text / Unicode

After the Eight Confirmations are approved and before outputting design_spec.md / spec_lock.md, if the confirmed formula policy is mixed or render-all and the content contains formula-worthy expressions, Strategist MUST:

  1. Identify explicit LaTeX and any source expressions that should be faithfully structured as formulas.
  2. Write <project_path>/images/formula_manifest.json with only the formulas selected for rendering.
  3. Run:
    python3 ${SKILL_DIR}/scripts/latex_render.py <project_path>
    
  4. Include the rendered formula PNGs as Acquire Via: formula, Status: Rendered, Type: Latex Formula rows in design_spec.md §VIII Image Resource List; also list them in spec_lock.md images with | no-crop.

The formula renderer uses a provider fallback chain by default: codecogs,quicklatex,mathpad,wikimedia. The first three are color-aware; Wikimedia is an availability fallback. Formula PNGs are transparent by default: manifest background is the temporary render matte and transparency-removal reference, not a retained final background unless transparent: false is set for that item. Do not scan spec_lock.md for $...$ or $$...$$. Dollar-delimited math in source material is only a signal for Strategist; the renderer consumes the explicit manifest.

If the user provided images or formula PNGs were rendered, run analysis before outputting the design spec:

python3 ${SKILL_DIR}/scripts/analyze_images.py <project_path>/images

⚠️ Image handling: NEVER directly read / open / view image files (.jpg, .png, etc.). All image info comes from analyze_images.py output or the Design Spec's Image Resource List.

Output:

  • <project_path>/design_spec.md — human-readable design narrative
  • <project_path>/spec_lock.md — machine-readable execution contract (skeleton: templates/spec_lock_reference.md); Executor re-reads before every page

✅ Checkpoint — Phase deliverables complete, auto-proceed to next step:

## ✅ Strategist Phase Complete
- [x] Eight Confirmations completed (user confirmed)
- [x] Split-mode note appended below the eight items (heavy or normal variant)
- [x] Design Specification & Content Outline generated
- [x] Execution lock (spec_lock.md) generated
- [ ] **Next**: Auto-proceed to [Image_Generator / Executor] phase

Step 5: Image Acquisition Phase (Conditional)

🚧 GATE: Step 4 complete; Design Specification & Content Outline generated and user confirmed. Any formula rows already have Acquire Via: formula and Status: Rendered.

Trigger: At least one row in the resource list has Acquire Via: ai and/or Acquire Via: web. If every row is user, formula, or placeholder, skip to Step 6.

Always load the common framework:

Read references/image-base.md

Shortened here. Read the whole file on GitHub.

Signals

GitHub stars
24
Forks
5
Last commit
Sep 2026

ahel review

  • K6low
    bundled executables the agent is told to run
  • K1binfo
    installs-packages (in scripts/analyze_images.py)
  • K1binfo
    installs-packages (in scripts/gemini_watermark_remover.py)
  • K1binfo
    installs-packages (in scripts/image_backends/backend_gemini.py)
  • K1binfo
    installs-packages (in scripts/image_backends/backend_openai.py)
  • K1binfo
    installs-packages (in scripts/image_gen.py)
  • K1binfo
    installs-packages (in scripts/latex_render.py)
  • K1binfo
    installs-packages (in scripts/notes_to_audio.py)

Automated review, not a security audit. Ruleset v1+k2.

Advanced
Item type
skill
Key
ppt-master
Source
github.com/hepai-lab/drsai