H3 Prompt Writing

SkillMedia

Write MiniMax H3 video generation prompts for T2VA, I2VA, FL2VA, L2VA, and Ref2VA. Use when rewriting multimodal requests into H3 prompt structures, composing integrated_multimodal_description, overall_soundscape, and non_diegetic_music, aligning keyframes, or defining reference labels for images, videos, and audio.

Available today. Use it from your connected AI after setup.

Add ahel to your AI once: Claude, ChatGPT, Cursor, Claude Code or Codex. Then ask it to use this.

Then ask your AI: use the H3 Prompt Writing skill

What this skill tells your AI

The instructions your AI receives, as published by swotatmk/infinite-creation in skills/h3-prompt-writing/SKILL.md and read by ahel’s review.

Workflow

  1. Identify the input mode: T2VA, I2VA, FL2VA, L2VA, or full-reference Ref2VA.
    • If the model supports image input, view each unique character/scene/prop reference image ONCE via view_asset and reuse it across shots — do NOT re-view the same image per shot; use view_shot_references(shot_id) only when a shot introduces an asset you have not yet seen.
  2. For base text/keyframe modes, read references/base-en.txt and follow its final prompt structure.
  3. For full-reference mode, read references/ref-en.txt and follow its six-section rewrite format.
  4. Preserve the exact field names, section order, labels, and timing notation from the selected guide.

Base Modes

  • T2VA: build the full audiovisual timeline from text.
  • I2VA: start from the first frame and develop forward from it.
  • FL2VA: describe the continuous path between the first and last frames.
  • L2VA: infer a plausible opening and converge to the supplied last frame.

Use integrated_multimodal_description, overall_soundscape, and non_diegetic_music in the order shown in references/base-en.txt.

Full-Reference Mode

Ref2VA rewrites use subject_definitions, summary, retention_analysis, detailed_description, overall_soundscape, and non_diegetic_music in that order. Reference labels stay consistent across all sections.

Read references/ref-en.txt for label rules, retention analysis, and complete examples.

Output Rules

  • Write rewrite sections in English; preserve dialogue, lyrics, and visible scene text in their original language.
  • Describe each shot by composition, subjects, environment, actions, camera, sound, and the exact point where referenced content appears.
  • Avoid plot summaries, unresolved reference labels, and timing that does not match the requested duration.

Tips for Better Results

  • Always match the total duration of the description to the requested video length (2–15 seconds), driven by content rhythm: a single short line of dialogue needs only 2–4 seconds — do not pad short shots to 10+ seconds.
  • Keep reference labels consistent (e.g. <Picture 1>, <Video 1>, <Audio 1>) across every section.
  • Prefer concrete visual and audio details over abstract words like "cinematic" or "beautiful".
  • When using keyframes (I2VA / FL2VA / L2VA), clearly state how the first and/or last frame connects to the timeline.

Signals

GitHub stars
28
Forks
3
Last commit
Sep 2026
Advanced
Item type
skill
Key
h3-prompt-writing-swotatmk
Source
github.com/swotatmk/infinite-creation