Video Gen

SkillMedia

Lets your agent generate AI video via a video Claude skill: text-to-video, image-to-video, and clip edits.

Available today. Use it from your connected AI after setup.

Connect ahel once, and every AI you use reads what you have installed.

Then ask your AI: use the Video Gen skill

About this capability

AI video generation via Seedance, Kling, Gemini Omni, MiniMax H3, and MiniMax H3 Max. Use when the user wants to generate a video clip, text-to-video, image-to-video, first/last-frame transitions, reference-guided generation, or wants to modify / edit / extend an existing generated clip.

What this skill tells your AI

The instructions your AI receives, as published by aqm857886159/nomi in tests/fixtures/skills/chatcut-video-gen/SKILL.md and read by ahel’s review.

Submits one video generation job per call and returns a jobId. Job status and later check-backs belong to track_progress; this skill does not place videos on the timeline automatically.

When to Use

Any time the user wants to generate a video clip — text-to-video, image-to-video, first-last-frame transition, reference-based generation, or generatively editing / extending an existing video (producing new generated footage based on a source clip; not timeline trimming).

Models

ModelReferenceStrengths
seedance-2-5references/seedance25.mdDefault. 4-30s, up to 50 multimodal references, audio-only reference, timestamp control, and mp4/mov output. Strongest Seedance choice for new generation, edit, and extend.
seedance2references/seedance2.mdSeedance 2.0 full model. 4-15s, rich multimodal references, and up to 1080p output.
seedance2fastreferences/seedance2.mdSeedance 2.0 Fast — same inputs and 4-15s range as seedance2, returns in minutes instead of ~ten, ~20% cheaper, capped at 720p. Use when the user is iterating and turnaround matters more than peak fidelity.
seedance2minireferences/seedance2.mdSeedance 2.0 mini — same inputs and 4-15s range as seedance2, ~50% cheaper, capped at 720p. Use for cheap drafts and high-volume batches where the full model's fidelity isn't required.
klingreferences/kling.mdCamera-control language via prompt; strong on fine emotional / performance control for character shots. 1080p via mode:"pro".
minimax-h3references/minimax-h3.mdMiniMax H3. 4-15s, 768p/2k, first/last frames, and up to 9 image + 3 video + 3 audio references (12 total). Strong multimodal reference following.
minimax-h3-maxreferences/minimax-h3-max.mdMiniMax H3 Max. 5-15s, 480p/768p, text or first/last-frame generation. No reference mode or 2K output.
omnireferences/omni.mdGemini Omni 1.1 Flash. Targeted edits/extensions via continueFrom, first/last frames, and cheap 360p drafts. 3-10s; 720p default, upscaled 1080p/4k.

IMPORTANT: Before generating, READ the chosen model's reference for capabilities, input channels, modes, prompt structure, and model-specific behavior.

Model Selection

For a new generation, seedance-2-5 is the default. Only switch when one of:

  • User explicitly named a model ("用 Kling", "use Kling", "用 Omni", "用 mini", "/kling") — switch, no need to re-ask.
  • User explicitly asks for MiniMax H3 or needs combined image/video/audio reference guidance with optional 2K output — use minimax-h3.
  • User explicitly asks for MiniMax H3 Max or wants its highest-fidelity text/frame generation without multimodal references — use minimax-h3-max.
  • Seedance clearly can't or won't do it well — when you hit a known case where Seedance struggles, propose switching to kling and confirm before submitting.
  • User explicitly asked for a fast / cheap draft they expect to revise — propose omni (360p drafts, ≤10s, supports targeted edits) and confirm before submitting.
  • User needs Seedance output at 1080p — prefer seedance-2-5; it and full seedance2 support 1080p.
  • User wants Seedance's 2.0 inputs cheaply, or many clips at once, and is fine with 720p — propose seedance2mini (4-15s, ~50% of the 2.0 full-model cost) and confirm before submitting.
  • User is iterating on 2.0-style output and cares about turnaround — propose seedance2fast (480p or 720p, 4-15s, ~80% of the 2.0 full-model cost, returns much sooner) and confirm before submitting.

Otherwise stay on seedance-2-5.

For revising a clip that already exists, the choice is made in Step 4 — not here. A revision request is not a reason to change the default for future fresh generations. Omni 1.1 replaces the previous Omni choice; keep the omni alias and do not invent a second legacy choice.

Access note: all eight models require ChatCut paid video-generation entitlement (subscription or paid credits). Never present another model as a free workaround for a subscription gate; offer free Motion Graphic animation instead when the user asks for a free path.

Tool Params

ParamValuesDefault
promptvideo description (required)
modelseedance-2-5, seedance2, seedance2fast, seedance2mini, kling, minimax-h3, minimax-h3-max, omniseedance-2-5
durationSecondsseconds5
ratiosee model docs16:9
resolutionmodel-specific; H3: 768p/2k; H3 Max: 480p/768p; Seedance: 480p/720p/1080p; Omni: 360p/720p/1080p/4k720p (H3 models: 768p)
outputFormatmp4, mov (seedance-2-5 only)mp4
taskModegenerate, edit, extend (seedance-2-5 or omni)generate
namedescriptive asset name (required)
firstFrameproject asset ref
lastFrameproject asset ref
continueFromvideo asset ref — omni only, source to edit/extend

Model-specific params (e.g., Kling mode, Seedance refImages / refVideos / refAudios) — see the model's reference.

Input Resolution

firstFrame / lastFrame / refImages / refVideos / refAudios all take a project asset reference. Prefer a full UUID or short prefix from browse_assets; asset://<id> and same-project asset URLs returned by asset tools are also accepted. Per-slot type: frame slots and refImages → image; refVideos → video; refAudios → audio.

External URLs and base64 are not accepted. If the source is a public URL, download it to a readable local path on this computer, register it with Desktop push_asset, and pass the resulting asset id.

Workflow

Four-step loop. For each new generation, restart from Step 1 if the user's intent has shifted.

Step 1 — Align scope with the user

Before writing any prompt, align on three dimensions:

  1. Duration & segments — total length, how many shots, and whether they live in one clip or several.

    If the user has already stated a direction ("做一段", "in one video", "分别生成", "split into N shots", etc.), follow it — don't second-guess.

    Otherwise, surface the two paths and let the user pick:

    • Multi-shot within one clip (see model ref) — single inference, subject / lighting / style physically consistent across sub-shots; fits a coherent narrative within the per-clip duration cap.
    • Multiple clips — each clip is independently controllable and re-rollable, but identity and style continuity have to be carried by anchors; fits durations beyond the cap or hard scene breaks.

    Offer the trade-off; do not pick for the user.

  2. Content — what each clip depicts. Summarize back what you understood, segment by segment. When content is vague (e.g. "generate a video of a girl dancing"), the user typically hasn't specified one or more of:

    • Subject: who / what is the main subject (appearance, outfit, defining features)?
    • Action: what are they doing? (For talking / emotional shots, what micro-expression?)
    • Scene: where — setting, time of day, environmental details?
    • Lighting / color mood: what atmosphere?
    • Camera: any shot-size / angle / movement preference?
    • Style: visual style or reference (cinematic / anime / documentary / ...).

    Focus on the items that matter for this specific request and can't be safely inferred — don't turn this into a blank-filling exercise. Summarize the understood parts back to the user before proceeding.

  3. Consistency anchors — only when multiple shots reuse a character, object, or scene: identify which anchor (reference image or video) to pin across shots. For sourcing rules, see §Visual consistency across shots below.

For each dimension, check the user's words:

  • Clear — proceed.
  • Ambiguous or missing — ASK the user. Do not guess, do not default to your own interpretation. A round-trip confirmation is cheaper than a wasted generation.
What NOT to do
  • Do not "tell then submit" — announcing "I'll make this as 2 clips" and immediately submitting is not alignment, it's a unilateral decision with announcement.
  • Do not default to splitting a single-video request into multiple clips. A single clip can carry multiple sub-shots (see model ref), with subject / lighting / style physically consistent across them. Surface the trade-off, then let the user choose.
  • Do not skip the ask because you think the answer is obvious.
Hard overrides (user's explicit word wins)
  • "one clip / single clip / 一条 / 一个镜头 / in 1 clip" → never split, even if the description is objectively long.
  • "N shots / N 段 / N 个镜头" → generate exactly N.
  • "use this image / 用这张图" → use as reference, don't substitute.

Step 2 — Write the prompt

See the chosen model's reference for prompt structure and param combinations (e.g., Seedance's 8-element structure and modes; Kling's prompt tips). Before submitting, check:

  • name is a descriptive asset name — descriptive enough for the user (and you in later turns) to recognize this asset in the project library. Avoid vague names like "Untitled" or "clip 1".
  • Param combination matches the user's intent — see the Modes section in the model's reference.
  • Generated video audio is not a tool parameter. Seedance and Kling are submitted with audio enabled by the backend.
  • On validation failure, read the error and fix the inputs — do not blindly retry the same invalid arguments.

Step 3 — Submit one, wait, confirm

Submit one generation job at a time. Unless the user explicitly asked for multiple clips in parallel, do not submit the next clip until the current one completes and the user has reviewed it. Parallel submission hides problems: if the first shot has drift or wrong framing, the user would rather redo it once than have several misaligned shots to discard.

  • submit_video.ratio controls the generated asset only; it does not change the project timeline canvas. If the user requested a final output aspect ratio (for example "9:16 vertical" or "16:9 landscape"), set the timeline canvas to the same ratio with manage_timelines action=update (e.g. ratio:"9:16") before placing the completed asset. If the user asked for no black bars / full-bleed, pass fit:"cover" when setting the canvas or updating/adding the visual item.
  • Do not use this skill for job management — use Desktop's reactive track_progress action=wait for generation completion, or action=status for an immediate snapshot.
  • After submitting, call track_progress with target=generation action=wait so the result returns to this conversation. End the turn with only the job ID only when the user explicitly asked to queue it without waiting.
  • When the job finishes, surface the result to the user for review before proceeding to the next shot.
  • For Omni source edits, follow the first/middle/end frame review in its reference before claiming the requested change succeeded.
  • Model-specific failure handling — see the model's reference.

Step 4 — Iterate

When the user wants a next clip, a revision, or a continuation, first decide what the existing clip is in the next call:

The existing clip is…Path
The thing being modified — a targeted change; preservation is not pixel-exactomni + continueFrom
A source to re-generate from — new motion / camera, whole-clip style shift, extend, bridgeseedance-2-5 + refVideos; set taskMode for edit/extend
Not reusable — the intent changedregenerate, restart at Step 1

A localized change defaults to omni + continueFrom. For an explicitly chosen Omni continuation use taskMode:"extend" + continueFrom; for transitions between two images use firstFrame + lastFrame. Keep Seedance 2.5 as the general edit/extend path otherwise. Do not wait for the user to say "keep everything else" — "把气球改成黄色" / "remove the text in the corner" already means it. Route to seedance-2-5 only when the change genuinely needs re-inference (motion, camera, duration, whole-clip style), not because Seedance is more familiar.

omni source edits/extensions require a project video of up to 10 seconds. For longer sources use seedance-2-5 + refVideos or ask to trim first. This integration is stateless: do not promise the upstream 40-second multi-turn workflow.

No generative edit is lossless. Omni 1.1 supports upscaled 1080p/4k, but higher resolution does not guarantee detail or unchanged pixels. Review the result before replacing a timeline item; the original asset remains available.

Then apply the standing rules:

  • If it's the next shot in a multi-shot sequence — reuse the established anchor (see §Visual consistency across shots below for principles, model ref for flag-level details).
  • If the user's feedback is ambiguous ("it doesn't feel right") — ask what specifically to change before regenerating.
  • If the same text-prompt adjustment has failed twice — stop adjusting text. Switch to reference images, or switch to an edit path (omni continueFrom for targeted changes, seedance refVideos for re-generation).
  • Each new generation restarts the loop at Step 1 — realign if scope shifted.
  • submit_video reports the omni edit round; when it warns about chain depth, relay the suggestion to the user instead of silently continuing.

Visual consistency across shots

Text alone cannot reliably maintain visual identity across shots; visual references constrain output far more precisely than words.

Anchors: the cornerstone of consistency

An anchor is a reference image or video pinned across every shot that shares the same character, object, or style. Any multi-shot sequence with recurring visual elements needs an anchor — don't try to reproduce them from text.

Sourcing an anchor

Have reference awareness. When the user's request involves a recurring character / object / scene, think about what anchor to use before writing prompts:

  • Check the project first. What has the user already provided or approved? Uploaded images, previously generated and approved shots, or earlier project assets can all serve as anchors.
  • Match the user's intent. If the user pointed to a specific asset ("use this photo", "像上一段那样"), use that. If they described a character only in words, no anchor exists yet and one must be established.
  • When in doubt, ask the user. Don't guess which asset to pin, and don't silently generate a new anchor when the user may already have one in mind.

Establishing a new anchor (with user consent)

When no existing asset fits and one must be generated, propose it to the user first — it costs credits and shapes every downstream shot. Model-specific paths — see the chosen model's ref.

Using the anchor

  • Pass the anchor in every shot that shares the character / object / style. The specific flag(s) to use depend on the model — see the model's ref.
  • Describe the anchor by appearance in the prompt, not by name: "The BLACK RACING CAR with chrome exhaust" constrains far more than "Fleetmaster". When role confusion is likely, add explicit negations: "The motorcycle does NOT transform."
  • Refer to the anchor with @Image1 / @Video1 in the prompt — not vague phrases like "the same car as before".
  • When a shot depends on a previous generation, call track_progress with target=generation action=wait, then use its terminal outputAssetId as the anchor reference. Desktop keeps the wait open reactively; do not submit dependent shots in parallel.

Multi-character projects

When a project has multiple named characters with distinct attributes (e.g. Faz with fire energy, Kev with ice energy), treat each character as a separate anchor — one reference asset per character. In every prompt:

  • Name the active character and attach their distinctive attributes ("Kev has blue ice electric energy").
  • Add explicit negations for the others to prevent attribute leakage ("NOT red fire energy, NOT Faz's look").
  • Pin the correct character's anchor (model-specific flag — see model ref). Do not reuse another character's anchor by accident.

Missing either explicit attribution or negation causes cross-character attribute mixing.

Multiple characters in the same frame. For shots where multiple characters appear together (especially facing the camera), the model is prone to face-swap or body-clipping. Add strong positional + outfit anchors to each character and prefer a fixed camera for that shot:

  • "the character on the LEFT wears a grey-blue tactical jacket, short beard, silver earring"
  • "the character on the RIGHT wears a red cape with gold trim, long braided hair"
  • "fixed camera, medium shot, both characters clearly separated"

Positional words (left / right / foreground / background) + distinctive outfit colors give the model enough signal to keep the characters apart.

Escalate when text adjustments fail

If a visual-identity issue (wrong character, drift, color mismatch) persists after two text-prompt adjustments on the same shot, stop adjusting text. Text is not a substitute for an anchor. Escalate to:

  • Adding or switching the anchor.
  • Edit mode where the model supports it (see model ref for how to invoke).

Do not submit a third text-only retry on the same consistency issue.

When to skip anchoring

Simple, one-off, or exploratory requests do not need anchors — generate directly.

Run

// Text-to-video (Seedance 2.5 default)
submit_video({
  model: "seedance-2-5",
  prompt: "A cat walks across a sunny windowsill",
  name: "Cat on windowsill",
});

// Image-to-video with Seedance 2.5 — pass the project asset id directly
submit_video({
  model: "seedance-2-5",
  prompt: "The scene comes to life, gentle breeze rustles the curtains",
  firstFrame: "abc12345",
  name: "Living room animation",
});

// Kling text-to-video — only after Model Selection check
submit_video({
  model: "kling",
  prompt: "A sports car drifts around a wet corner",
  name: "Car drift shot",
});

After submission, call track_progress with target=generation action=wait jobIds=<jobId> by default. Desktop subscribes to the live job, so completion or failure enters this same conversation and any dependent step can continue with outputAssetId. If the user explicitly asked only to queue the job, return the jobId and end the turn instead; do not busy-poll.

Config Mode

For complex multimodal jobs, build the full args object up front and pass it in a single call:

submit_video({
  model: "seedance-2-5",
  prompt: "...",
  name: "...",
  refImages: ["def67890", "ghi24680"],
  refVideos: ["abc99999"],
  refAudios: ["jkl55555"],
  durationSeconds: 8,
  ratio: "9:16",
});

Rules

  • Always provide name with a descriptive asset name.
  • Default to submit-only. End your turn after submitting unless a follow-up task is queued.
  • Do not try to manage jobs through submit_video — status and later check-backs belong to track_progress.
  • Generation costs credits. Before submitting, briefly tell the user what you're about to generate.

Signals

GitHub stars
515
Forks
114
Last commit
Sep 2026
Advanced
Catalog kind
skill
Gateway key
video-gen-aqm857886159
Source
github.com/aqm857886159/nomi