UI Motion

SkillWeb & browsing

Lets your agent add preset animation styles like hover, entrance, and toggle effects to UI components.

Available today. Use it from your connected AI after setup.

Add ahel to your AI once: Claude, ChatGPT, Cursor, Claude Code or Codex. Then ask it to use this.

Then ask your AI: use the UI Motion skill

About this skill

Creates UI motion effects, app interface animations, SaaS hero animations, or product demos based on real brand visuals, app screenshots, web pages, data dashboards, or explicit brand descriptions, along with accurate copy and interface interaction goals, covering brand analysis, style selection, t

What this skill tells your AI

The instructions your AI receives, as published by qxryz/workflowgenerator in skills/library/ui-motion/SKILL.md and read by ahel’s review.

This skill produces UI motion graphics that look like the brand, not like a stock template. The architecture has three load-bearing ideas:

  1. Brand is the source of truth. Colors, typography, photography mood, end-card mark all come from the user's brand assets. The 8 style files contribute motion language only.
  2. Two duration modes. A video up to and including 15s is one generation: the user's image is an all-purpose reference, not a first frame, and no extra BGM is generated by default. A video longer than 15s is a numbered continuation chain: segment 1 uses the all-purpose reference, while every later segment uses the actual tail frame produced by the immediately preceding segment as its first frame.
  3. Continuations are stateful, not repeated openings. A later segment is valid only when its sequence and continuation_of point to the preceding generated clip and its input file is that clip's extracted tail frame. Calling the opening generation procedure again never creates a continuation segment.

Inputs to collect

Ask the user in this order. Items 1–3 are required; 4–7 unlock a better output.

  1. Product / concept — what the video is for.
  2. Brand visual — at least one of:
    • Brand images (1–3 files): logo, product screenshot, marketing banner, brand guideline
    • Text description (2–5 sentences): the user's own words about the brand
    • Both is best. If the user has neither, ask once more, then default to a generic palette and call it out in the chat reply.
  3. Brand text — product name, tagline, store badges. English by default; ask if the user wants something else.
  4. "What good looks like" reference (optional) — a reference video the user admires. The original ui动画.zip upload is a perfect example. This informs motion language (pacing, easing, music mood) without touching brand identity.
  5. Aspect ratio — 9:16 (Reels/Shorts/小红书/抖音) or 16:9 (B站/YouTube). Default 9:16.
  6. Length — 6s, 10s, 15s (default), or >15s (uses numbered continuation segments). The model natively supports 6s / 10s / 15s in one segment.
  7. Music mood — only needed for >15s or when the user explicitly requests extra music. Use cinematic, lofi, electronic, ambient, corporate, or freeform. Default: derived from the brand profile.

If the user provides everything except the brand visual, do not proceed with a stock template. Push back once: "I need at least one brand image or a 2-sentence brand description to make this look like you, not like a stock reel."

Procedure

1. Lock the brief

Write a 1-paragraph brief to <workspace>/ui-motion/brief.md capturing: product/concept, the brand visual sources (file paths or quoted text), brand text, optional reference video, aspect ratio, length.

2. Extract the brand profile (REQUIRED)

This is the heart of the skill. The brand profile is a structured JSON capturing the brand's visual DNA in model-actionable form. See references/brand-analysis.md for the full procedure.

Steps:

  1. For each brand image, read it (the model can see images). Extract:
    • Background base (light / dark / gradient / photographic)
    • Dominant colors (3–5 with descriptive names + hex)
    • Typography feel (geometric / humanist / serif, weight, tracking)
    • Photography mood (lifestyle, abstract, technical, warm, clinical)
    • Visual density (spacious / balanced / dense)
    • Motion hints from composition
  2. For any text description, parse mood words into the same fields.
  3. Resolve conflicts (images win over text; logos win over marketing banners for color; ask the user once if unresolvable).
  4. 将结果写入 <workspace>/ui-motion/brand_profile.json,字段参照 brand-analysis.md。

Do not skip this step even if the user only provided one image. One image is enough.

3. Lock inspiration (optional)

If the user provided a "what good looks like" reference, analyze it for pacing, easing, transition vocabulary, music mood. Write to <workspace>/ui-motion/inspiration_profile.json. The 8 style files are pre-built inspiration profiles — if the user's reference matches one, adopt that style's profile directly (see style-matching.md).

If the user did not provide a reference, derive the inspiration profile from the brand profile alone: warm/wellness brands → slow + spring; energetic brands → fast + snap; corporate/SaaS → medium + smooth-ease.

4. Choose a style anchor (optional)

The 8 style files are reference points for motion language. They contribute: motion rhythm, easing, transition vocabulary, music mood template, geometry vocabulary. They do NOT contribute colors, typography, or photography mood — those come from the brand profile.

Procedure:

  1. If the user named a style, anchor to it.
  2. Otherwise, score each of the 8 against the brand profile. See style-matching.md.
  3. If a strong match (>70%): suggest, ask once, anchor.
  4. If no match or all weak: do not anchor to any style. Derive a custom style from the brand + inspiration profiles (see custom-style.md).
  5. Record in <workspace>/ui-motion/style_anchor.json with anchor: "A" | "B" | … | "H" | "custom". If custom, include a custom_spec block.

5. Build the storyboard

The storyboard for a single-segment video has 1 optional reference_prompt + 1 complete motion_prompt. The reference prompt selects or creates an all-purpose visual reference; it never defines a first frame. Before writing the motion field, read references/motion-prompt-writing.md.

Write the motion_prompt as a timed director's treatment, not a 2–4 sentence summary. For a 15s take, use 5–7 time-coded beats and never fewer than 1000 Chinese characters or 600 English words of non-repetitive, executable direction. Each beat must specify: visible trigger → primary UI response → secondary spatial response → camera response → exact transition bridge → readable settle. Use a coherent mix of push-in/push-through, pull-back, orbit, tilt, focus pull, path-following, parallax, and foreground occlusion; choose only the moves that serve the concept.

Define one recurring motion cause (cursor, route, card edge, ring, line, waveform, or brand glyph) and let it physically connect the whole take. In a 15s prompt, choreograph at least three explicit continuous transitions using shared-element expansion, shape morphing, path continuation, natural occlusion, depth pass-through, fold/unfold, orientation change, or material propagation. “Smooth transition”, “the page changes”, and lists of camera keywords are incomplete unless the prompt names the carrier, direction, preserved object, and landing state.

The short “Single motion prompt template” inside a style file is only a verb/motif seed. Never copy it as the final motion_prompt; expand it through references/motion-prompt-writing.md using the user's product, brand profile, interaction logic, and spatial hierarchy.

超过 15 秒时,storyboard 包含至少 2 个编号片段。第 1 段使用全能参考,第 N+1 段只能使用第 N 段已核验的尾帧,每段保留自己的 motion prompt。字段说明见 references/storyboard-schema.md。

The reference prompt and motion prompt must be derived from the brand profile first, with the style anchor's motion language applied second. Specifically:

  • Colors in the prompts come from brand_profile.colors (descriptive names, not hex).
  • Typography keywords come from brand_profile.typography.
  • Motion verbs come from the style anchor's motion signature (or the custom spec).
  • Camera and transition choices come from references/motion-prompt-writing.md and must remain causally linked to the UI interaction.
  • The end card uses the real brand mark from the user's logo, in the brand's font and color.

将结果写入 <workspace>/ui-motion/storyboard.json,遵循当前故事板结构。

6. Prepare the reference inputs

For a video up to and including 15s, keep the user's uploaded image as an all-purpose reference. Pass it as a reference image; it is not renamed or treated as first_frame.

For a video longer than 15s, prepare segment 1 with the same all-purpose reference. Do not manufacture a separate opening still unless the user explicitly asks for one. After segment N finishes, extract and probe its real tail frame; only that file may become segment N+1's opening-frame reference. Record the source and dimensions in the storyboard before the next generation call.

7. Generate the video

单段视频使用全能参考路径,并把用户图片作为参考素材附加,不能当作首帧。请求时长、画幅、分辨率和完整 motion_prompt 一起保留在故事板元数据中;执行时不得摘要、缩短或合并时间码节拍。

连续片段先用全能参考生成第 1 段,再严格按顺序用第 N 段已核验的尾帧生成第 N+1 段。每段保留自己的 sequence、承接说明、Prompt、输入文件、输出片段和提取的尾帧;后续段不得重复开场流程。

8. Generate the music (only for >15s)

不超过 15 秒的视频默认不增加音乐,保留片段原生声音。

超过 15 秒的视频必须等全部编号片段成功且总时长已知后,再根据品牌画像和全片能量时间线制作一条器乐 BGM,不写歌词也不强行指定名义时长,并记录实际时长为 audio/bg.mp3。

9. Assemble the result

For a single segment (≤15s), deliver the video directly after verifying its duration and actual dimensions. Do not add a generated BGM track.

超过 15 秒时先按 sequence 顺序完整拼接片段,再混入唯一 BGM,并保持请求的总时长;不得按先结束的音视频流提前截断。若本地无法完成合成,交付原始片段和 brand_profile.json,明确说明还需要最终合成步骤。

10. Deliver and verify brand consistency

Send final.mp4 to the user with <deliver-assets>. Run through references/qa-checklist.md with brand consistency as the top check:

  • Are the rendered colors from the brand profile (not the style anchor's defaults)?
  • Is the real brand mark in the end card, or did the model fall back to a generic glyph?
  • Any stray colors that don't appear in the brand palette?

Mention any brand-consistency flag in the chat reply. Don't ship a video that doesn't look like the brand.

The chat reply follows this structure:

  1. Creative interpretation (2–3 sentences). What you extracted from the brand and how it shaped the design. Quote one or two specific brand signals you leaned on.
  2. What the viewer sees (1 sentence per beat, 5 total).
  3. Brand-consistency flags + next steps (1–2 sentences).

Output contract

  • One file: <workspace>/ui-motion/final.mp4, H.264, 9:16 or 16:9, with the requested duration. Audio is unchanged for ≤15s; >15s includes the single post-generated BGM mix unless the user explicitly requests another audio policy.
  • Process artifacts: brief.md, brand_profile.json, inspiration_profile.json (if used), style_anchor.json, storyboard.json, intermediate frames/ and clips/; audio/bg.mp3 exists only for >15s outputs.
  • Chat reply contains the final asset + the 3-paragraph summary.

Failure handling

  • No brand visual at all (no image, no text) → ask once. If the user still can't provide, push back: "I can produce a generic Style X video, but it won't look like your brand. Confirm you want generic, or send me a logo / screenshot / 2-sentence description."
  • read of a brand image fails → fall back to the text description. If both fail, see previous bullet.
  • Brand colors and style anchor colors conflict → brand wins. Always.
  • Reference-image generation rejects a prompt (safety / content) → soften the reference prompt and regenerate only the all-purpose reference; do not relabel it as a first frame.
  • The selected model cannot generate the requested ≤15s duration in one segment → report the model constraint or select another user-approved model. Do not silently turn a ≤15s request into a continuation chain.
  • Official Music 3.0 is unavailable or rejects the >15s BGM request → report the actual constraint and ask before switching 模型s. Never silently fall back to ElevenLabs Music v2.
  • Continuation join has visible drift → verify that segment N+1 uses the extracted tail frame from segment N, not the segment-1 reference or a regenerated opening. If the lineage is correct, regenerate only segment N+1 with a tighter local continuation prompt and end segment N on a more static moment next time.
  • A later segment looks like a new opening → stop delivery, inspect sequence, continuation_of, first_frame_file, and last_frame_file, and do not silently relabel the output as a continuation.
  • The >15s BGM is generated too early → finish and verify every numbered segment first, then generate exactly one BGM and mux it after ordered concatenation.
  • Video feels static, slideshow-like, or transition-poor → keep the same all-purpose reference or verified continuation frame, rewrite only the local motion_prompt using the causal beat chain in references/motion-prompt-writing.md, then regenerate that segment.
  • Final >15s audio ends before the video → remove -shortest, set -t to the requested total duration, and verify both stream durations with ffprobe.
  • 编辑器合成失败 → 交付原始片段与品牌档案。
  • The user wants > 15s → numbered continuation segments with tail-frame handoff and one post-generated BGM. See long-form.md.
  • The brand doesn't match any style → use the custom path. Don't force a fit.

Windows (win32) platform notes

Windows 下使用随 Skill 提供的合成流程并遵循平台路径分隔符;音乐默认仍只用于超过 15 秒的成片,参考素材和尾帧承接规则不变。

Examples

Input: "帮我做一个 15 秒的 UI 动效视频,产品叫 Notely,是个 AI 笔记 App"

→ No brand visual. Push back once for logo + screenshot. If provided, extract the brand profile and pass the screenshot as an all-purpose reference. Storyboard: 1 reference input + 1 motion prompt. Make one 15s video generation, do not force the screenshot to be the first frame, and do not generate extra BGM.

Input: "15s video for AURA smart home app, here's my logo and the app icon"

→ Brand intake: 2 images. Brand extraction: warm cream + orange + green + soft photography. Style anchor: custom. Storyboard: one 15s segment using the images as all-purpose references, with a five-beat motion prompt and no extra BGM. End card: real AURA logo.

Input: "15s Vision Pro style, here's my logo"

→ Brand intake: 1 image used as an all-purpose reference, not a first frame. Style anchor: G (user named "Vision Pro style"). Generate one 15s segment using style G's motion language but the brand's colors, with no extra BGM.

Input: "我想要 30s 的"

→ Continuation chain: segment 1 uses the all-purpose reference; segment 2 uses segment 1's extracted tail frame as its first frame. Only after both segments succeed, generate one 30s BGM and run the ordered concat + mux. See long-form.md.

Signals

GitHub stars
26
Last commit
Sep 2026
Advanced
Item type
skill
Key
ui-motion
Source
github.com/qxryz/workflowgenerator