Split Animated Short — user talking head (taller live band)

SkillDev tools

Turn a user-supplied talking-head clip (webcam, phone selfie, or recorded speaker) plus an optional script or source URL into a split-format vertical short, custom-generated animations in the top ~50%, the provided talking-head cropped into a taller rounded band across the bottom, kinetic captions synced to the clip's own voice, real product logos/screenshots whenever a named tool is spoken, real portraits of the notable figures behind any spoken niche (Claude/Anthropic, SpaceX, OpenAI, marketing, business, sales…), Giphy reaction GIFs on punchlines, a hyper-edit layer (keyword punch-ins, hyperframe inserts, whip cuts, shake) No HeyGen. Use when the user provides a talking-head video, webcam recording, selfie clip, or asks for the split animated format using their footage instead of a Megan/Danny/Kevin avatar.

Use Split Animated Short — user talking head (taller live band) in Claude, ChatGPT or Ahel Desktop

Free. Sign in, add Split Animated Short — user talking head (taller live band) and connect your AI. About a minute.

Also: Claude Code · Cursor · Codex

Then ask your AI: use the Split Animated Short skill

Details

Instructions available. Your AI can read the instructions. Execution depends on the setup they require.

Add Ahel to your AI once: Claude, ChatGPT, Cursor, Claude Code or Codex. Then ask it to use this.

Split Animated Short — user talking head (taller live band)Start free

What this skill tells your AI

The instructions your AI receives, as published by kevinbadi/open-edits in skills/split-animated-talking-head/SKILL.md and read by Ahel’s review.

HARD RULE (Kevin 2026-09-03): Final video is 1080×1920. Never 720×1280. Never scale the user's talking-head down to 704×640. Their clip is already 1080p. Instagram/TikTok/YouTube Shorts display at 1080×1920; a 720p master gets upscaled and looks like mud. If ffprobe on renders/final.mp4 is not 1080x1920, delete it and start over. Do not publish 720p.

Same flagship format as split-animated-short, except the bottom band is the clip the user hands you, not a HeyGen talking-photo. The clip's audio is the master timeline. Animations are authored per script — zero pixels from any reference reel except its words, if one is given.

Live talking-head footage needs more band height than the HeyGen 45% lock so hats/hair aren't clipped (Kevin 2026-09-01). Default band is 1056×960 at y=936 on the 1080×1920 canvas (caption baseline rests ON the accent line at y=BAND_Y-5, letters fully above the head; Kevin 2026-09-18/19). Round all four corners of the band mask — never BAND_H+60 (that square-cuts the bottom flush to the canvas). ~24px of canvas stays under the card.

PACK LOCK (Kevin 2026-09-03, opensource-100 recut): animations sit just above the captions, not up against the top of the frame. A top-hugging pack leaves a vacant band in the middle of the screen. Locked numbers:

AH = 770          # animation layer height
LAYER_Y = 90      # composite at (0, 90) → layer covers canvas y=90..860
PACK_DY = 210     # add to every place()/conveyor/orbit/ticker/ghost cy
CAP_Y = 867       # top of the caption zone (pack reference); the word itself is
                  # baseline on the accent line at BAND_Y-5, anchor="ls" (Kevin 2026-09-19)
BAND_Y = 936      # talking-head band (Kevin 2026-09-15: 960 was too low / clipped the bottom)

Layer ends ~7px above the captions. Pack numbers from clone-projects/opensource-100-animated/render.py (historical, not bundled). Never LAYER_Y=42 + AH=840 with heroes at cy≈180–280 — that was the vacant-middle bug. Never design at 853 and downscale, and never render the whole video at 720.

BEST EDIT YET (Kevin 2026-09-26): examples/social-agents/render.py is now the top reference (it outranks gitnexus). What made it: the user's b-roll used three ways (banner crop as the reveal on the product name; a live terminal card that starts zoomed on the command, then pulls back and scrolls with the output; single-line "LIVE FROM THE REPO" crops, left part of the line only so the text stays readable). Spoken feature lists become a ticking checklist card with one bespoke visual per item (real reel covers/thumbnails from brand-content + Downloads, mock inbox / caption / hashtag / SEO-climb cards). GIFs whose burned-in text matches the spoken line word for word. Stamps only on each section's dramatic beat. Reuse its Broll / crop_img / strip_card / broll_view(zoom=) helpers whenever b-roll is supplied.

STATUS QUO (Kevin 2026-09-19, "SOOOOO GOOOOD that should be the new status quo"): examples/gitnexus/render.py is the current gold reference. Everything below still applies; what made it land: Luckiest Guy captions resting on the accent line with a tight glow, the user's own screen recording as the b-roll hook (mystery "?" over it) and as per-word crop cuts later, a meme beat matched to the spoken line (Hangover calculation meme on "understand any piece of code") followed by a custom motion piece on the next word (spinning wireframe globe_card on "planet"), terminals that type the spoken action, a legend card lighting up per component, stamps on the punch words, a kinetic headline + lower third + counter. Match that energy on every cut.

LOOK LOCK (Kevin 2026-09-10, china-robots density + github-research-agent wash): gold-standard density / titles / GIFs: china-robots (render.py no longer on disk; use social-agents / gitnexus as code refs). Slot brightness lock: clone-projects/github-research-agent-animated/render.py (historical, not bundled) (Kevin 2026-09-10: "background is too dark" on the near-black CHAR wash). Copy china-robots energy, but use the lifted graphite textured_bg — lift patches, stronger accent glow, slot_vignette(..., a=22). Do not ship Megan paper/sage/coral, and do not use near-black CHAR=(14,15,18) as the slot wash. Locked colors (override after exec of animation_assets_and_beats.py). Type lock (Kevin 2026-09-15 + 2026-09-19): overlay / component text is Oswald Bold (templates/fonts/Oswald-Bold.ttf, loaded via templates/fonts.py as AR). Title cards, chips, huge_word, counters, name plates, ghosts. Kinetic captions are Luckiest Guy (templates/fonts/LuckiestGuy-Regular.ttf, CAP in fonts.py): CapCut-style chunky comic caps, white fill (KEY_COL fill on keywords), 7px black stroke, no drop shadow (Kevin 2026-09-19 "DOG" screenshot). Copy both ttfs. Never Arial in the slot. Terminals / code rain / tickers stay Menlo (ML) — Menlo.ttc must be opened with index=0 or glyphs double-print. Oswald is the condensed display face (the big AND over a UI). Copy fonts.py + fonts/Oswald-Bold.ttf into every workdir.

INK=(10,11,14)  SILVER=(198,204,214)  STEEL=(92,98,108)
GRAPHITE=(36,38,46)  CHAR=(48,50,58)
NEON=(57,255,132)  ORANGE=(255,122,24)  RED=(255,48,64)
CORAL=ORANGE  PAPER=GRAPHITE
AR=Oswald Bold    CAP=Luckiest Guy (captions only)    ML=Menlo (terminals only)

Captions: silver body, KEY_COL neon/orange/red on keywords, black shadow, Oswald Bold ~66px (same face as the title cards). Chips: graphite + accent bar, Oswald. Cards: graphite, Oswald. Slot accent bars like china-robots (orange/neon/red per section). Hero titles are ~960px wide. User B-roll in the slot should fill ~1000×500, not a small inset.

TOPIC LOGO CONVEYOR (Kevin 2026-09-10 — staple). The horizontal scroll of brand icons tied to the video (voice-agents: OpenAI / Cursor / Docker / Cloudflare / … under the faces) is required on every cut. Not optional decoration. See 5d below.

INTRO BILLBOARD (Kevin 2026-09-14): the hook / cold-open must read from across the room. Giant numbers, giant figure circles, giant logos — not a contact sheet of 150–220px heads under a title. See 5b size lock. 150px rounds are banned anywhere.

Do not call HeyGen. Do not use this skill when the user wants Megan, Danny, or Kevin generated. That is split-animated-short.

Reference compositor + asset library: this skill's templates/. Per video, create projects/<slug>-animated/ (slug from the clip filename or topic) and adapt.

Env (only if publishing / staging a feed item): INSFORGE_* in .env.local (repo or project root). Band prep is local: ffmpeg, OpenCV, Whisper.

Required input

The user must provide a talking-head clip. Accept:

  • an absolute/relative file path (~/Desktop/hook.mov, source/talking.mp4)
  • a file they dropped into the chat / workspace
  • a URL that is their talking-head (download with yt-dlp / curl immediately)

Refuse to start until the clip exists on disk. If they also drop a reference reel URL (the thing to recreate visually), that is optional and only supplies the script — never its pixels.

Vertical 9:16 is ideal. Horizontal/webcam is fine; prep will cover-crop.

Pipeline

  1. Workdir. mkdir -p projects/<slug>-animated/{source,audio,renders,frames} and cd there. Copy this skill's templates/ in (including apple_counter.py, light_fx.py, fonts.py, and fonts/Oswald-Bold.ttf).

  2. Prep the talking head (do this first — it is the timeline):

python3 "${CLAUDE_SKILL_DIR}/scripts/prep_talking_head.py" \
  --clip "/ABS/PATH/TO/CLIP" \
  --workdir . \
  --whisper

Writes source/talking_head.mp4 (native resolution, never downscaled), audio/narration.wav + .m4a, renders/talking_band45.mp4 (1056×960 @ 30fps, rounded later in compose), source/meta.json (DUR, crop), and audio/words.json.

Inspect band frames at 3 to 4 timestamps (the recorder re-frames the head mid-clip). Crop rule for live heads: keep the full hat / hairline in frame with a sliver of space above (backward caps sit higher than Haar's face box). Default --head-pad is 0.35. If the crown is clipped, lower --crop-top (and raise --crop-h to match the 1056:960 band aspect), but never below the clip's own black region: prep scans the whole clip for the first lit row and clamps the crop under it (content_top), and refuses a band with black top rows (verify_band_top). A black strip above the head is a failed prep, not a look. Do not letterbox the band. Do not scale the band below 1056×960.

  1. Get the script (words to animate).

    • Default: the talking-head transcript (audio/words.json) is the script.
    • Extra script given → use it, but time beats to the clip's Whisper stamps, not to the written script's imagined pacing.
    • Extra URL given (IG/YT/TikTok to recreate) → scrape caption / Whisper that source for copy only. IG: Apify apify~instagram-scraper (directUrls + resultsType: "details"), download videoUrl immediately. YouTube/TikTok: yt-dlp. Then persona-swap / product-name-fix against the caption. The talking-head clip still owns picture + audio.
  2. Guard on-screen text. Em/en dashes are banned from UGC overlays (AI tell). Censor spoken profanity on-screen Nick-thumbnail style (sh*t); leave the clip's audio uncensored. If the talking head is not Kevin and the optional source script names Kevin/Kev/Nick/Saraev, strip or swap those names in overlay copy.

  3. SOURCE REAL BRAND ART (do this before choreography). Walk the transcript for every named product, tool, model, or company (Claude, Google, GitHub, GitNexus, OmniRoute, GPT, Gemini, Llama, …). For each one, fetch a real mark into assets/logos/<slug>.png and a social/product card into assets/shots/<slug>.png when it exists:

python3 "${CLAUDE_SKILL_DIR}/scripts/fetch_brand_asset.py" \
  --name claude
python3 "${CLAUDE_SKILL_DIR}/scripts/fetch_brand_asset.py" \
  --name gitnexus --github abhigyanpatwari/GitNexus

Priority: Claude family is locked (Kevin 2026-09-03). Do not fetch a favicon for Claude, Claude Code, Opus, Sonnet, Haiku, Fable, or Anthropic. Copy the bundled marks and dance them:

python3 "${CLAUDE_SKILL_DIR}/scripts/fetch_brand_asset.py" --name claude
# writes assets/logos/claude.png (sphinx) plus claude-invader.png + claude-sphinx.png
  • Invader (templates/claude_marks.py → place_claude_invader): block Space-Invader. Idle dance + walk cycle. Look: pass look_at=(x,y) in layer pixels and the pupils slide toward that point (blinks on their own). Grab: grab="right"|"left"|"both", grab_at=(x,y), grab_t=0..1 (0 rest, ~0.55 reached, 1 holding/lifted). Pass hold= an RGBA (use skill_chip("SKILL.md")) to stick a card to the hand once grab_t > 0.55. Morph with morph=0..1 (starburst scatter).
  • Sphinx (place_claude_sphinx): official Anthropic starburst badge. Use as the spinning mark for Opus / Sonnet / Haiku / Anthropic, and as a secondary badge next to the invader. Call place_claude(layer, cx, cy, lt, word="claude", look_at=..., grab="right", grab_at=..., grab_t=...). Copy templates/claude_marks.py and assets/claude-*.png into the project.

For every other named product: official SVG/PNG (Wikimedia / simple-icons / site favicon) → GitHub avatar + OpenGraph card → Google s2 favicon. Convert SVG with qlmanage -t on macOS. A named product must appear as a logo, icon, screenshot, or orbiting mark — never as a text-only pill if a mark can be fetched. Punch near-white backgrounds to alpha; keep dark full-bleed logos as round badges.

5b. SOURCE THE FACES (Kevin 2026-09-05). Whenever the transcript names a company, model, or niche that has famous people behind it, put those people on screen — viewers who recognize Dario, Sam, Elon, GaryVee, Hormozi lock in harder than they do on a logo. Curated niche map lives in scripts/fetch_figure.py (--list): claude/anthropic/opus/… → Anthropic founders (Dario + Daniela Amodei, Jack Clark, Jared Kaplan, Chris Olah, Tom Brown); spacex → Musk, Shotwell, Mueller; openai/chatgpt → Altman, Brockman, Sutskever, Murati…; marketing → GaryVee, Hormozi, Godin, Neil Patel, Brunson, Kotler, Ogilvy; business/founders/money → Buffett, Bezos, Jobs, Musk, Cuban, Hormozi, Naval, Codie, Martell, Thiel; sales → Cardone, Belfort, Hormozi, Ziglar, Tracy, Robbins; plus google, nvidia, meta, microsoft, apple, tesla, amazon, xai.

python3 "${CLAUDE_SKILL_DIR}/scripts/fetch_figure.py" --niche claude --top 4
python3 "${CLAUDE_SKILL_DIR}/scripts/fetch_figure.py" --niche openai --niche marketing
python3 "${CLAUDE_SKILL_DIR}/scripts/fetch_figure.py" --name "Jensen Huang" --wiki Jensen_Huang --x nvidia --role "NVIDIA CEO"
# → assets/figures/<slug>.png (800x800 face-cropped) + <slug>.json {name, role, credit}

Sources: Wikipedia/Commons original → current X avatar (unavatar) → Wikidata name-match last (it has returned the wrong "Tom Brown"; eyeball any warn line). Always view the fetched portraits (contact sheet) before choreographing; swap with --name/--x --force if a crop is a stranger or a decade old. Choreography (templates/media_marks.py): figure_row(layer, slugs, cx, cy, lt, t0, place, names=True) for a founders lineup (2 faces on the intro, 2–3 later, staggered, alternating tilt), figure_spotlight(layer, slug, cx, cy, lt, t0, place) when ONE person is the beat (hero portrait + name plate + role), figure_card(slug, w=380) / figure_round(slug, 360, ring=CORAL) as marks. Face lands on the spoken word (word_onset(words, "anthropic")), never early. Pair the row with the brand mark (Claude → invader + sphinx beside the Amodeis). Names in plates are first names for rows, full name + role for spotlights.

Figure circle size lock (Kevin 2026-09-14: "make them much larger… intro screens should be eye catching with big #s, figures, logos"): the 150–220px rounds that kept shipping (phone-farm, anything-to-explain, open-source-trap hook) read as thumbnails, not icons. Floor:

  • Hook / intro / cold-open: these ARE the billboard. 1 face → figure_spotlight(..., size=560) or figure_round(..., 520–620). 2 faces → figure_row(..., size=480, gap=12). Never 3+ tiny heads on the first section — pick the two most iconic. Pair with a 960-wide title or a odometer_card / counter_card(..., h=440) or a 400–520px logo badge; the slot should feel huge, not busy. Conveyor marks stay 110–140 and sit under the giants, they do not replace them.
  • Body sections: single circle ≥280; 2–3-up row ≥320. Spotlight ≥480.
  • Banned: figure_round(..., 150) / size=170 / size=190 anywhere. If a 3-up row cannot fit at 320, drop to 2 faces and go bigger. Max one lineup per section. Intro faces may be the hero for up to 2.5s; later sections treat extra faces as supporting marks, still at the floor above — never shrunk to make room for a fourth.

5c. PULL REACTION GIFs (Giphy API, Kevin 2026-09-05). Punchlines, flexes, shocks and "I'm sorry CapCut" jokes get a real reaction GIF, not a text pill. GIPHY_API_KEY is in .env.local (repo or project root); the dashboard's favorited allow-list (gifs table) is available via --favorites.

python3 "${CLAUDE_SKILL_DIR}/scripts/fetch_giphy.py" --query "mind blown" --list   # look first
python3 "${CLAUDE_SKILL_DIR}/scripts/fetch_giphy.py" --query "mind blown" --slug mind-blown --pick 1
python3 "${CLAUDE_SKILL_DIR}/scripts/fetch_giphy.py" --favorites --query money --slug money-rain --mark-used
# → assets/gifs/<slug>/f0001.png … @30fps (≤4s, 720px wide) + meta.json

Pick by title from --list (Kramer / Jon Stewart "mind blown" over a 12-second anime loop). Eyeball frame 1 before using: drop GIFs with slurs/subtitles, the wrong person, or a joke that misses the line. Play with gif_card(..., w=400, frame="round") for pairs under a title. Rules (china-robots lock): one GIF per beat, 5–8 per video, never the hero for >2.5s, never covering a 960px title card. Later-section stack is title at cy≈125, pair of ~400px GIFs at cy≈450 (left/right). Hook frame 0 is the figure/number/logo billboard — do not cover it with a GIF pair unless the hook is B-roll. GIF audio is never mixed in. Rating pg-13.

5d. TOPIC LOGO CONVEYOR — staple (Kevin 2026-09-10). Every video gets a full-width horizontal scroll of the icons associated with this edit (the stack, products, companies, tools in the story — not random logos). Voice-agents showed OpenAI / Cursor / Docker / Cloudflare / … sliding under the faces. China-robots used Unitree / Tesla / NVIDIA / Meta / Boston Dynamics on "hardware + AI". That lane is the touch Kevin wants every time. Conveyor marks stay 110–140px; they are the stream, not the intro hero. On the hook, also place() 1–2 giant logos (360–520px) as billboard pieces next to the big figures / number (Kevin 2026-09-14).

Fetch 5–8 real marks in step 5 (fetch_brand_asset.py). Punch white pads. Then one lane, one y-band, via guarded_conveyor:

from media_marks import topic_marks
marks = topic_marks(["openai", "cursor", "docker", "cloudflare", "vapi"], size=120)
# y is authored cy; helper adds PACK_DY like china-robots
guarded_conveyor(layer, 400 + PACK_DY, max(0, lt - t0), marks,
                 conveyor_marks, name="topic-lane", speed=180, gap=22)

Rules:

  • At least one topic lane per video. Cold-open it on the hook when the stack is the point, or enter on the spoken product/category beat.
  • 5–8 marks, ~110–130px, graphite badges (dark slot). Full slot width.
  • One moving stream at a time (already locked). The lane owns its y-band; no second conveyor / orbit on the same band; lanes.reset() every frame.
  • Marks must match the video. Do not fill with leftover icons from another project. Eyeball each PNG (wrong logo / broken SVG → drop).
  • Do not sit the lane on the caption line. Center ≤ 770 - size/2 - 10 after PACK_DY (half-sliced icons are a known fail).
  1. AUTHOR THE ANIMATION (fresh per script): split narration into sections at spoken anchors using audio/words.json + source/meta.json DUR. Density + look are locked to the china-robots recut (china-robots (render.py no longer on disk; use social-agents / gitnexus as code refs), Kevin 2026-09-10): that is how animated every video in this format should feel. Target 6–10 elements on screen at once, not 2–3 lonely cards. Prefer motion graphics over copy: spinning logos, GitHub/product chrome (shot_card), terminals with ticking lines, skill-file grids, computer mockups, scan HUDs, force-directed graphs, brand badges, code rain, locks/skulls/bombs when the topic is hacks. Text cards caption a graphic, they are not the graphic. Do not default every beat to orbit_marks. The topic logo conveyor is mandatory (see 5d). Rotate the other motion: zigzag, explode, ticker, cascade, loop-arrows, Ken Burns. One orbit per video is plenty. Scene content must MATCH what's being said. Set DUR from meta.json, not the template's 62.9s.

    Cold open (locked, Kevin 2026-09-02): frame 0 of the video must already be packed. Scroll-stop on TikTok/Reels shows the first frame. Atmosphere (terminal, skill tiles, ghost word, a logo that is legal to show) uses place(..., t0=-0.4, dur=0.01) so the ease-in has already finished. Never start the first beat at t0=0 with dur=0.3 — that leaves frame 0 blank. Skip the white flash on the hook section for the same reason.

    Beat t0 is the spoken word, not the section start. For every punchline (a product name, a number, a claim), look up that word's start in audio/words.json and set t0 = word_start - section_start. Max lead is 0.1s. Never put "300" / "2B" / a named plugin on screen while he is still winding up to say it — that spoils the line. Do not preview later named products in an earlier section (e.g. don't flash OmniRoute during "three core plugins"). Cold-open atmosphere is the exception: generic (computer, skills, "installing") can be up before the first named product.

6b. EVERY SPOKEN NUMBER GETS A COUNTER (Kevin 2026-09-05). Money, GitHub stars, downloads, views, followers, counts, years, percentages, "10 p.m.", "3 times a day": if the transcript says a number, a count-up animation lands on it. Run the detector while planning beats:

python3 "${CLAUDE_SKILL_DIR}/scripts/number_beats.py"   # reads audio/words.json
#  26.40s   100  socials  <- "100 socials"
#  48.46s  $20/mo         <- "20 bucks a month"
#  51.02s  $2B            <- "$2 billion"

number_beats(words) (in hyper_edits.py) understands digits, spelled numbers ("twenty five", "a hundred"), scales ("2 billion", "9.3k"), money ("$20", "20 bucks"), "%", "3 times", "10 p.m.", "per month" and a trailing label word (stars, downloads, views, socials, agents…). It skips bare version digits ("Gemma 4", "LTX 2.5"). Fix Whisper mishears in words first (see the nano-banana / "2 .5" fixes) so the detector sees the real number.

Apple counter lock (Kevin 2026-09-29: "find and implement a modern apple inspired odometer that'll replace the one we currently have, it's bad"). Every hero number, big or small, renders through templates/apple_counter.py (counter_card and odometer_card in media_marks both delegate to it, so old call sites just work). It is modeled on iOS contentTransition(.numericText()) and keynote stat slides:

Shortened here. Read the whole file on GitHub.

Signals

GitHub stars
35
Forks
9
Last commit
Oct 2026

Ahel review

  • K6low
    bundled executables the agent is told to run
  • K1binfo
    installs-packages (in scripts/fetch_giphy.py)

Automated review, not a security audit. Ruleset v1+k2.

Advanced
Item type
skill
Key
split-animated-talking-head
Source
github.com/kevinbadi/open-edits