Split Animated Short — user talking head (taller live band)
SkillDev toolsTurn a user-supplied talking-head clip (webcam, phone selfie, or recorded speaker) plus an optional script or source URL into a split-format vertical short, custom-generated animations in the top ~50%, the provided talking-head cropped into a taller rounded band across the bottom, kinetic captions synced to the clip's own voice, real product logos/screenshots whenever a named tool is spoken, real portraits of the notable figures behind any spoken niche (Claude/Anthropic, SpaceX, OpenAI, marketing, business, sales…), Giphy reaction GIFs on punchlines, a hyper-edit layer (keyword punch-ins, hyperframe inserts, whip cuts, shake) No HeyGen. Use when the user provides a talking-head video, webcam recording, selfie clip, or asks for the split animated format using their footage instead of a Megan/Danny/Kevin avatar.
Use Split Animated Short — user talking head (taller live band) in Claude, ChatGPT or Ahel Desktop
Free. Sign in, add Split Animated Short — user talking head (taller live band) and connect your AI. About a minute.
Also: Claude Code · Cursor · Codex
Then ask your AI: use the Split Animated Short skill
Details
Instructions available. Your AI can read the instructions. Execution depends on the setup they require.
Account requirements not reviewed. Check the skill instructions before use; Ahel provides instructions and does not run this skill.
No other account needed.
Add Ahel to your AI once: Claude, ChatGPT, Cursor, Claude Code or Codex. Then ask it to use this.
What this skill tells your AI
The instructions your AI receives, as published by kevinbadi/open-edits in skills/split-animated-talking-head/SKILL.md and read by Ahel’s review.
HARD RULE (Kevin 2026-09-03): Final video is 1080×1920. Never 720×1280.
Never scale the user's talking-head down to 704×640. Their clip is already
1080p. Instagram/TikTok/YouTube Shorts display at 1080×1920; a 720p master
gets upscaled and looks like mud. If ffprobe on renders/final.mp4 is not
1080x1920, delete it and start over. Do not publish 720p.
Same flagship format as split-animated-short, except the bottom band is the
clip the user hands you, not a HeyGen talking-photo. The clip's audio is
the master timeline. Animations are authored per script — zero pixels from any
reference reel except its words, if one is given.
Live talking-head footage needs more band height than the HeyGen 45% lock
so hats/hair aren't clipped (Kevin 2026-09-01). Default band is 1056×960 at
y=936 on the 1080×1920 canvas (caption baseline rests ON the accent line at
y=BAND_Y-5, letters fully above the head; Kevin 2026-09-18/19). Round all four
corners of the band mask — never BAND_H+60 (that square-cuts the bottom
flush to the canvas). ~24px of canvas stays under the card.
PACK LOCK (Kevin 2026-09-03, opensource-100 recut): animations sit just above the captions, not up against the top of the frame. A top-hugging pack leaves a vacant band in the middle of the screen. Locked numbers:
AH = 770 # animation layer height
LAYER_Y = 90 # composite at (0, 90) → layer covers canvas y=90..860
PACK_DY = 210 # add to every place()/conveyor/orbit/ticker/ghost cy
CAP_Y = 867 # top of the caption zone (pack reference); the word itself is
# baseline on the accent line at BAND_Y-5, anchor="ls" (Kevin 2026-09-19)
BAND_Y = 936 # talking-head band (Kevin 2026-09-15: 960 was too low / clipped the bottom)
Layer ends ~7px above the captions. Pack numbers from
clone-projects/opensource-100-animated/render.py (historical, not bundled). Never LAYER_Y=42 +
AH=840 with heroes at cy≈180–280 — that was the vacant-middle bug. Never
design at 853 and downscale, and never render the whole video at 720.
BEST EDIT YET (Kevin 2026-09-26): examples/social-agents/render.py
is now the top reference (it outranks gitnexus). What made it: the user's b-roll used three ways
(banner crop as the reveal on the product name; a live terminal card that starts zoomed on the
command, then pulls back and scrolls with the output; single-line "LIVE FROM THE REPO" crops,
left part of the line only so the text stays readable). Spoken feature lists become a ticking
checklist card with one bespoke visual per item (real reel covers/thumbnails from
brand-content + Downloads, mock inbox / caption / hashtag / SEO-climb cards). GIFs whose burned-in
text matches the spoken line word for word. Stamps only on each section's dramatic beat.
Reuse its Broll / crop_img / strip_card / broll_view(zoom=) helpers whenever b-roll is supplied.
STATUS QUO (Kevin 2026-09-19, "SOOOOO GOOOOD that should be the new status quo"):
examples/gitnexus/render.py is the current gold reference. Everything
below still applies; what made it land: Luckiest Guy captions resting on the accent
line with a tight glow, the user's own screen recording as the b-roll hook (mystery
"?" over it) and as per-word crop cuts later, a meme beat matched to the spoken line
(Hangover calculation meme on "understand any piece of code") followed by a custom
motion piece on the next word (spinning wireframe globe_card on "planet"), terminals
that type the spoken action, a legend card lighting up per component, stamps on the
punch words, a kinetic headline + lower third + counter. Match that energy on every cut.
LOOK LOCK (Kevin 2026-09-10, china-robots density + github-research-agent wash):
gold-standard density / titles / GIFs: china-robots (render.py no longer on disk; use social-agents / gitnexus as code refs).
Slot brightness lock: clone-projects/github-research-agent-animated/render.py (historical, not bundled)
(Kevin 2026-09-10: "background is too dark" on the near-black CHAR wash).
Copy china-robots energy, but use the lifted graphite textured_bg — lift
patches, stronger accent glow, slot_vignette(..., a=22). Do not ship
Megan paper/sage/coral, and do not use near-black CHAR=(14,15,18) as the
slot wash. Locked colors (override after exec of animation_assets_and_beats.py).
Type lock (Kevin 2026-09-15 + 2026-09-19): overlay / component text is Oswald Bold
(templates/fonts/Oswald-Bold.ttf, loaded via templates/fonts.py as AR).
Title cards, chips, huge_word, counters, name plates, ghosts. Kinetic captions
are Luckiest Guy (templates/fonts/LuckiestGuy-Regular.ttf, CAP in fonts.py):
CapCut-style chunky comic caps, white fill (KEY_COL fill on keywords), 7px black
stroke, no drop shadow (Kevin 2026-09-19 "DOG" screenshot). Copy both ttfs.
Never Arial in the slot. Terminals / code rain / tickers stay Menlo (ML)
— Menlo.ttc must be opened with index=0 or glyphs double-print. Oswald is
the condensed display face (the big AND over a UI). Copy fonts.py +
fonts/Oswald-Bold.ttf into every workdir.
INK=(10,11,14) SILVER=(198,204,214) STEEL=(92,98,108)
GRAPHITE=(36,38,46) CHAR=(48,50,58)
NEON=(57,255,132) ORANGE=(255,122,24) RED=(255,48,64)
CORAL=ORANGE PAPER=GRAPHITE
AR=Oswald Bold CAP=Luckiest Guy (captions only) ML=Menlo (terminals only)
Captions: silver body, KEY_COL neon/orange/red on keywords, black shadow,
Oswald Bold ~66px (same face as the title cards). Chips: graphite + accent
bar, Oswald. Cards: graphite, Oswald. Slot accent bars like
china-robots (orange/neon/red per section). Hero titles are ~960px wide.
User B-roll in the slot should fill ~1000×500, not a small inset.
TOPIC LOGO CONVEYOR (Kevin 2026-09-10 — staple). The horizontal scroll
of brand icons tied to the video (voice-agents: OpenAI / Cursor / Docker /
Cloudflare / … under the faces) is required on every cut. Not optional
decoration. See 5d below.
INTRO BILLBOARD (Kevin 2026-09-14): the hook / cold-open must read from
across the room. Giant numbers, giant figure circles, giant logos — not a
contact sheet of 150–220px heads under a title. See 5b size lock. 150px
rounds are banned anywhere.
Do not call HeyGen. Do not use this skill when the user wants Megan,
Danny, or Kevin generated. That is split-animated-short.
Reference compositor + asset library: this skill's templates/. Per video,
create projects/<slug>-animated/ (slug from the clip filename or topic)
and adapt.
Env (only if publishing / staging a feed item): INSFORGE_* in
.env.local (repo or project root). Band prep is local: ffmpeg, OpenCV, Whisper.
Required input
The user must provide a talking-head clip. Accept:
- an absolute/relative file path (
~/Desktop/hook.mov,source/talking.mp4) - a file they dropped into the chat / workspace
- a URL that is their talking-head (download with yt-dlp / curl immediately)
Refuse to start until the clip exists on disk. If they also drop a reference reel URL (the thing to recreate visually), that is optional and only supplies the script — never its pixels.
Vertical 9:16 is ideal. Horizontal/webcam is fine; prep will cover-crop.
Pipeline
-
Workdir.
mkdir -p projects/<slug>-animated/{source,audio,renders,frames}andcdthere. Copy this skill'stemplates/in (includingapple_counter.py,light_fx.py,fonts.py, andfonts/Oswald-Bold.ttf). -
Prep the talking head (do this first — it is the timeline):
python3 "${CLAUDE_SKILL_DIR}/scripts/prep_talking_head.py" \
--clip "/ABS/PATH/TO/CLIP" \
--workdir . \
--whisper
Writes source/talking_head.mp4 (native resolution, never downscaled),
audio/narration.wav + .m4a, renders/talking_band45.mp4 (1056×960 @
30fps, rounded later in compose), source/meta.json (DUR, crop), and
audio/words.json.
Inspect band frames at 3 to 4 timestamps (the recorder re-frames the
head mid-clip). Crop rule for live heads: keep the full hat /
hairline in frame with a sliver of space above (backward caps sit higher
than Haar's face box). Default --head-pad is 0.35. If the crown is clipped,
lower --crop-top (and raise --crop-h to match the 1056:960 band aspect),
but never below the clip's own black region: prep scans the whole clip for
the first lit row and clamps the crop under it (content_top), and refuses
a band with black top rows (verify_band_top). A black strip above the
head is a failed prep, not a look. Do not letterbox the band. Do not scale
the band below 1056×960.
-
Get the script (words to animate).
- Default: the talking-head transcript (
audio/words.json) is the script. - Extra script given → use it, but time beats to the clip's Whisper stamps, not to the written script's imagined pacing.
- Extra URL given (IG/YT/TikTok to recreate) → scrape caption / Whisper
that source for copy only. IG: Apify
apify~instagram-scraper(directUrls+resultsType: "details"), downloadvideoUrlimmediately. YouTube/TikTok: yt-dlp. Then persona-swap / product-name-fix against the caption. The talking-head clip still owns picture + audio.
- Default: the talking-head transcript (
-
Guard on-screen text. Em/en dashes are banned from UGC overlays (AI tell). Censor spoken profanity on-screen Nick-thumbnail style (
sh*t); leave the clip's audio uncensored. If the talking head is not Kevin and the optional source script names Kevin/Kev/Nick/Saraev, strip or swap those names in overlay copy. -
SOURCE REAL BRAND ART (do this before choreography). Walk the transcript for every named product, tool, model, or company (Claude, Google, GitHub, GitNexus, OmniRoute, GPT, Gemini, Llama, …). For each one, fetch a real mark into
assets/logos/<slug>.pngand a social/product card intoassets/shots/<slug>.pngwhen it exists:
python3 "${CLAUDE_SKILL_DIR}/scripts/fetch_brand_asset.py" \
--name claude
python3 "${CLAUDE_SKILL_DIR}/scripts/fetch_brand_asset.py" \
--name gitnexus --github abhigyanpatwari/GitNexus
Priority: Claude family is locked (Kevin 2026-09-03). Do not fetch a favicon for Claude, Claude Code, Opus, Sonnet, Haiku, Fable, or Anthropic. Copy the bundled marks and dance them:
python3 "${CLAUDE_SKILL_DIR}/scripts/fetch_brand_asset.py" --name claude
# writes assets/logos/claude.png (sphinx) plus claude-invader.png + claude-sphinx.png
- Invader (
templates/claude_marks.py→place_claude_invader): block Space-Invader. Idle dance + walk cycle. Look: passlook_at=(x,y)in layer pixels and the pupils slide toward that point (blinks on their own). Grab:grab="right"|"left"|"both",grab_at=(x,y),grab_t=0..1(0 rest, ~0.55 reached, 1 holding/lifted). Passhold=an RGBA (useskill_chip("SKILL.md")) to stick a card to the hand oncegrab_t > 0.55. Morph withmorph=0..1(starburst scatter). - Sphinx (
place_claude_sphinx): official Anthropic starburst badge. Use as the spinning mark for Opus / Sonnet / Haiku / Anthropic, and as a secondary badge next to the invader. Callplace_claude(layer, cx, cy, lt, word="claude", look_at=..., grab="right", grab_at=..., grab_t=...). Copytemplates/claude_marks.pyandassets/claude-*.pnginto the project.
For every other named product: official SVG/PNG (Wikimedia / simple-icons /
site favicon) → GitHub avatar + OpenGraph card → Google s2 favicon. Convert
SVG with qlmanage -t on macOS. A named product must appear as a logo,
icon, screenshot, or orbiting mark — never as a text-only pill if a mark
can be fetched. Punch near-white backgrounds to alpha; keep dark
full-bleed logos as round badges.
5b. SOURCE THE FACES (Kevin 2026-09-05). Whenever the transcript names a
company, model, or niche that has famous people behind it, put those people
on screen — viewers who recognize Dario, Sam, Elon, GaryVee, Hormozi lock in
harder than they do on a logo. Curated niche map lives in
scripts/fetch_figure.py (--list): claude/anthropic/opus/… →
Anthropic founders (Dario + Daniela Amodei, Jack Clark, Jared Kaplan, Chris
Olah, Tom Brown); spacex → Musk, Shotwell, Mueller; openai/chatgpt →
Altman, Brockman, Sutskever, Murati…; marketing → GaryVee, Hormozi, Godin,
Neil Patel, Brunson, Kotler, Ogilvy; business/founders/money → Buffett,
Bezos, Jobs, Musk, Cuban, Hormozi, Naval, Codie, Martell, Thiel; sales →
Cardone, Belfort, Hormozi, Ziglar, Tracy, Robbins; plus google, nvidia,
meta, microsoft, apple, tesla, amazon, xai.
python3 "${CLAUDE_SKILL_DIR}/scripts/fetch_figure.py" --niche claude --top 4
python3 "${CLAUDE_SKILL_DIR}/scripts/fetch_figure.py" --niche openai --niche marketing
python3 "${CLAUDE_SKILL_DIR}/scripts/fetch_figure.py" --name "Jensen Huang" --wiki Jensen_Huang --x nvidia --role "NVIDIA CEO"
# → assets/figures/<slug>.png (800x800 face-cropped) + <slug>.json {name, role, credit}
Sources: Wikipedia/Commons original → current X avatar (unavatar) →
Wikidata name-match last (it has returned the wrong "Tom Brown"; eyeball
any warn line). Always view the fetched portraits (contact sheet)
before choreographing; swap with --name/--x --force if a crop is a
stranger or a decade old. Choreography (templates/media_marks.py):
figure_row(layer, slugs, cx, cy, lt, t0, place, names=True) for a
founders lineup (2 faces on the intro, 2–3 later, staggered, alternating tilt),
figure_spotlight(layer, slug, cx, cy, lt, t0, place) when ONE person is
the beat (hero portrait + name plate + role), figure_card(slug, w=380) /
figure_round(slug, 360, ring=CORAL) as marks. Face lands on the spoken
word (word_onset(words, "anthropic")), never early. Pair the row with
the brand mark (Claude → invader + sphinx beside the Amodeis). Names in
plates are first names for rows, full name + role for spotlights.
Figure circle size lock (Kevin 2026-09-14: "make them much larger… intro screens should be eye catching with big #s, figures, logos"): the 150–220px rounds that kept shipping (phone-farm, anything-to-explain, open-source-trap hook) read as thumbnails, not icons. Floor:
- Hook / intro / cold-open: these ARE the billboard. 1 face →
figure_spotlight(..., size=560)orfigure_round(..., 520–620). 2 faces →figure_row(..., size=480, gap=12). Never 3+ tiny heads on the first section — pick the two most iconic. Pair with a 960-wide title or aodometer_card/counter_card(..., h=440)or a 400–520px logo badge; the slot should feel huge, not busy. Conveyor marks stay 110–140 and sit under the giants, they do not replace them. - Body sections: single circle ≥280; 2–3-up row ≥320. Spotlight ≥480.
- Banned:
figure_round(..., 150)/size=170/size=190anywhere. If a 3-up row cannot fit at 320, drop to 2 faces and go bigger. Max one lineup per section. Intro faces may be the hero for up to 2.5s; later sections treat extra faces as supporting marks, still at the floor above — never shrunk to make room for a fourth.
5c. PULL REACTION GIFs (Giphy API, Kevin 2026-09-05). Punchlines, flexes,
shocks and "I'm sorry CapCut" jokes get a real reaction GIF, not a text pill.
GIPHY_API_KEY is in .env.local (repo or project root); the dashboard's favorited
allow-list (gifs table) is available via --favorites.
python3 "${CLAUDE_SKILL_DIR}/scripts/fetch_giphy.py" --query "mind blown" --list # look first
python3 "${CLAUDE_SKILL_DIR}/scripts/fetch_giphy.py" --query "mind blown" --slug mind-blown --pick 1
python3 "${CLAUDE_SKILL_DIR}/scripts/fetch_giphy.py" --favorites --query money --slug money-rain --mark-used
# → assets/gifs/<slug>/f0001.png … @30fps (≤4s, 720px wide) + meta.json
Pick by title from --list (Kramer / Jon Stewart "mind blown" over a
12-second anime loop). Eyeball frame 1 before using: drop GIFs with
slurs/subtitles, the wrong person, or a joke that misses the line.
Play with gif_card(..., w=400, frame="round") for pairs under a title.
Rules (china-robots lock): one GIF per beat, 5–8 per video, never the
hero for >2.5s, never covering a 960px title card. Later-section stack is
title at cy≈125, pair of ~400px GIFs at cy≈450 (left/right). Hook
frame 0 is the figure/number/logo billboard — do not cover it with a GIF
pair unless the hook is B-roll. GIF audio is never mixed in. Rating pg-13.
5d. TOPIC LOGO CONVEYOR — staple (Kevin 2026-09-10). Every video gets a
full-width horizontal scroll of the icons associated with this edit
(the stack, products, companies, tools in the story — not random logos).
Voice-agents showed OpenAI / Cursor / Docker / Cloudflare / … sliding
under the faces. China-robots used Unitree / Tesla / NVIDIA / Meta /
Boston Dynamics on "hardware + AI". That lane is the touch Kevin wants
every time. Conveyor marks stay 110–140px; they are the stream, not the
intro hero. On the hook, also place() 1–2 giant logos (360–520px)
as billboard pieces next to the big figures / number (Kevin 2026-09-14).
Fetch 5–8 real marks in step 5 (fetch_brand_asset.py). Punch white
pads. Then one lane, one y-band, via guarded_conveyor:
from media_marks import topic_marks
marks = topic_marks(["openai", "cursor", "docker", "cloudflare", "vapi"], size=120)
# y is authored cy; helper adds PACK_DY like china-robots
guarded_conveyor(layer, 400 + PACK_DY, max(0, lt - t0), marks,
conveyor_marks, name="topic-lane", speed=180, gap=22)
Rules:
- At least one topic lane per video. Cold-open it on the hook when the stack is the point, or enter on the spoken product/category beat.
- 5–8 marks, ~110–130px, graphite badges (dark slot). Full slot width.
- One moving stream at a time (already locked). The lane owns its
y-band; no second conveyor / orbit on the same band;
lanes.reset()every frame. - Marks must match the video. Do not fill with leftover icons from another project. Eyeball each PNG (wrong logo / broken SVG → drop).
- Do not sit the lane on the caption line. Center ≤
770 - size/2 - 10after PACK_DY (half-sliced icons are a known fail).
-
AUTHOR THE ANIMATION (fresh per script): split narration into sections at spoken anchors using
audio/words.json+source/meta.jsonDUR. Density + look are locked to the china-robots recut (china-robots (render.py no longer on disk; use social-agents / gitnexus as code refs), Kevin 2026-09-10): that is how animated every video in this format should feel. Target 6–10 elements on screen at once, not 2–3 lonely cards. Prefer motion graphics over copy: spinning logos, GitHub/product chrome (shot_card), terminals with ticking lines, skill-file grids, computer mockups, scan HUDs, force-directed graphs, brand badges, code rain, locks/skulls/bombs when the topic is hacks. Text cards caption a graphic, they are not the graphic. Do not default every beat toorbit_marks. The topic logo conveyor is mandatory (see5d). Rotate the other motion: zigzag, explode, ticker, cascade, loop-arrows, Ken Burns. One orbit per video is plenty. Scene content must MATCH what's being said. SetDURfrommeta.json, not the template's 62.9s.Cold open (locked, Kevin 2026-09-02): frame 0 of the video must already be packed. Scroll-stop on TikTok/Reels shows the first frame. Atmosphere (terminal, skill tiles, ghost word, a logo that is legal to show) uses
place(..., t0=-0.4, dur=0.01)so the ease-in has already finished. Never start the first beat att0=0withdur=0.3— that leaves frame 0 blank. Skip the white flash on the hook section for the same reason.Beat
t0is the spoken word, not the section start. For every punchline (a product name, a number, a claim), look up that word's start inaudio/words.jsonand sett0 = word_start - section_start. Max lead is 0.1s. Never put "300" / "2B" / a named plugin on screen while he is still winding up to say it — that spoils the line. Do not preview later named products in an earlier section (e.g. don't flash OmniRoute during "three core plugins"). Cold-open atmosphere is the exception: generic (computer, skills, "installing") can be up before the first named product.
6b. EVERY SPOKEN NUMBER GETS A COUNTER (Kevin 2026-09-05). Money, GitHub stars, downloads, views, followers, counts, years, percentages, "10 p.m.", "3 times a day": if the transcript says a number, a count-up animation lands on it. Run the detector while planning beats:
python3 "${CLAUDE_SKILL_DIR}/scripts/number_beats.py" # reads audio/words.json
# 26.40s 100 socials <- "100 socials"
# 48.46s $20/mo <- "20 bucks a month"
# 51.02s $2B <- "$2 billion"
number_beats(words) (in hyper_edits.py) understands digits, spelled
numbers ("twenty five", "a hundred"), scales ("2 billion", "9.3k"), money
("$20", "20 bucks"), "%", "3 times", "10 p.m.", "per month" and a trailing
label word (stars, downloads, views, socials, agents…). It skips bare
version digits ("Gemma 4", "LTX 2.5"). Fix Whisper mishears in words
first (see the nano-banana / "2 .5" fixes) so the detector sees the real
number.
Apple counter lock (Kevin 2026-09-29: "find and implement a modern apple
inspired odometer that'll replace the one we currently have, it's bad").
Every hero number, big or small, renders through templates/apple_counter.py
(counter_card and odometer_card in media_marks both delegate to it, so
old call sites just work). It is modeled on iOS contentTransition(.numericText())
and keynote stat slides:
Shortened here. Read the whole file on GitHub.
Signals
- GitHub stars
- 35
- Forks
- 9
- Last commit
- Oct 2026
Ahel review
K6low
bundled executables the agent is told to runK1binfo
installs-packages (in scripts/fetch_giphy.py)
Automated review, not a security audit. Ruleset v1+k2.
Advanced
- Item type
- skill
- Key
split-animated-talking-head- Source
- github.com/kevinbadi/open-edits
github.com/kevinbadi/open-edits
Related picks
Skill · wshobson
The pick for Pythonpython-pro
Skill · jeffallan
The pick for Pythonanimate-expo
Skill · emilkowalski
The pick for Expoexpo-dev-client
Skill · expo
The pick for Expoon-page-seo-checker
Skill · aaron-he-zhu
The pick for SEOseo-audit
Skill · aeonfun
The pick for SEO