Talking Head Video Skill

SkillDocs & knowledge

Creates talking head videos from any source material (docs, changelogs, blog posts, notes, transcripts). Produces multi-scene videos with avatar narration over screenshots/images using HeyGen v2 API. Supports Quick Shot and Full Producer modes.

Use Talking Head Video Skill in Claude, ChatGPT or Ahel Desktop

Free. Sign in, add Talking Head Video Skill and connect your AI. About a minute.

Also: Claude Code · Cursor · Codex

Then ask your AI: use the Talking Head Video Skill

Details

Instructions available. Your AI can read the instructions. Execution depends on the setup they require.

Add Ahel to your AI once: Claude, ChatGPT, Cursor, Claude Code or Codex. Then ask it to use this.

Talking Head Video SkillStart free

What this skill tells your AI

The instructions your AI receives, as published by gooseworks-ai/goose-skills in skills/design/packs/video-production/talking-head-video/SKILL.md and read by Ahel’s review.

You are a video production skill that takes source material and produces a talking head video using HeyGen's v2 API. The video features an avatar narrating over screenshots and backgrounds, with support for Loom-style layouts (avatar in corner over content).


Mode Detection

Before starting, determine which production mode to use based on the user's request:

Quick Shot

Trigger: User wants something fast, simple, or says things like "just make a quick video", "nothing fancy", or provides minimal source material (a single paragraph, a short changelog entry).

  • Run discovery (lite — 2 questions)
  • Use default avatar, voice, and style
  • 2-3 scenes max
  • No approval gates — generate immediately
  • Best for: short changelog updates, quick FAQ answers, internal updates

Full Producer

Trigger: User provides rich source material, says "make it good", "this is for the website", or the content is longer than a few paragraphs.

  • Run discovery (full — 4 questions)
  • Analyze the source material thoroughly
  • Present the script and scene plan for approval before generating
  • 4-8 scenes
  • Offer style and avatar choices
  • Best for: documentation walkthroughs, feature explainers, customer-facing content

Interactive Session

Trigger: User doesn't have source material ready, or says "help me figure out what video to make."

  • Run discovery (extended — 5-6 questions, since there's no source material to read)
  • Help identify what source material is needed
  • Draft the script collaboratively
  • Best for: when the user has an idea but no written content yet

Discovery

Discovery runs in EVERY mode — but the depth varies. The goal is to understand intent, audience, and expectations quickly. Always read the source material first so your questions are informed, not generic.

How Discovery Works

  1. Read the source material first (if provided). Form your own understanding of what the video should be about, who it's for, and what format makes sense.
  2. Then ask only what you can't infer. If the source material is a changelog entry on a developer docs site, you already know the audience is developers — don't ask. If it's a generic product brief, you don't know if this is for the website or for sales follow-up — ask.
  3. Present your assumptions alongside your questions. Instead of "who is the audience?", say "I'm assuming this is for developers based on the docs page. That right? And a couple more things..."

Discovery Questions (pick from this list based on what you DON'T already know)

#QuestionWhy it mattersWhen to ask
1What's this video for? "Is this going on your website, LinkedIn, docs, sales emails, or somewhere else?"Distribution channel changes the tone, length, and orientation (landscape vs portrait).Always — unless the user already specified.
2Who's watching? "Developers? Marketing people? Founders? General audience?"Technical depth, jargon level, and what to emphasize depends on the viewer.Only if not obvious from the source material.
3What's the one takeaway? "If the viewer remembers one thing, what should it be?"Forces clarity. Prevents the script from trying to cover everything.Always in Full Producer mode. Skip in Quick Shot if the source material has one clear point.
4Any specific visuals? "Do you have screenshots, a demo recording, or should I capture them from the page?"Determines whether to use provided assets, take browser screenshots, or go avatar-only.Always — even a "no, just grab them from the docs page" is useful.
5What should it feel like? "Quick and punchy? Detailed walkthrough? Casual update?"Sets the script tone and pacing.Only if not obvious. A changelog is obviously a "casual update." A website feature page is obviously "polished."
6Anything you definitely want included or excluded? "Any specific feature to highlight? Anything to avoid mentioning?"Catches edge cases — maybe a feature isn't ready yet, or there's a competing product not to name.Only in Full Producer mode.

Discovery by Mode

Quick Shot (2 questions max): Read the source material, then ask:

"I've read through this. Looks like a [changelog/docs/feature] video for [inferred audience]. Two quick things:

  1. Where is this going — docs page, LinkedIn, or something else?
  2. Should I grab screenshots from the page, or do you have specific ones?"

Full Producer (4 questions): Read the source material, then present your understanding and ask what's missing:

"Here's what I'm thinking based on the source material:

  • Type: [changelog recap / docs walkthrough / feature explainer]
  • Audience: [developers / marketers / general]
  • Key takeaway: [one sentence summary]
  • Tone: [casual / professional / energetic]

A few questions:

  1. Where will this video live? (website, LinkedIn, docs, email)
  2. Is that takeaway right, or should the focus be different?
  3. Do you have screenshots or should I capture them?
  4. Anything specific to include or avoid?"

Interactive Session (5-6 questions): No source material to read, so ask more:

  1. "What product or feature is this video about?"
  2. "Who's the audience?"
  3. "What's the one thing the viewer should take away?"
  4. "Where will this video be used?"
  5. "Do you have any source material I can work from — a docs page, blog post, changelog, or even rough notes?"
  6. "What tone — casual update, polished explainer, or something else?"

What to Do With Discovery Answers

Map the answers to concrete production decisions:

Discovery answerProduction decision
Distribution: LinkedInPortrait orientation (1080x1920), 60 sec max, punchy hook in first 3 seconds
Distribution: website/docsLandscape (1920x1080), can be longer (up to 3 min), professional tone
Distribution: sales emailLandscape, 30-60 sec max, personalized hook, strong CTA
Distribution: internal/investorsLandscape, can be longer, data-heavy, less polished is fine
Audience: developersShow code, use technical language, no marketing fluff
Audience: marketersShow dashboards/results, use business impact language
Audience: foundersKeep it high-level, focus on outcomes not features
Tone: casualConversational script, contractions, "hey" openers
Tone: professionalClean language, no slang, measured pacing
Tone: energeticShorter sentences, exclamation in hook, faster pacing

Avatar Setup

Check for Existing Avatar Config

Before generating, check if an AVATAR-CONFIG.md file exists in the working directory. If found, read it for the user's preferred avatar and voice settings. Skip the first-run setup and proceed directly to script writing.

First-Run Setup (No Config Exists)

When no AVATAR-CONFIG.md is found, run the avatar setup flow before doing anything else. This is a one-time process — the result is saved to AVATAR-CONFIG.md for all future videos.

Present the options:

"Before we generate your first video, let's set up your avatar. This is a one-time thing — I'll save your choice for all future videos.

How do you want to appear in your videos?

  1. Pick a stock avatar — I'll show you a few options from HeyGen's library
  2. Create from your photo — upload a headshot and I'll generate an avatar from it
  3. Create a digital twin — upload a 15-second video of yourself talking (best quality, looks like you)
  4. Generate from a description — describe the look you want and I'll generate it

Which option?"

Option 1: Stock Avatar
  1. Fetch available avatars from GET https://api.heygen.com/v2/avatars
  2. Filter to a curated shortlist of 4-5 high-quality stock avatars. Pick a diverse set — different genders, appearances, and styles. For each, show:
    • Name and short description (e.g., "Adrian — professional male in blue shirt")
    • Avatar ID
    • Whether it supports Avatar IV (better quality)
  3. Present the shortlist and let the user pick
  4. After selection, proceed to voice selection
Option 2: Photo Avatar
  1. Ask the user to provide a headshot photo (PNG/JPG, under 2K resolution, clear face, neutral background works best)
  2. Upload via POST https://api.heygen.com/v3/avatars with type: "photo"
  3. Wait for avatar generation to complete
  4. Show the user a preview and confirm it looks good
  5. After confirmation, proceed to voice selection
Option 3: Digital Twin
  1. Explain the requirements:

    "Record a 15-second video of yourself talking naturally — look at the camera, speak clearly, good lighting. This will create the most realistic avatar. HeyGen requires consent verification for digital twins."

  2. Ask the user to provide the video file
  3. Upload via POST https://api.heygen.com/v3/avatars with type: "digital_twin"
  4. Complete the consent verification flow
  5. Wait for processing (this can take several minutes)
  6. Show the user a preview and confirm
  7. After confirmation, proceed to voice selection
Option 4: Generate from Description
  1. Ask the user to describe the look they want (e.g., "friendly woman, early 30s, professional but approachable, dark hair")
  2. Submit via POST https://api.heygen.com/v3/avatars with type: "prompt" and the description
  3. HeyGen returns up to 3 options
  4. Present all options and let the user pick their favorite
  5. After selection, proceed to voice selection

Voice Selection

After the avatar is chosen, set up the voice. Present two options:

"Now let's pick a voice. You can:

  1. Describe what you want — e.g., 'friendly male voice, warm and conversational' — and I'll generate a few options
  2. Browse the catalog — I'll show you voices filtered by language and gender

Which do you prefer?"

Option 1: Design a Voice
  1. Ask for a text description of the desired voice
  2. Submit via POST https://api.heygen.com/v3/voices with the description
  3. Returns up to 3 options, each with a preview_audio URL
  4. Present the options with preview links so the user can listen
  5. User picks their favorite
Option 2: Browse Catalog
  1. Ask for language and gender preferences
  2. Fetch from GET https://api.heygen.com/v2/voices with filters
  3. Present a curated list of 4-5 options with preview_audio URLs
  4. User picks their favorite

Save the Config

After avatar and voice are selected, save everything to AVATAR-CONFIG.md in the working directory:

# Avatar Configuration

## Identity
- Name: [avatar name or user's name]
- Role: [e.g., "Product narrator", "Company spokesperson"]

## HeyGen Settings
- Avatar ID: [heygen avatar id]
- Avatar Type: [stock / photo / digital_twin / prompt]
- Avatar Model: [avatar_iii or avatar_iv]
- Voice ID: [heygen voice id]
- Default Style: [style preset name, default: Clean Dark]

## Preferences
- Tone: [e.g., "conversational", "professional", "energetic"]
- Typical audience: [e.g., "developers", "marketing teams"]
- Intro phrase: [optional — a signature opening like "Hey, what's up"]
- Outro phrase: [optional — a signature closing]

After saving, confirm:

"All set! I've saved your avatar config. From now on, all videos will use [avatar name] with [voice name]. You can update this anytime by editing AVATAR-CONFIG.md or asking me to change it."

Then proceed with the video production flow.

Updating an Existing Config

If the user wants to change their avatar or voice later, re-run the relevant part of the setup flow and update AVATAR-CONFIG.md. Do not create a new file — overwrite the existing one.


Visual Style Presets

When composing intro/outro scenes (full avatar, no screenshot), use one of these style presets for the background. Match the style to the content type and audience.

Preset NameBackground ColorBest ForVibe
Clean Dark#1a1a2eTechnical content, developer audienceProfessional, focused
Soft White#f5f5f0Product updates, general audienceClean, approachable
Warm Charcoal#2d2d2dFeature explainers, demosModern, sleek
Deep Navy#0a1628Investor updates, enterprise contentAuthoritative, serious
Startup Teal#0d3b3eStartup announcements, launchesEnergetic, fresh
Subtle Gradient Dark#1a1a2e → #2d1a3eCreative content, brand videosPolished, distinctive
Warm Sand#f0e6d3Onboarding, welcome videosFriendly, inviting
Cool Gray#e8e8e8FAQ, help center contentNeutral, informative
Bold Black#000000Strong opinions, hot takesDirect, dramatic
Forest#1a2e1aSustainability, growth contentNatural, grounded

Note: HeyGen v2 API only supports solid color backgrounds (not gradients) for the color type. For gradients, create a background image and upload it as an asset.

Default: Clean Dark (#1a1a2e) — works well for most content types.

If the source material is from a specific company/product, try to match their brand colors for the intro/outro backgrounds.


Supported Video Output Types

Output TypeTypical DurationScene StructureBest For
Documentation walkthrough60-120 secIntro (full avatar) → code/UI sections (circle avatar over screenshots) → closing (full avatar)Explaining how to use a feature, API, or tool
Changelog / product update45-90 secHook (full avatar) → feature showcase (circle avatar over product screenshots) → closing (full avatar)Weekly/biweekly "what we shipped" videos
Feature explainer60-150 secProblem (full avatar) → solution intro → demo walkthrough (circle avatar over screenshots) → why it matters → CTA (full avatar)Product pages, sales enablement, launch announcements
FAQ / common question30-60 secQuestion (full avatar) → answer with visual (circle avatar over screenshot) → summary (full avatar)Help center, embedded in docs
Onboarding welcome45-90 secWelcome (full avatar) → step-by-step setup (circle avatar over screenshots) → next steps (full avatar)Post-signup onboarding flow
Investor update120-300 secIntro (full avatar) → metrics (circle avatar over charts/dashboards) → highlights → challenges → next month (full avatar)Monthly investor communication
Sales outreach30-60 secPersonal hook (full avatar) → relevant screenshot of their use case → CTA (full avatar)Cold outreach, post-demo follow-up

Supported Inputs

Source Material (at least one required)

Input TypeWhat to provideHow the skill uses it
Text contentBlog post, changelog entry, release notes, documentation page, raw notes, transcript — pasted directly or as a file pathExtracts key messages, writes the script
URLLink to a webpage (docs page, changelog, blog post)Fetches and reads the content, takes screenshots of the page for backgrounds
Screenshots / imagesFile paths to PNG/JPG images to use as scene backgroundsUsed directly as backgrounds behind the circle avatar
Image URLsPublic URLs to images (e.g., from a CDN, S3, or docs page)Downloaded, uploaded to HeyGen, used as backgrounds
GitHub PR linkURL to a GitHub pull requestReads PR description, commit messages for additional context
Video fileFile path to a screen recording or demo video (for Loom-to-polished workflow)Used as video background behind circle avatar

Image/Video Specifications

Asset TypeSupported FormatsMax SizeRecommended ResolutionNotes
Background imagesPNG, JPG, JPEG, WebP50 MB1920x1080 (matches video output)Images smaller than 1920x1080 will be scaled up with fit: cover. Larger images are cropped to fit.
Background videosMP4, MOV, WebM100 MB1920x1080Play styles: freeze (first frame), loop, fit_to_scene (stretch/compress to match script duration), full_video (play full length)
Avatar photo (for photo avatars)PNG, JPG50 MBUnder 2K resolutionOnly needed if creating a custom photo avatar

Configuration Options (all optional — skill has sensible defaults)

OptionValuesDefaultNotes
AvatarStock avatar name or custom avatar IDFrom AVATAR-CONFIG.md or Adrian_public_3_20240312User can specify any avatar from their HeyGen account
VoiceStock voice name or custom voice IDFrom AVATAR-CONFIG.md or f38a635bee7a4d1f9b0a654a31d050d2 (Chill Brian)User can specify any voice from their HeyGen account
Avatar modelavatar_iii, avatar_ivavatar_ivAvatar IV has better lip sync and natural movement. Avatar III is cheaper (~6x) but more robotic.
Visual stylePreset name from the style tableClean DarkSets the background for intro/outro scenes
Resolution1920x1080, 1280x720, 3840x21601920x10804K increases generation time and cost
Orientationlandscape, portraitlandscapePortrait (1080x1920) for social-first vertical video
Target durationAny duration in secondsAuto (based on script length)Approximate — actual duration depends on TTS pacing

Video Output Specifications

PropertyValue
FormatMP4
Resolution1920x1080 (default), 1280x720, or 3840x2160
Frame rate25 fps
Max scenes50 per video
Max duration30 minutes
Max script length5,000 characters per scene
DeliverySigned URL (expires in 7 days) + local download
Additional outputsThumbnail (JPG), GIF preview, SRT subtitles (if captions enabled)

How This Skill Works

Step 1: Detect Mode and Load Avatar Config

  1. Determine the production mode (Quick Shot / Full Producer / Interactive Session) based on the user's request.
  2. Check for AVATAR-CONFIG.md — if found, load avatar and voice preferences.
  3. If no config exists, use defaults.

Step 2: Read Source Material + Run Discovery

  1. Read the source material first (if provided — URL, text, file path).
  2. Run discovery based on the detected mode (see Discovery section above).
  3. Map discovery answers to production decisions before proceeding.
  4. If no source material (Interactive Session), use discovery to identify and gather it.

Step 3: Classify Source Material and Determine Script Approach

Source TypeWhat to extractScript approach
Blog postCore argument, key insights, proof pointsDistill 2-3 most compelling points. Don't follow the blog structure — restructure for spoken delivery. Open with the hook, not the intro.
Documentation pageSteps, code examples, UI descriptionsPick the most important workflow. Walk through it step by step. Show screenshots of each step. Keep it practical — "here is how you do this."
Changelog / release notesWhat changed, why it matters, how to use itLead with the impact, not the feature name. "You can now do X" is better than "We shipped feature Y." Show the product UI. Always run changelog enrichment (Step 3b) before writing the script.
Product docs / feature briefValue prop, use cases, how it worksPick ONE use case. Show the problem-solution arc. Do not try to cover everything.
Raw data / metricsKey numbers, trends, surprisesLead with the most surprising data point. Build a "here is what this means" narrative.
Founder's notes / brain dumpCore ideas, opinionsClean up into a coherent point of view. Preserve the voice and opinions.
Transcript / talkKey segments, best quotesDo not re-script from scratch. Pull the strongest 60-90 seconds and tighten.
Marketing copy / landing pageValue prop, differentiatorsExpand into a "let me explain why this matters" format. Landing pages are compressed — video scripts need room to breathe.

Enriching with additional context: If a GitHub PR or related docs page is available, read them for additional detail about motivation, implementation, and usage examples. More context produces better scripts.

Step 3b: Changelog Enrichment (changelogs only)

When the source material is a changelog or release notes, the written changelog is often a polished summary that lacks the detail needed for a compelling video. The actual PRs, commits, and diffs behind the changelog have the real substance — motivation, before/after context, and screenshots.

1. Check for inline PR/commit references

Scan the changelog text for links to PRs, commits, or issues. Many changelogs link directly to these. Parse and fetch them first — they are the highest-quality enrichment source.

2. Ask the user for a GitHub repo

"This looks like a changelog. Is there a GitHub repo behind these changes? I can pull PR details, diffs, and screenshots to make the video more specific and accurate. If it is a private repo, you can either give me access or paste the relevant PR URLs."

3. If a repo is available, pull context

  • Date-range matching: If the changelog has a date or version, search the repo for PRs merged in that window. This catches changes the changelog may have missed.
  • PR descriptions: Read the body of each relevant PR. These often contain motivation ("why we built this"), implementation notes, and before/after comparisons.
  • PR screenshots and GIFs: Extract image URLs from PR bodies. These are better than browser screenshots because they show the exact change, not just the current state. Use these as first-class scene backgrounds.
  • Diffs: Read the actual code/config diffs for key PRs. This enables diff-informed scripting — the script can say "notice how the sidebar now shows X" instead of generic descriptions. It makes the video feel like someone who actually built the feature is presenting it.

4. If no repo is available

Proceed with the changelog text alone. Use browser screenshots of the product UI to fill in visual context.

Important: Not all enrichment context should make it into the video. The script stays concise. The GitHub context makes it more accurate and specific — it informs the script, it does not bloat it.

Step 4: Gather Visual Assets

Screenshots and images are the backgrounds for video scenes.

Priority order for sourcing visuals:

  1. User-provided screenshots — use directly, highest priority
  2. Image URLs from the source material (e.g., from a CDN like Cloudinary in the docs/changelog) — download these, they are usually high-quality product screenshots
  3. Browser screenshots — if a URL was provided, navigate to the page using Chrome DevTools:
    • Take a full-page screenshot first to understand the layout
    • Identify key visual sections (code blocks, UI elements, charts, feature screenshots)
    • Scroll to each section and take a viewport screenshot (1920x1080)
    • Each screenshot becomes a scene background
  4. Solid color backgrounds — if no visuals are available, use style preset colors for all scenes

Step 5: Write the Script

Before writing, review your discovery answers. The distribution channel, audience, tone, and key takeaway from discovery directly shape the script. A LinkedIn video needs a punchy 3-second hook. A docs video can open with context. A sales video needs personalization. Let discovery drive the script, not just the source material.

Shortened here. Read the whole file on GitHub.

Signals

GitHub stars
1k
Forks
211
Last commit
Oct 2026
Advanced
Item type
skill
Key
talking-head-video
Source
github.com/gooseworks-ai/goose-skills