Scenario UGC Creator Video

SkillMedia

Lets your agent create short creator-style videos like ads, testimonials, demos, and social clips from a photo and a script.

Available today. Use it from your connected AI after setup.

Add ahel to your AI once: Claude, ChatGPT, Cursor, Claude Code or Codex. Then ask it to use this.

Then ask your AI: use the Scenario UGC Creator Video skill

About this skill

Use when producing UGC-style creator video with Scenario: a talking-head ad or testimonial from a portrait and a script, a founder clip, a product demo or unboxing, a reaction or before-after cut, a faceless voiceover over b-roll, or vertical social video for TikTok, Reels, or Shorts that must feel

What this skill tells your AI

The instructions your AI receives, as published by scenario-labs/skills in skills/scenario-ugc/SKILL.md and read by ahel’s review.

Overview

UGC is a register, not a length: content that reads as a person talking into their own phone, not a brand talking through a camera crew. Everything in this skill serves that register, and most failures come from importing ad craft into it. This skill routes the production; mechanics live in the sibling skills named per lane. Connection and the core loop: see the scenario skill in this repo. If a sibling skill named here is missing from your available skills, ask the user to install it (npx skills add scenario-labs/skills --skill <name>); unattended, proceed from tool schemas and flag the gap.

Two rules are non-negotiable. The product is never generated: demo shots start from an uploaded photo or footage of the real product (scenario-product-shots for stills). And the words are never invented: no fabricated testimonials, review counts, metrics, medical or financial claims, or legal copy; speak only lines the user supplied or approved, and prefer observable statements ("the texture looks lighter") over claims ("this cures acne").

Quick reference: route by speaker lane

Discover members with recommend, passing the lane's capability and the user's own words: it ranks by measured cost and latency and names the purpose-built pick, where a capability-worded search returns hundreds of keyword hits with nothing to choose between them. Read next_step before taking a pick, per the scenario skill. Keep search for a member you can already name. Never assert a generative model's id as a constant. Scenario's own single-purpose tool models are named outright below: there is exactly one of each, so discovering them would only re-derive a constant.

LaneRouteContract
Portrait plus speech audioTalking-avatar members (img2video)scenario-kling
Existing footage, new wordsLipsync members (video2video), audio or text, never bothscenario-kling
Recorded delivery, different faceMotion-control members (video2video, say "motion control" in the recommend prompt or lipsync members come back): character still plus the recording as driving video, keepOriginalSound keeps the voice; output lands at the member's own size, the still lending only its orientationscenario-kling
Generated creator speaking nativelyNative-audio video families (txt2video), dialogue in quotes with the delivery named (pace, tone, one gesture)scenario-veo, scenario-seedance, scenario-kling
Faceless voiceoverB-roll clips (txt2video) plus TTS narration (txt2audio)scenario-video, scenario-elevenlabs
Product demo insertsImage-to-video off the uploaded product stillscenario-product-shots, scenario-video
Captions, cut, 9:16 masterAssembly tool models; text cards as image layersscenario-video-assembly, scenario-text-overlay

Script in six spoken beats: hook (one concrete tension or result), context (why this speaker cares), product moment, proof (visible demo or a user-supplied fact), turn (objection answered or before-after), close (soft CTA). Write for the mouth, not the page: contractions, false starts allowed, no taglines. Spoken pace runs near 2.5 words a second, so a 25-second ad is roughly 60 words, but that is a first guess and avatar members undershoot it: observed rates run 1.76 to 1.94 words a second, and one member padded a 60-word script with 21s of silence instead. Time the returned clip before assembling, then trim the silences where there are any and cut the script where the delivery is simply slow. The compositor has no cut inside a layer, so a trim is one layer per kept segment, each with its own trimStart and duration.

Keep the register in every visual prompt: phone-height framing, available light, a real location with clutter, natural skin texture, one handheld drift at most. Cinematic grammar (dolly moves, golden-hour rim light, shallow anamorphic looks, graded color, retouched skin) reads as an ad and kills belief. Compose 9:16 natively, reframing any non-vertical still a lane consumes (the avatar portrait, a motion-control character still, the product photo) per scenario-formats before generating; a cropped 16:9 master frames like television.

Worked example: 25-second founder ad from a portrait

  1. Brief once: offer, platform, runtime, the facts the founder may claim, tone. Collect the portrait and the real product photo, then run without stopping.
  2. Create this run's collection before the first generation (collection_create, catalog write lane, name only), then collection_add_assets each keeper as it lands; its returned itemCount is the receipt.
  3. Draft the six beats at about 60 words and confirm the wording with the user; the script is a claims surface, not just copy. Unattended, keep every line to wording the brief already supplied and cut any beat that would need a new claim.
  4. Voice: upload_asset the founder's recorded narration, or generate TTS per scenario-elevenlabs when they want a stand-in voice they approved.
  5. recommend with capability="img2video" and the brief in the user's words; pick from its ranked list by input contract (a script-taking member when the speech exists only as text, an audio-taking one when a recording exists), then by whether the member can be told not to burn in captions: some hallucinate gibberish subtitles that no negative prompt suppresses, and the ones with a switch cost several times more, which recommend prices for you.
  6. upload_asset the portrait. Prompt the speaker's hands empty and the set free of products: avatar members invent props and label type. model_run with dry_run=true first: avatar members sit far apart on price, so quote before spending.
  7. Run for real with wait=false, then jobs_wait with the returned job id, re-called with pending_job_ids on timeout, never a second model_run.
  8. Demo insert: animate the uploaded product photo with an image-to-video member (scenario-video), 3 to 5 seconds, one micro-move. Whether aspectRatio survives a start image is per member and the schema note can be wrong either way, so read width and height off the returned asset rather than assuming the ratio held.
  9. Assemble with model_scenario-compose-video per scenario-video-assembly: talking head as the spine, insert cut over beats three and four, product-name card from scenario-text-overlay as an image layer, dropped outright when no approved product name exists, since inventing one is the fabrication the brief forbids. Every overlay layer needs an explicit width and height or the compositor scales it, and durationMode: "custom" pins the master's length against image layers that would stretch it. A succeeded compose job is not proof the layers drew or landed where sent: pull a frame from the composite and confirm the card is present and placed before delivering. Captions come last, on the finished cut, from the captioner scenario-video-assembly discovers with recommend (two exist, so no fixed id), positioned inside the platform's safe zone: the compositor has no text layer.
  10. asset_display the master. Gate every generated clip, the talking head included: take the frames from the platform, the free firstFrame and lastFrame off asset_get first and model_scenario-video-to-image-seq when a mid-clip frame is what settles it, then verify them against the uploads (scenario-asset-analysis); a local extraction is not a substitute at any budget, because it yields no asset to file or audit and the gate stops being traceable; an invented product in the speaker's hands or legible generated type fails a clip exactly like label drift on the insert. Report spend by summing this run's own job records by job id: jobs_list is project-scoped and over-reports, and usage's headline figure is project-lifetime.

Common mistakes

  • Writing ad copy and handing it to a mouth: alliterative taglines collapse on a talking head; read the script aloud before generating.
  • Fabricating social proof: an invented "10,000 five-star reviews" is a claim the user never made; keep numbers and testimonials to supplied wording.
  • Prompting the creator like a commercial: tripod framing, perfect light, and a spotless studio kitchen read as an ad; imperfection is the format.
  • Passing both audio and text to a lipsync member: exclusive inputs; pick one.
  • Skipping dry_run on avatar and lipsync runs: per-member pricing varies too widely to guess.
  • Generating the product or any on-screen text: the product comes from the uploaded photo, type is overlaid in assembly. Native-speech members can caption the quoted line, and those captions collide with the ones the captioner places in the safe zone, so end every dialogue prompt with "no subtitles, no on-screen text". There the line is the only lever and costs nothing, and it is not a guarantee (one native-speech member captioned anyway at authoring time), so pull a frame before compositing; avatar members can hallucinate subtitles no prompt suppresses, which is why step 5 prices their caption switch.
  • One 40-second b-roll or generated-creator take: generate per-beat clips and cut on the beat turns; single long takes drift and cost more to retry. A talking-head script stays one take, trimmed in assembly.
  • Cropping a landscape master to 9:16: heads and captions land outside the safe zone; compose vertical from the start.

Signals

GitHub stars
681
Forks
82
Last commit
Sep 2026
Advanced
Item type
skill
Key
scenario-ugc
Source
github.com/scenario-labs/skills