Scenario Caption Studio
SkillMediaUse when a video needs its spoken words on screen through Scenario via MCP: burned-in styled captions for a TikTok, Reels, or Shorts cut, ad captions for sound-off feeds, YouTube subtitles, an SRT sidecar, transcription of a clip's audio, captions translated into another language, karaoke or word-by-word styles, or restyling and correcting an existing transcript. Keywords: captions, subtitles, SRT, transcribe, karaoke, word-by-word, burn in, closed captions, translate video, caption style.
Available today. Use it from your connected AI after setup.
No other account needed.
Add ahel to your AI once: Claude, ChatGPT, Cursor, Claude Code or Codex. Then ask it to use this.
Then ask your AI: use the Scenario Caption Studio skill
What this skill tells your AI
The instructions your AI receives, as published by scenario-labs/skills in skills/scenario-caption-studio/SKILL.md and read by ahel’s review.
Overview
Caption Studio is one tool model, model_scenario-caption-studio: a video in, its speech transcribed (Whisper) or an existing SRT applied, styled captions out, burned into the picture or delivered as a soft track and an .srt sidecar. It translates into 18 languages and styles captions three ways. Running that one member is this skill's whole purpose, so the id is named rather than discovered and model_schema_get starts the flow directly; availability differs per team, so a member the team lacks is a gap to flag, not a cue to substitute. Connection and the core loop: see the scenario skill.
Captioning is the last pass on a finished cut: assemble first (scenario-video-assembly), then caption the master once. Captions carry the transcript only; text that must appear letter-perfect without being spoken (CTAs, prices, legal supers) is scenario-text-overlay territory. If a sibling skill named here is missing from your available skills, ask the user to install it (npx skills add scenario-labs/skills --skill <name>); unattended, proceed from tool schemas and flag the gap.
Ask once, then caption
The destination decides nearly every parameter, so collect one round of answers before touching the schema: where the video ships (a sound-off mobile feed, a paid placement, a seated long-form viewer), whether the spoken language stays or translates (targetLanguage, auto keeps it), brand colors if any, and which deliverable the platform wants. The deliverable is three switches: burned-in pixels are outputSubtitles: "video_image" (the default), the toggleable track is "video_data", the sidecar file is outputSrt: true, and an SRT-only pass is that plus outputVideo: false, the first pass when the words must be letter-perfect: it priced the same as a burn-in at authoring time, so it buys certainty rather than savings, proving the words before any pixels are paid for. Then map the answers:
| Destination | Style | Segmentation | Position | Output |
|---|---|---|---|---|
| Social mobile short (9:16 Shorts, Reels) | tiktok-bouncy or word-pop; karaoke-fill when music drives | maxSegmentWords 3 to 5; 1 with a karaoke preset | middle: platform UI and native auto-captions own the bottom | Burn in |
| Ad short, performance cut | modern-chip or minimal-underline, accents set to brand color | 3 to 7 words per cue | bottom, or top when an end card or overlay sits below | Burn in for sound-off feeds; add outputSrt for the platform's caption upload |
| YouTube long-form, tutorial, interview | Default look or cinematic-fade; restraint reads as professionalism | maxLines 2, maxSegmentChars 84 (two 42-character lines, the broadcast convention) | bottom | outputSrt for the platform's closed captions; burn in only for re-embeds |
| Cinematic piece, trailer, festival cut | cinematic-fade | Sentence-length cues, maxSegmentDuration about 6 | bottom | Burn in |
Rows are authoring-time starting points to confirm with the user, not platform contracts; unattended, the task's own instructions answer the interview and the matching row's defaults fill what they leave unsaid. The per-placement safe zones behind the position column live in scenario-formats. middle renders at frame center, which on a centered talking head is the mouth: caption a selfie cut only once its face sits in the upper third (reframe in assembly), because the schema offers no lower-third position. An uploaded track beats burn-in for long-form because viewers toggle and restyle it, assistive tech reads it, and platforms index it for search; burn-in wins wherever the style is the point or a track cannot travel with the file.
The style ladder
Three tiers: stylePreset picks a ready-made look (7 presets, empty for the default); stylePrompt describes a look in plain words and builds a matching style (it carries cost_impact); themeTsx supplies a full custom theme that replaces the preset, with stylePrompt then refining that theme. There is no font parameter: type rides inside the tiers, and it is the strongest signal a style sends, so when the type itself must carry the mood or the brand, put the intent into stylePrompt in plain words (the weight, the letterform class, the feeling: "heavy condensed sans, high-energy", "light geometric sans, quiet and premium"); an exact brand face is themeTsx territory, and scenario-text-overlay chooses faces by meaning for the text cards around the captions. Presets also restyle the words themselves: an authoring-time run of tiktok-bouncy uppercased every caption, and the other presets are unverified for casing, so when exact casing matters (a product name, "LoRA") steer with stylePrompt or themeTsx and verify a frame before delivering. fontColor sets the body text (contrast beats aesthetics: white body text survives every backdrop the presets put behind it), and accentColorStart/accentColorEnd drive the highlight animation (karaoke fills, pops): spend the accent on one thing, usually the brand color, with equal values for a solid and different values for a gradient. Auto sizing (fontSizePx empty) rendered words about 30 pixels tall on a 1080x1920 frame at authoring time, unreadable on a phone feed: for vertical social set fontSizePx explicitly (60 to 90 on a 1920-tall frame is the authoring-time starting point) and judge a frame at phone scale. outputTsx: true returns the theme a run used, so a look that landed can be replayed exactly on the next video.
Getting the words right
transcriptionPromptis a spelling hint, not a style field: list the names, brands, and jargon the audio contains. The hint raises the odds without guaranteeing them (an authoring-time run misspelled a hinted name twice), so check the transcript for every required name before trusting a burn-in, and on a miss retry withlarge-v3or a sharper hint.modelSizetrades accuracy for speed and cost (cost_impact): themediumdefault is fine for clean voiceover; step up tolarge-v3for noisy audio, accents, or dense terminology;.envariants are English-only.- The model transcribes whatever audio the master carries, so balance the music bed against dialogue before captioning (
scenario-video-assembly), never after. - Corrections and restyles ride the
subtitlesinput, whose contract is inline content, not a reference (authoring-time): pass the SRT text itself, base64-encoded, as the value. Anasset_...id is not dereferenced there; the id string is base64-decoded as if it were content, and the run still reports success, bills, and renders zero captions (segment_count: 0in the job record is the tell). Reuse is therefore:outputSrt: truereturns the transcript as an asset,asset_downloadit, correct spellings locally if needed, and feed the edited text back base64-encoded, which also coversupload_assethaving no text kind (authoring-time fact). The SRT inherits the run's segmentation (a karaoke run returns word-per-cue), so re-chunk it locally before an ads upload or a calmer restyle. Never deliver asubtitlesrun on job status alone: sweep the output as the worked example reviews it, and when captions are missing,asset_getthe subtitles asset the job consumed to see what it received. The schema does not say whether segmentation caps re-chunk a supplied SRT, so set segmentation when transcribing and omit the caps alongsidesubtitles.
Worked example: a vertical social short, then a Spanish variant
asset_getthe assembled master: confirm duration and that it is the finished cut, since the price tracks the footage (videocarriescost_impact; trim first,scenario-video-editing).model_schema_getonmodel_scenario-caption-studio.- Price it:
model_runwithdry_run: trueandparameters={"video": "<asset_id>", "stylePreset": "tiktok-bouncy", "maxSegmentWords": 4, "textPosition": "middle", "transcriptionPrompt": "Scenario, LoRA, Flux", "outputSrt": true}. Re-estimate after changingtargetLanguage,modelSize,stylePrompt, oroutputSubtitles: all carrycost_impact. - Run it with
wait: false, thenjobs_waitwith thejob_id; on timeout re-call with the returnedpending_job_ids, neverjob_getin a loop. - Review before deriving, since a successful job proves nothing about what rendered: download the SRT sidecar with
asset_download(noformat: it converts images only; a run that returns both video and SRT lists the two asset ids in no fixed order, andasset_gettells them apart bymimeType) and check its text for every name thetranscriptionPromptcarries; then download the video and sweep it into contact sheets (ffmpeg -vf "fps=2"), reading the burned captions for spelling, casing, and placement, because a defect that appears mid-cue survives a spot check. Pop presets reveal words within a cue, so also sample the last frames of each cue: a flash under 100 ms survives anfps=2sweep. - Spanish variant: add
targetLanguage: "es"to the same parameters,transcriptionPromptincluded, anddry_runagain (it moves the price) before running. Translation happens inside the run, so one master yields a variant per market. Latin-script segmentation caps do not transfer to Chinese, Japanese, or Korean (streaming style guides run them at a third of the characters per line), so revisitmaxSegmentCharsper target language. - File the master, variants, and SRT assets in a collection (
scenarioskill) so the delivery set stays findable.
Common mistakes
- Captioning each clip before assembly: caption the finished master once, or cues drift across cuts and every edit orphans its captions.
- Expecting the preset look on the soft track:
video_datarenders as plain text the viewer toggles; styling survives only when burned in. - Styling legal lines or CTAs as captions: disclosures carry locked wording, size, and dwell time that caption logic would re-chunk and retime, and a toggleable track fails "visible without viewer action" rules outright; exact unspoken text is a
scenario-text-overlaycard composited in assembly. - Setting
maxSegmentWords: 1without a karaoke or pop preset: one-word cues flash as a slideshow unless the style animates them. With a pop preset (the ones that reveal words inside a cue:tiktok-bouncy,word-pop, and the karaoke pair at authoring time), never end a cue on a one-letter word: pop-in timing follows character count, so a trailing "I" or "a" showed for about 70 ms at authoring time. Re-chunking a supplied SRT means owning its timings too: a supplied SRT carries cue times only, so a pop preset spreads its words evenly across each cue and the highlight drifts from the voice wherever a cue outlasts the speech; keep each cue's end on the spoken span, and extend only the last cue to the clip's end. - Treating a
jobs_waittimeout as failure: re-call withpending_job_ids; video tool jobs outlast the wait window routinely.
Signals
- GitHub stars
- 681
- Forks
- 82
- Last commit
- Sep 2026
Advanced
- Item type
- skill
- Key
scenario-caption-studio- Source
- github.com/scenario-labs/skills