Scenario Caption Studio

SkillMedia

Use when a video needs its spoken words on screen through Scenario via MCP: burned-in styled captions for a TikTok, Reels, or Shorts cut, ad captions for sound-off feeds, YouTube subtitles, an SRT sidecar, transcription of a clip's audio, captions translated into another language, karaoke or word-by-word styles, or restyling and correcting an existing transcript. Keywords: captions, subtitles, SRT, transcribe, karaoke, word-by-word, burn in, closed captions, translate video, caption style.

Available today. Use it from your connected AI after setup.

Add ahel to your AI once: Claude, ChatGPT, Cursor, Claude Code or Codex. Then ask it to use this.

Then ask your AI: use the Scenario Caption Studio skill

What this skill tells your AI

The instructions your AI receives, as published by scenario-labs/skills in skills/scenario-caption-studio/SKILL.md and read by ahel’s review.

Overview

Caption Studio is one tool model, model_scenario-caption-studio: a video in, its speech transcribed (Whisper) or an existing SRT applied, styled captions out, burned into the picture or delivered as a soft track and an .srt sidecar. It translates into 18 languages and styles captions three ways. Running that one member is this skill's whole purpose, so the id is named rather than discovered and model_schema_get starts the flow directly; availability differs per team, so a member the team lacks is a gap to flag, not a cue to substitute. Connection and the core loop: see the scenario skill.

Captioning is the last pass on a finished cut: assemble first (scenario-video-assembly), then caption the master once. Captions carry the transcript only; text that must appear letter-perfect without being spoken (CTAs, prices, legal supers) is scenario-text-overlay territory. If a sibling skill named here is missing from your available skills, ask the user to install it (npx skills add scenario-labs/skills --skill <name>); unattended, proceed from tool schemas and flag the gap.

Ask once, then caption

The destination decides nearly every parameter, so collect one round of answers before touching the schema: where the video ships (a sound-off mobile feed, a paid placement, a seated long-form viewer), whether the spoken language stays or translates (targetLanguage, auto keeps it), brand colors if any, and which deliverable the platform wants. The deliverable is three switches: burned-in pixels are outputSubtitles: "video_image" (the default), the toggleable track is "video_data", the sidecar file is outputSrt: true, and an SRT-only pass is that plus outputVideo: false, the first pass when the words must be letter-perfect: it priced the same as a burn-in at authoring time, so it buys certainty rather than savings, proving the words before any pixels are paid for. Then map the answers:

DestinationStyleSegmentationPositionOutput
Social mobile short (9:16 Shorts, Reels)tiktok-bouncy or word-pop; karaoke-fill when music drivesmaxSegmentWords 3 to 5; 1 with a karaoke presetmiddle: platform UI and native auto-captions own the bottomBurn in
Ad short, performance cutmodern-chip or minimal-underline, accents set to brand color3 to 7 words per cuebottom, or top when an end card or overlay sits belowBurn in for sound-off feeds; add outputSrt for the platform's caption upload
YouTube long-form, tutorial, interviewDefault look or cinematic-fade; restraint reads as professionalismmaxLines 2, maxSegmentChars 84 (two 42-character lines, the broadcast convention)bottomoutputSrt for the platform's closed captions; burn in only for re-embeds
Cinematic piece, trailer, festival cutcinematic-fadeSentence-length cues, maxSegmentDuration about 6bottomBurn in

Rows are authoring-time starting points to confirm with the user, not platform contracts; unattended, the task's own instructions answer the interview and the matching row's defaults fill what they leave unsaid. The per-placement safe zones behind the position column live in scenario-formats. middle renders at frame center, which on a centered talking head is the mouth: caption a selfie cut only once its face sits in the upper third (reframe in assembly), because the schema offers no lower-third position. An uploaded track beats burn-in for long-form because viewers toggle and restyle it, assistive tech reads it, and platforms index it for search; burn-in wins wherever the style is the point or a track cannot travel with the file.

The style ladder

Three tiers: stylePreset picks a ready-made look (7 presets, empty for the default); stylePrompt describes a look in plain words and builds a matching style (it carries cost_impact); themeTsx supplies a full custom theme that replaces the preset, with stylePrompt then refining that theme. There is no font parameter: type rides inside the tiers, and it is the strongest signal a style sends, so when the type itself must carry the mood or the brand, put the intent into stylePrompt in plain words (the weight, the letterform class, the feeling: "heavy condensed sans, high-energy", "light geometric sans, quiet and premium"); an exact brand face is themeTsx territory, and scenario-text-overlay chooses faces by meaning for the text cards around the captions. Presets also restyle the words themselves: an authoring-time run of tiktok-bouncy uppercased every caption, and the other presets are unverified for casing, so when exact casing matters (a product name, "LoRA") steer with stylePrompt or themeTsx and verify a frame before delivering. fontColor sets the body text (contrast beats aesthetics: white body text survives every backdrop the presets put behind it), and accentColorStart/accentColorEnd drive the highlight animation (karaoke fills, pops): spend the accent on one thing, usually the brand color, with equal values for a solid and different values for a gradient. Auto sizing (fontSizePx empty) rendered words about 30 pixels tall on a 1080x1920 frame at authoring time, unreadable on a phone feed: for vertical social set fontSizePx explicitly (60 to 90 on a 1920-tall frame is the authoring-time starting point) and judge a frame at phone scale. outputTsx: true returns the theme a run used, so a look that landed can be replayed exactly on the next video.

Getting the words right

  • transcriptionPrompt is a spelling hint, not a style field: list the names, brands, and jargon the audio contains. The hint raises the odds without guaranteeing them (an authoring-time run misspelled a hinted name twice), so check the transcript for every required name before trusting a burn-in, and on a miss retry with large-v3 or a sharper hint.
  • modelSize trades accuracy for speed and cost (cost_impact): the medium default is fine for clean voiceover; step up to large-v3 for noisy audio, accents, or dense terminology; .en variants are English-only.
  • The model transcribes whatever audio the master carries, so balance the music bed against dialogue before captioning (scenario-video-assembly), never after.
  • Corrections and restyles ride the subtitles input, whose contract is inline content, not a reference (authoring-time): pass the SRT text itself, base64-encoded, as the value. An asset_... id is not dereferenced there; the id string is base64-decoded as if it were content, and the run still reports success, bills, and renders zero captions (segment_count: 0 in the job record is the tell). Reuse is therefore: outputSrt: true returns the transcript as an asset, asset_download it, correct spellings locally if needed, and feed the edited text back base64-encoded, which also covers upload_asset having no text kind (authoring-time fact). The SRT inherits the run's segmentation (a karaoke run returns word-per-cue), so re-chunk it locally before an ads upload or a calmer restyle. Never deliver a subtitles run on job status alone: sweep the output as the worked example reviews it, and when captions are missing, asset_get the subtitles asset the job consumed to see what it received. The schema does not say whether segmentation caps re-chunk a supplied SRT, so set segmentation when transcribing and omit the caps alongside subtitles.

Worked example: a vertical social short, then a Spanish variant

  1. asset_get the assembled master: confirm duration and that it is the finished cut, since the price tracks the footage (video carries cost_impact; trim first, scenario-video-editing).
  2. model_schema_get on model_scenario-caption-studio.
  3. Price it: model_run with dry_run: true and parameters={"video": "<asset_id>", "stylePreset": "tiktok-bouncy", "maxSegmentWords": 4, "textPosition": "middle", "transcriptionPrompt": "Scenario, LoRA, Flux", "outputSrt": true}. Re-estimate after changing targetLanguage, modelSize, stylePrompt, or outputSubtitles: all carry cost_impact.
  4. Run it with wait: false, then jobs_wait with the job_id; on timeout re-call with the returned pending_job_ids, never job_get in a loop.
  5. Review before deriving, since a successful job proves nothing about what rendered: download the SRT sidecar with asset_download (no format: it converts images only; a run that returns both video and SRT lists the two asset ids in no fixed order, and asset_get tells them apart by mimeType) and check its text for every name the transcriptionPrompt carries; then download the video and sweep it into contact sheets (ffmpeg -vf "fps=2"), reading the burned captions for spelling, casing, and placement, because a defect that appears mid-cue survives a spot check. Pop presets reveal words within a cue, so also sample the last frames of each cue: a flash under 100 ms survives an fps=2 sweep.
  6. Spanish variant: add targetLanguage: "es" to the same parameters, transcriptionPrompt included, and dry_run again (it moves the price) before running. Translation happens inside the run, so one master yields a variant per market. Latin-script segmentation caps do not transfer to Chinese, Japanese, or Korean (streaming style guides run them at a third of the characters per line), so revisit maxSegmentChars per target language.
  7. File the master, variants, and SRT assets in a collection (scenario skill) so the delivery set stays findable.

Common mistakes

  • Captioning each clip before assembly: caption the finished master once, or cues drift across cuts and every edit orphans its captions.
  • Expecting the preset look on the soft track: video_data renders as plain text the viewer toggles; styling survives only when burned in.
  • Styling legal lines or CTAs as captions: disclosures carry locked wording, size, and dwell time that caption logic would re-chunk and retime, and a toggleable track fails "visible without viewer action" rules outright; exact unspoken text is a scenario-text-overlay card composited in assembly.
  • Setting maxSegmentWords: 1 without a karaoke or pop preset: one-word cues flash as a slideshow unless the style animates them. With a pop preset (the ones that reveal words inside a cue: tiktok-bouncy, word-pop, and the karaoke pair at authoring time), never end a cue on a one-letter word: pop-in timing follows character count, so a trailing "I" or "a" showed for about 70 ms at authoring time. Re-chunking a supplied SRT means owning its timings too: a supplied SRT carries cue times only, so a pop preset spreads its words evenly across each cue and the highlight drifts from the voice wherever a cue outlasts the speech; keep each cue's end on the spoken span, and extend only the last cue to the clip's end.
  • Treating a jobs_wait timeout as failure: re-call with pending_job_ids; video tool jobs outlast the wait window routinely.

Signals

GitHub stars
681
Forks
82
Last commit
Sep 2026
Advanced
Item type
skill
Key
scenario-caption-studio
Source
github.com/scenario-labs/skills