Transcribe Video
SkillFiles & storageUse when the user wants to transcribe a video or audio file to text with word-level timestamps — spins up a self-contained local Whisper (whisper.cpp) Docker service bundled in this skill and returns transcript.txt, transcript.srt, and word-level transcript.words.json. Also used as a building block by /produce-video. Triggers on "transcribe this video", "get a transcript", "transcribe-video".
Available today. Use it from your connected AI after setup.
No other account needed.
Connect ahel once, and every AI you use reads what you have installed.
Then ask your AI: use the Transcribe Video skill
What this skill tells your AI
The instructions your AI receives, as published by alemtuzlak/skills in skills/transcribe-video/SKILL.md and read by ahel’s review.
Transcribes any video/audio file locally using a bundled whisper.cpp service. No external project or cloud API. Returns word-level timestamps (needed for synced overlays/captions).
When to use
- "Transcribe this clip / video / audio", "get me a transcript with timestamps".
- As a sub-step of
/produce-video.
How it runs
Everything is driven by the bundled runner — never ask the user to manage Docker:
node scripts/transcribe.mjs <path-to-video-or-audio> [--out <dir>] [--port 9111] [--language en] [--no-word-ts] [--task transcribe|translate]
node scripts/transcribe.mjs --stop # stop the warm container
The runner:
- Checks the Docker daemon is reachable (fails loud with guidance if not).
- Builds the
transcribe-video-whisperimage fromassets/whisper-service/if it is missing (first build is slow: it compiles whisper.cpp and bakes theggml-base.en.binmodel). - Starts the
transcribe-video-whispercontainer (host port 9111 → container 9001); reuses it if already running,docker starts it if stopped. - Waits for
GET /healthzto report ok. - Extracts a 16 kHz mono WAV from the input with local ffmpeg (whisper.cpp only decodes WAV), then POSTs it to
/transcribewithword_ts=true. - Writes
transcript.txt,transcript.srt,transcript.words.jsonto--out(default: the input file's directory), and prints a JSON result line to stdout.
Outputs
transcript.txt— plain text.transcript.srt— subtitle text.transcript.words.json—[{ "word", "start", "end" }], seconds, time-ordered. This is the sync source for overlays/captions.
stdout result line (for programmatic callers like /produce-video):
{ "ok": true, "outDir": "...", "wordCount": 38, "segmentCount": 2, "files": { "txt": "transcript.txt", "srt": "transcript.srt", "words": "transcript.words.json" } }
Requirements
- Docker is REQUIRED — the Whisper service runs in a container, so Docker Desktop must be installed AND running. The runner checks this first and fails loud with install/start guidance if Docker is missing or its daemon is down; there is no non-Docker fallback.
- ffmpeg is REQUIRED (used to extract a 16 kHz mono WAV for Whisper) — the runner checks for it up front and fails loud if it is not on PATH.
- Node ≥ 22.
Notes
- The container is left running (
--restart unless-stopped) for warm reuse; stop with--stop. - If word timestamps are requested but none come back, the runner fails loud (overlay sync depends on them).
- See
references/for the Docker lifecycle, the API shape, and the vendoring provenance.
Signals
- GitHub stars
- 39
- Forks
- 1
- Last commit
- Aug 2026
Advanced
- Catalog kind
- skill
- Gateway key
transcribe-video- Source
- github.com/alemtuzlak/skills