Analyzing narrated recordings with talkthrough
SkillFiles & storageAnalyze narrated screen recordings and audio files through the talkthrough MCP server — triage feedback into findings, extract specs/backlogs/action items from recordings, and correlate spoken remarks with logs via wall-clock timestamps. Use when the user mentions a screen recording, screencast, narrated video/audio file, or asks to "watch" a recording and act on it.
Available today. Use it from your connected AI after setup.
No other account needed.
Connect ahel once, and every AI you use reads what you have installed.
Then ask your AI: use the Analyzing narrated recordings with talkthrough skill
What this skill tells your AI
The instructions your AI receives, as published by korovin-aa97/talkthrough-mcp in .agents/skills/talkthrough/SKILL.md and read by ahel’s review.
The talkthrough MCP server turns a local video/audio file into queryable structured data: timestamped transcript segments, scene keyframes, OCR'd on-screen text, and wall-clock anchoring. No LLM inside — you bring the reasoning; it brings the evidence. Everything is lazy and token-budgeted: never ask for more than the moment you are analyzing.
Prerequisite
The talkthrough MCP server must be connected (tools like
process_media / get_transcript are visible). If not, tell the user to
install it: claude mcp add -s user talkthrough -- uvx --python ">=3.11,<3.14" "talkthrough-mcp[diarization,url]"
(see the repository README for other clients).
Core workflow
- Ingest once:
process_media(path)— idempotent by content hash; re-calls on the same file return instantly. Given a public video/audio URL instead of a file, callprocess_url(url): the source is downloaded once (the only network step; YouTube needs the[url]extra) and kept inside the job, then everything below is identical and local — never download twice; a repeat call serves the stored job. Long videos take minutes and stream progress. The summary gives youjob_id, counts, wall-clock, and a transcript preview — do NOT dump anything else eagerly. Multi-person recording (meeting/interview)? Adddiarize=true— even when the ask is just "summarize", speaker structure is part of meeting analysis — and — whenever the headcount is known —num_speakers=N(the main accuracy lever): segments getS1/S2/… labels and the summary a talk-time roster. On an already-processed job the amend re-runs ONLY diarization (no re-transcription) — still minutes on long recordings. - Orient:
get_transcript(job_id)(paginate vianext_start_mswhentruncated) orsearch(job_id, "<distinctive word>")to jump straight to the relevant moments (searches speech AND on-screen OCR text). Multi-word search defaults tomatch_mode="all_words"; use"any_word"for broader lexical recall. - Evidence per remark:
get_moment(job_id, t0-2000, t1+2000)— one call returns the transcript slice + up to 3 unique frames + their OCR text + the wall-clock range. This is the workhorse; describeobservedfrom the returned pixels, never from imagination. - Precision when needed:
get_frames(at_ms=...)for nearby keyframes;extract_frame(job_id, at_ms, crop={x,y,w,h})for an exact instant at native resolution (keyframes capture scene changes + a 1 fps floor, so sub-second moments can fall between them). - Keep verified names: after proving an anonymous label's identity,
call
label_speakers(job_id, labels={"S1":"Name"}, evidence={"S1":"intro or frame proof"}). Saved names appear in later transcript, moment, and search calls while rawS1/S2labels remain. If a diarization amend changes labels, those names move tospeaker_names_pending_reviewand stop being identities. Use the stored old-roster context anchors to re-check them. A pending label still in the roster can be confirmed, replaced, or removed; a stale pending label can only be removed withlabels={"Sx":null}. Never use a pending name in minutes or search as though it were active. A fullforce=truerebuild of a job with active or pending names must also usediarize=true; it rebuilds safely and moves every old identity to pending review, while omitting diarization is refused without changing the stored job. - Recall across sessions:
list_jobs()— the store persists; a file processed yesterday (even via CLI) is queryable byjob_idtoday.
Timestamps
Every timestamped result carries t_ms (video-relative) and, when the
recording start is known, t_wall (ISO 8601 real time). Copy t_wall
VERBATIM from the payload — never compute it from t_ms yourself
(hand-derived wall-clocks drift by whole hours). Use t_wall to
correlate remarks with server/app logs (±30 s grep window). If
wall_clock is null or low-confidence, ask the user when the recording
started and re-anchor: process_media(path, recorded_at="<ISO 8601>", force=true); when the job already has speaker identities, include
diarize=true as required by the safe-rebuild contract.
Packaged workflows (server prompts)
Prefer the server prompts when the task matches — they encode the full
method: bug (one recording → evidence-backed GitHub issue draft; silent,
narration-free recordings welcome), triage-recording (screencast →
findings JSON per the contract in examples/output-contract.schema.json),
spec-from-workshop, backlog-from-demo, meeting-actions (audio-only
friendly), correlate-with-logs.
Rules of thumb
- Audio-only jobs (.m4a/.mp3/…): transcript tools work; frame tools error by design — that error is expected, not a failure.
- Speaker labels are anonymous (
S1/S2, ordered by first voice). Mapping them to names is YOUR job: self-introductions, vocatives, the attendees list — and on video jobs the screen check is MANDATORY: for every label you map,get_frames(at_ms=<that label's longest_turn_at_ms from the roster>)and read the meeting-app name plates, the recording's title card, the active-speaker highlight BEFORE asserting the mapping. STT homophones lie about name spellings (spoken "profit" vs on-screen "Prophet") — trust OCR/frames over the transcript for names. State the mapping explicitly and mark unmapped labels "unidentified". Rostername_candidatesare raw OCR hints, not identities: they may be UI text, a job title, or somebody else's name. Inspect the cited frame and persist only defensible mappings withlabel_speakers; never auto-save a candidate. Aname_candidates_noteon a pre-0.3.1 video job explains that its legacy flat OCR may not yield hints. The job remains readable; regenerate only when useful, withforce=true, diarize=true, so old identities become pending review instead of being lost.diarize=trueneeds the[diarization]extra — its absence produces an actionable install-hint error. - Findings/quotes must cite the narrator's exact words +
t_ms(+t_wallwhen known) + the frame files you actually inspected. - Low STT/vision confidence → surface a question; never silently guess.
- Any narration language works (Whisper auto-detects; the summary reports
language+language_probability). Garbled transcript or low/wrong detection → re-callprocess_media(path, model="large-v3-turbo", force=true)(best multilingual quality) or pinlanguage="…"; domain jargon → passvocabulary="Term1, Term2". - Write digests/summaries for the recording author in the narrator's language; keep quotes verbatim in the original — translate in your own prose only, never inside a quote.
Signals
- GitHub stars
- 29
- Forks
- 4
- Last commit
- Sep 2026
Advanced
- Catalog kind
- skill
- Gateway key
talkthrough-korovin-aa97- Source
- github.com/korovin-aa97/talkthrough-mcp